How claims are checked before they appear here.
Product facts come from generated release records. Research claims point to their test data and limits. When a published trial or public claim is corrected, the prior record and correction remain available.
How a claim gets on this site
Every number renders from a single source file (src/lib/facts.ts), where each constant carries a pointer to the artifact it was measured from. Three gates run in CI when their affected surfaces change and during explicit release validation: a claims gate that fails the build if a correlation appears without its label basis or if a banned overclaim phrase appears anywhere in the source (the banned list is enforced to stay identical to the one in Bulla’s own test suite); a facts gate that fails if a constant drifts from the Bulla source it cites; and a dead-route gate for links. The site cannot state a claim the package does not back.
What this check cannot tell you
The gates reject registered drift and known relational overclaims; they do not prove that every possible misleading formulation has been anticipated.
Validated
- The mathematics. The coherence fee is a genuine coboundary-rank invariant with an exact additive decomposition over convention dimensions. The repository includes machine-checked Lean proofs for the stated witness-geometry results. The Per-Dimension Additivity Theorem is also exercised on all 703 corpus compositions. The PyPI package does not vendor Lean; it implements the measurement and receipt layers in Python.
What this check cannot tell you
Machine checking applies to the stated formal object and hypotheses; it does not establish empirical identification, execution truth, or operational recourse.
- The wire format. ActionReceipt v0.2 has a zero-dependency checker (standard library only, no Bulla imports) that reproduces all four hashes and the modality rule against 17 published examples. Those examples let another implementation test its reading of the specification.
What this check cannot tell you
The browser and packaged checker agree on the published examples. That does not show how a separate implementation or live system will behave.
Technical evidence
Passing the supplied vectors establishes fixture parity, not independent implementation or live-system correctness.
- The calibration corpus. The calibration corpus (38 servers, 703 compositions) is frozen at bulla 0.33.0.
What this check cannot tell you
The frozen calibration labels are schema-derived annotations; their correlation with the coherence fee is not execution-derived and does not estimate runtime failure.
- The benchmark. BABEL ships frozen dev/test/hidden splits with method-neutral ground truth; public scores come from its reproducibility split, not the hidden official-ranking holdout.
What this check cannot tell you
Public scores are frozen dev-split reference results; official ranking uses a hidden holdout and controlled benchmark performance is not production safety.
Tested and falsified
Bulla was once described as if the coherence fee predicted execution failure — “catches this before execution.” A pre-registered, execution-labelled kill-test checked that claim: a fee-vs-baselines battery scored by real Python round-trips the predictors never see. The fee’s fire was approximately independent of whether a break occurred — likelihood ratio ≈ 1.07, against 1.0 for no information. Two supporting negatives, kept distinct: a pre-registered result that on the real corpus a cheap depth-3 baseline recovers the full obstruction, and an execution-labelled probe in which fee=0 compositions breached under real file I/O (30 of 36 — on an authored tool set, so evidence the blind spot is broad, not a corpus estimate). The claim was retired; the full record is in FALSIFICATIONS.md.
What replaced it: the fee is a disclosure measure — how much convention two composed tools leave undisclosed at their seam — and one field a receipt can carry. The record and recourse layer (receipt, registry, coverage, retention asymmetry) never depended on the fee predicting anything. The falsification retired a product claim, not the mathematics and not the receipt.
What this check cannot tell you
The coherence fee measures undisclosed conventions in a pinned composition model; it is not Bulla's safety foundation or an execution-failure predictor.
Open
- Independent retention. The verification ladder has three levels: file integrity, an authenticated issuer signature, and proof that a separate log retained the receipt. The published checker covers the first level. The later levels require accepted keys and an actual log.
- A second, independently operated log. Today there is one registry, operated by Glyph Standard. The standing model (ADR-001) commits to any-log verification and records this as its named closing condition — until a second log exists, log plurality is a commitment, not a fact.
What this check cannot tell you
An operated composition-deed log is not independent ActionReceipt witness plurality; the current independent-witness count is zero.
- Evidence-source classes. Spec v0.2 (NORMATIVE) labels how each evidence item is supported — self_asserted, counterparty_signed, third_party_anchored, execution_verified — so a receipt’s displayed strength is the minimum over its necessary evidence. The classes are normative in the shipped v0.2 wire; each evidence item still needs support for the class it claims.