Measured before claimed.
The benchmark is not marketing — it is how the platform is governed. Champions are elected by measurement, weaknesses are hunted on purpose, and a number is published only after it survives pre-registered evaluation.
Honest states, not empty promises.
Most numbers here are in the publication pipeline. Rather than advertise internal results, each dimension shows its true state — the same discipline the platform applies to knowledge.
| State | Meaning |
|---|---|
| published | survived pre-registered evaluation; number, n and confusion matrix available |
| measured · internal | measured against internal gold sets; publication pending external validation |
| publication pending | gold set or evaluation in construction |
| UNKNOWN · fail-closed | the capability abstains in production rather than guess |
The parser, dimension by dimension.
| Dimension | Unit | Status | Notes |
|---|---|---|---|
| Reading order | F1 vs gold | measured · internal | publication after pre-registered eval |
| Tables | structure F1 | measured · internal | gold set expanding — n reported with every number |
| OCR | CER / WER | measured · internal | champions measured per document family |
| Forms | field F1 | publication pending | gold set in construction |
| Long documents | fidelity at 60+ pages | publication pending | long-document bench in progress |
| Charts | chart-to-table | UNKNOWN · fail-closed | abstains instead of guessing |
Every product is a measured surface.
| Dimension | Unit | Status | Notes |
|---|---|---|---|
| Audio | canonical audio pipeline | measured · internal | acquisition → canonical, certified |
| Canonical determinism | identical re-runs | measured · internal | byte-level round-trip checks |
| Chunking | structure alignment | publication pending | projections measured against gold |
| Retrieval | grounded recall | publication pending | measured on governed knowledge |
| Claim extraction | precision / recall | measured · internal | abstention counted, never hidden |
| Reasoning grounding | answers traceable to claims | measured · internal | adversarial sets included |
| Latency | p50 / p95 per unit | publication pending | per product, per deployment |
| Cost | per 1k units | publication pending | published with credit rates |
Designed so we cannot flatter ourselves.
Every vendor wins their own benchmark. Ours is built to find our weaknesses first.
Pre-registered
Metrics, thresholds and datasets are fixed before the run. No metric shopping after the fact.
Gold is corrigible
Gold sets are scientific artifacts: versioned, adjudicated, corrected through a recorded process — never silently edited.
Abstention is scored
UNKNOWN is a measured outcome. A system that hides its abstentions is optimizing the wrong thing.
Champions re-elected
Every capability runs a permanent championship. Champions keep their place only while they win — including ours.
Scoped numbers
No naked scores: every published number carries its n, its dataset version and its confusion matrix.
Detection proven
Evaluation layers are validated by deliberate fault injection — we prove the benchmark can catch failure before trusting its green.
A benchmark that never stops.
Every release re-enters the championship. Regressions are caught by the benchmark, not by customers.
Gold sets grow by adjudicated annotation. Coverage gaps are tracked as debt, visibly.
Published results with datasets and methodology, verifiable end to end.