Benchmarks

Measured before claimed.

The benchmark is not marketing — it is how the platform is governed. The best measured system ships, weaknesses are hunted on purpose, and a number is published only after it survives pre-registered evaluation.

pre-registered metrics versioned gold sets n always reported abstention counted
Reading this page

Honest states, not empty promises.

Most numbers here are still in publication. Rather than advertise internal results, each dimension shows its true state — the same discipline the platform applies to knowledge.

StateMeaning
publishedsurvived pre-registered evaluation; number, n and confusion matrix available
measured · internalmeasured against internal gold sets; publication pending external validation
publication pendinggold set or evaluation in construction
UNKNOWN · fail-closedthe capability abstains in production rather than guess
Document parsing

Parsing, dimension by dimension.

DimensionUnitStatusNotes
Reading orderF1 vs goldmeasured · internalpublication after pre-registered eval
Tablesstructure F1measured · internalgold set expanding — n reported with every number
Text recognition accuracyCER / WERmeasured · internalmeasured per document family
Formsfield F1publication pendinggold set in construction
Long documentsfidelity at 60+ pagespublication pendinglong-document bench in progress
Chartschart-to-tableUNKNOWN · fail-closedabstains instead of guessing
Beyond parsing

Every product is a measured surface.

DimensionUnitStatusNotes
Audiotranscription fidelitymeasured · internalacquisition → evidence, certified
Determinismidentical re-runsmeasured · internalbyte-level round-trip checks
Structure alignmentoutput vs source structurepublication pendingmeasured against gold
Retrievalgrounded recallpublication pendingmeasured on governed knowledge
Claim extractionprecision / recallmeasured · internalabstention counted, never hidden
Reasoning groundinganswers traceable to claimsmeasured · internaladversarial sets included
Latencyp50 / p95 per unitpublication pendingper product, per deployment
Costper 1k unitspublication pendingpublished with credit rates
Method

Designed so we cannot flatter ourselves.

Every vendor wins their own benchmark. Ours is built to find our weaknesses first.

Pre-registered

Metrics, thresholds and datasets are fixed before the run. No metric shopping after the fact.

Scoped numbers

No naked scores: every published number carries its dataset and version, sample size (n), inclusion criteria, the comparators it was measured against, its run date, and known limitations.

Independent gold

Gold sets are held independent of the systems under test — never derived from the same process that produces the results being scored.

Gold is corrigible

Gold sets are scientific artifacts: versioned, adjudicated, corrected through a recorded process — never silently edited.

Abstention is scored

UNKNOWN is a measured outcome. A system that hides its abstentions is optimizing the wrong thing.

Reproducible evaluation

The evaluation itself is reproducible: same datasets, same metrics, same protocol — you can re-run the comparison, not just read the score.

Continuously re-measured

Every capability is re-measured as methods improve. The best measured result ships — and only keeps its place while the numbers hold, including ours.

Detection proven

Evaluation layers are validated by deliberate fault injection — we prove the benchmark can catch failure before trusting its green.

Method, not mechanism

What we publish is the proof, not the machine. Datasets, metrics and results are open; the implementation that achieves them is not part of what you audit — the number is.

Continuity

A benchmark that never stops.

continuous
New product version → automatic re-measurement

Every release is re-measured against the benchmark. Regressions are caught by the benchmark, not by customers.

continuous
New gold data → wider coverage

Gold sets grow by adjudicated annotation. Coverage gaps are tracked as debt, visibly.

upcoming
Public leaderboard

Published results with datasets and methodology, verifiable end to end.

Build AI systems that know where every answer came from.

Create your account with 15,000 free credits and generate your API keys today. ORIS is in private pilot; VERI and MESH open at launch.