Benchmarks

Measured before claimed.

The benchmark is not marketing — it is how the platform is governed. Champions are elected by measurement, weaknesses are hunted on purpose, and a number is published only after it survives pre-registered evaluation.

pre-registered metrics versioned gold sets n always reported abstention counted
Reading this page

Honest states, not empty promises.

Most numbers here are in the publication pipeline. Rather than advertise internal results, each dimension shows its true state — the same discipline the platform applies to knowledge.

StateMeaning
publishedsurvived pre-registered evaluation; number, n and confusion matrix available
measured · internalmeasured against internal gold sets; publication pending external validation
publication pendinggold set or evaluation in construction
UNKNOWN · fail-closedthe capability abstains in production rather than guess
Document parsing

The parser, dimension by dimension.

DimensionUnitStatusNotes
Reading orderF1 vs goldmeasured · internalpublication after pre-registered eval
Tablesstructure F1measured · internalgold set expanding — n reported with every number
OCRCER / WERmeasured · internalchampions measured per document family
Formsfield F1publication pendinggold set in construction
Long documentsfidelity at 60+ pagespublication pendinglong-document bench in progress
Chartschart-to-tableUNKNOWN · fail-closedabstains instead of guessing
Beyond parsing

Every product is a measured surface.

DimensionUnitStatusNotes
Audiocanonical audio pipelinemeasured · internalacquisition → canonical, certified
Canonical determinismidentical re-runsmeasured · internalbyte-level round-trip checks
Chunkingstructure alignmentpublication pendingprojections measured against gold
Retrievalgrounded recallpublication pendingmeasured on governed knowledge
Claim extractionprecision / recallmeasured · internalabstention counted, never hidden
Reasoning groundinganswers traceable to claimsmeasured · internaladversarial sets included
Latencyp50 / p95 per unitpublication pendingper product, per deployment
Costper 1k unitspublication pendingpublished with credit rates
Method

Designed so we cannot flatter ourselves.

Every vendor wins their own benchmark. Ours is built to find our weaknesses first.

Pre-registered

Metrics, thresholds and datasets are fixed before the run. No metric shopping after the fact.

Gold is corrigible

Gold sets are scientific artifacts: versioned, adjudicated, corrected through a recorded process — never silently edited.

Abstention is scored

UNKNOWN is a measured outcome. A system that hides its abstentions is optimizing the wrong thing.

Champions re-elected

Every capability runs a permanent championship. Champions keep their place only while they win — including ours.

Scoped numbers

No naked scores: every published number carries its n, its dataset version and its confusion matrix.

Detection proven

Evaluation layers are validated by deliberate fault injection — we prove the benchmark can catch failure before trusting its green.

Continuity

A benchmark that never stops.

continuous
New product version → automatic re-measurement

Every release re-enters the championship. Regressions are caught by the benchmark, not by customers.

continuous
New gold data → wider coverage

Gold sets grow by adjudicated annotation. Coverage gaps are tracked as debt, visibly.

upcoming
Public leaderboard

Published results with datasets and methodology, verifiable end to end.

Build AI systems that know where every answer came from.

Free sandbox — real API, sample corpus, verify a real certificate. Buy credits when it earns it.