We publish what we measure.
Orves is run as a measurement program: pre-registered evaluations, versioned gold sets, adversarial review. What survives becomes public — as benchmarks, methodology and datasets.
Why every AI stack will become canonical.
Models are converging. The durable differentiator is not the model — it is what the model is allowed to know, and whether anyone can trust it.
Improvised knowledge does not survive contact with responsibility. Ad-hoc chunks cannot say where an answer came from, cannot be rebuilt after a fix, and cannot enter a regulated decision.
So the substrate gets rebuilt — once, properly: canonical representation, permanent identity, provenance, determinism, verifiable transformation. That layer is what Orves builds as infrastructure.
What the research program runs.
| Line | What it measures | Status |
|---|---|---|
| Parsing benchmarks | reading order, tables, OCR, forms, long documents — per document family | measured · internal |
| Canonical determinism | byte-level round trips, reconstruction fidelity | measured · internal |
| Grounding & abstention | claims traceable to evidence; refusing when evidence is insufficient | measured · internal |
| Continuous acquisition | observation→update funnel, admission rubrics | measured · internal |
| Public leaderboard & datasets | published results, verifiable end to end | publication pending |
| Papers & technical reports | methodology, gold-set construction, adjudication | soon |
Current honest states, dimension by dimension: Benchmarks →
Designed so we cannot flatter ourselves.
Pre-registered metrics. Versioned, corrigible gold sets. Abstention scored. Detection power proven by deliberate fault injection. Every published number carries its n.