Research

We publish what we measure.

Orves is run as a measurement program: pre-registered evaluations, versioned gold sets, adversarial review. What survives becomes public — as benchmarks, methodology and datasets.

The thesis

Why every AI stack will become canonical.

Models are converging. The durable differentiator is not the model — it is what the model is allowed to know, and whether anyone can trust it.

Improvised knowledge does not survive contact with responsibility. Ad-hoc chunks cannot say where an answer came from, cannot be rebuilt after a fix, and cannot enter a regulated decision.

So the substrate gets rebuilt — once, properly: canonical representation, permanent identity, provenance, determinism, verifiable transformation. That layer is what Orves builds as infrastructure.

Program

What the research program runs.

LineWhat it measuresStatus
Parsing benchmarksreading order, tables, OCR, forms, long documents — per document familymeasured · internal
Canonical determinismbyte-level round trips, reconstruction fidelitymeasured · internal
Grounding & abstentionclaims traceable to evidence; refusing when evidence is insufficientmeasured · internal
Continuous acquisitionobservation→update funnel, admission rubricsmeasured · internal
Public leaderboard & datasetspublished results, verifiable end to endpublication pending
Papers & technical reportsmethodology, gold-set construction, adjudicationsoon

Current honest states, dimension by dimension: Benchmarks →

Method

Designed so we cannot flatter ourselves.

Pre-registered metrics. Versioned, corrigible gold sets. Abstention scored. Detection power proven by deliberate fault injection. Every published number carries its n.

Build AI systems that know where every answer came from.

Free sandbox — real API, sample corpus, verify a real certificate. Buy credits when it earns it.