Home/The Attestation Benchmark
Open benchmark / v0.1, published Sep 2026

We published the benchmark. Including where we lose.

Every vendor in this category publishes an extraction accuracy figure. None of them publishes whether the value can be defended — whether the citation points at the right place, whether the system declined when it should have, and whether an auditor could reconstruct the trail without calling anyone. So we wrote that measure, opened the methodology, and scored ourselves on it first.

Open methodology Corpus available on request Any vendor may run it Our weakest measure named
01 / why a new measure

An extraction score tells you the value was read. It does not tell you it can be relied on.

The published parsing benchmarks in this category are good at what they measure, and what they measure is whether the characters came off the page correctly. That was the binding problem in 2023. It is not the binding problem now.

The gap

A correct value with a wrong citation

Systems that produce a page reference are rarely measured on whether the reference is right. A field that reads correctly but cites the wrong page fails the only test that matters in an examination, and passes every extraction benchmark.

The gap

Confidence with nothing behind it

Accuracy averaged across a corpus hides the number a risk owner actually needs: of the values the system was confident enough to pass through without review, how many were wrong. That is the only population that reaches a client or a regulator unchecked.

The gap

No credit for saying “I don’t know”

A confident error costs far more than an abstention, and every headline accuracy figure punishes the system that abstains correctly. Until abstention is scored as a skill, vendors are rewarded for guessing.

Why we are publishing something we can be beaten on

Because a measure only means anything if the vendor who wrote it can lose on it. We build for defensibility, so we expect to do well here — but the methodology is open precisely so that claim can be checked rather than believed, and so a buyer can hold every vendor including us to the same test. If somebody beats us on it, that is a working benchmark, not a failed launch.

02 / the five measures

Five measures, each tied to a question somebody will be asked under oath.

Each is computed per field, then aggregated per document family. There is no single headline number, deliberately — a composite score would let a weak measure hide behind a strong one.

MeasureWhat is computedThe question it answers
Citation correctnessShare of asserted values whose cited page and region actually contain the value, verified against the labelled source“Show me where this number came from.”
Gate precisionError rate within the population that cleared the confidence gate and went through without human review“How many wrong values reached a client unchecked?”
Abstention correctnessBoth directions: correct declines on genuinely ambiguous fields, and wrong declines on fields that were legible“Does it know what it does not know?”
Reviewer agreementAgreement between two independent qualified reviewers and the produced record, on a sampled subset“Would an expert have recorded the same thing?”
Evidence completenessShare of sampled fields for which a full trail — source, version, transform, gate decision, approver — can be reconstructed with no human involvement“Reconstruct this record for the examiner.”
Deliberately excluded: cost per page and pages per second. Both are real operational numbers and neither belongs in a defensibility measure. Publishing them alongside these five would invite exactly the trade that this benchmark exists to expose — a cheaper, faster run that clears the gate more often because the gate was lowered. We report throughput and cost to customers, on their own estate, in the evaluation report.
03 / the corpus and the method

Published in full, so a result can be reproduced rather than trusted.

A benchmark whose corpus is private is a marketing asset. These are the parts we release, and the two limits we would rather state than have found.

Released

The document set

Synthetic and consented documents across the two shipping packs — wealth-management onboarding and insurance submissions — built to reproduce the failure modes real estates contain: superseded versions, handwritten amendments, scanned inserts, and values that appear twice with different meanings.

Released

The labels and the labelling guide

Field-level ground truth with the adjudication rules that produced it, so a disagreement about a score can be settled by reading the guide rather than by arguing about intent.

Released

The scoring harness

The scripts that compute all five measures from a vendor-neutral output format, plus the adapter we used for our own runs. If your system emits fields and citations, it can be scored.

Stated limit

We wrote it, and it is v0.1

A benchmark authored by a vendor deserves scepticism. Two document families, one author, first publication. It becomes credible when other people run it and argue with it — which is the entire reason the corpus and harness ship with it.

Stated limit

It is not your estate

No public corpus predicts your numbers. This benchmark makes vendors comparable to each other; the evaluation on your own documents is what tells you what you would actually get. Use both, in that order.

Versioning

Frozen and dated

Each version is frozen on publication and every result carries the version it was produced against. Corpus changes ship as a new version with a changelog, alongside the runtime releases — scores are never restated retroactively.

04 / our scores

Forge v1.0.0 against Attestation Benchmark v0.1.

Self-run, on the published corpus, with the harness released above. These are gothink measuring gothink — the correct amount of trust to place in them is “enough to check”.

MeasureWealth onboardingInsurance submissionsWhere this is weakest
Citation correctnessMulti-page tables where a value is restated in a summary
Gate precisionHandwritten amendments that clear the gate on a clean scan
Abstention correctnessOver-abstains on legible but unusual clause layouts
Reviewer agreementFields where two reviewers disagree with each other first
Evidence completenessRecords that crossed a schema version boundary mid-run

Publish the run before you publish the number

The figures above are withheld until the first run is complete and independently witnessed — putting numbers here that nobody outside gothink has seen would repeat the practice this page exists to replace. The measures, the weaknesses and the method are published now because those are the parts a buyer can hold us to. The scores land on the changelog the day the witnessed run finishes, with the raw outputs attached.

No competitor scores appear on this page, and none will be published by us. A vendor scoring its rivals is a marketing exercise regardless of how careful the method is. The harness is public so that buyers, analysts and the vendors themselves can produce those comparisons. We will link to any run of v0.1 that publishes its outputs, including runs where Forge comes second.
05 / running it

Three ways to use this, depending on who you are.

If you are buying

Put it in the RFP

Ask every shortlisted vendor for their five numbers against v0.1 and their raw outputs. It is a fair test, none of them can tune to a corpus they did not choose, and the answers will separate the field faster than a feature matrix. Take the wording straight from this page — you do not need our permission and we would rather you did not ask for it.

Request the corpus and harness
If you are a vendor

Run it and publish

The corpus, labels, adjudication guide and scoring harness are yours under a permissive licence, including for competitive marketing against us. The only condition is that a published score carries the version and the raw outputs. Disagreements about the method are welcome and belong in the changelog.

Tell us you are running it
If it is your own build

Score what you already have

Most teams that built extraction in-house have never measured gate precision or evidence completeness, because nothing prompted them to. Run the harness against your own pipeline before you decide whether to buy anything. If you score well, you have saved a procurement cycle; if you do not, you now know which two layers are missing.

The five layers, and what each one adds
The comparison that actually decides it / your estate

The benchmark makes vendors comparable. Your documents decide it.

A public corpus is the right way to shortlist and the wrong way to choose. Once the shortlist is down to two, run twenty-five of your own documents through each — labelled by your people, scored on the measures above, on the estate you actually have to defend.