Every vendor in this category publishes an extraction accuracy figure. None of them publishes whether the value can be defended — whether the citation points at the right place, whether the system declined when it should have, and whether an auditor could reconstruct the trail without calling anyone. So we wrote that measure, opened the methodology, and scored ourselves on it first.
The published parsing benchmarks in this category are good at what they measure, and what they measure is whether the characters came off the page correctly. That was the binding problem in 2023. It is not the binding problem now.
Systems that produce a page reference are rarely measured on whether the reference is right. A field that reads correctly but cites the wrong page fails the only test that matters in an examination, and passes every extraction benchmark.
Accuracy averaged across a corpus hides the number a risk owner actually needs: of the values the system was confident enough to pass through without review, how many were wrong. That is the only population that reaches a client or a regulator unchecked.
A confident error costs far more than an abstention, and every headline accuracy figure punishes the system that abstains correctly. Until abstention is scored as a skill, vendors are rewarded for guessing.
Because a measure only means anything if the vendor who wrote it can lose on it. We build for defensibility, so we expect to do well here — but the methodology is open precisely so that claim can be checked rather than believed, and so a buyer can hold every vendor including us to the same test. If somebody beats us on it, that is a working benchmark, not a failed launch.
Each is computed per field, then aggregated per document family. There is no single headline number, deliberately — a composite score would let a weak measure hide behind a strong one.
| Measure | What is computed | The question it answers |
|---|---|---|
| Citation correctness | Share of asserted values whose cited page and region actually contain the value, verified against the labelled source | “Show me where this number came from.” |
| Gate precision | Error rate within the population that cleared the confidence gate and went through without human review | “How many wrong values reached a client unchecked?” |
| Abstention correctness | Both directions: correct declines on genuinely ambiguous fields, and wrong declines on fields that were legible | “Does it know what it does not know?” |
| Reviewer agreement | Agreement between two independent qualified reviewers and the produced record, on a sampled subset | “Would an expert have recorded the same thing?” |
| Evidence completeness | Share of sampled fields for which a full trail — source, version, transform, gate decision, approver — can be reconstructed with no human involvement | “Reconstruct this record for the examiner.” |
A benchmark whose corpus is private is a marketing asset. These are the parts we release, and the two limits we would rather state than have found.
Synthetic and consented documents across the two shipping packs — wealth-management onboarding and insurance submissions — built to reproduce the failure modes real estates contain: superseded versions, handwritten amendments, scanned inserts, and values that appear twice with different meanings.
Field-level ground truth with the adjudication rules that produced it, so a disagreement about a score can be settled by reading the guide rather than by arguing about intent.
The scripts that compute all five measures from a vendor-neutral output format, plus the adapter we used for our own runs. If your system emits fields and citations, it can be scored.
A benchmark authored by a vendor deserves scepticism. Two document families, one author, first publication. It becomes credible when other people run it and argue with it — which is the entire reason the corpus and harness ship with it.
No public corpus predicts your numbers. This benchmark makes vendors comparable to each other; the evaluation on your own documents is what tells you what you would actually get. Use both, in that order.
Each version is frozen on publication and every result carries the version it was produced against. Corpus changes ship as a new version with a changelog, alongside the runtime releases — scores are never restated retroactively.
Self-run, on the published corpus, with the harness released above. These are gothink measuring gothink — the correct amount of trust to place in them is “enough to check”.
| Measure | Wealth onboarding | Insurance submissions | Where this is weakest |
|---|---|---|---|
| Citation correctness | — | — | Multi-page tables where a value is restated in a summary |
| Gate precision | — | — | Handwritten amendments that clear the gate on a clean scan |
| Abstention correctness | — | — | Over-abstains on legible but unusual clause layouts |
| Reviewer agreement | — | — | Fields where two reviewers disagree with each other first |
| Evidence completeness | — | — | Records that crossed a schema version boundary mid-run |
The figures above are withheld until the first run is complete and independently witnessed — putting numbers here that nobody outside gothink has seen would repeat the practice this page exists to replace. The measures, the weaknesses and the method are published now because those are the parts a buyer can hold us to. The scores land on the changelog the day the witnessed run finishes, with the raw outputs attached.
Ask every shortlisted vendor for their five numbers against v0.1 and their raw outputs. It is a fair test, none of them can tune to a corpus they did not choose, and the answers will separate the field faster than a feature matrix. Take the wording straight from this page — you do not need our permission and we would rather you did not ask for it.
Request the corpus and harnessThe corpus, labels, adjudication guide and scoring harness are yours under a permissive licence, including for competitive marketing against us. The only condition is that a published score carries the version and the raw outputs. Disagreements about the method are welcome and belong in the changelog.
Tell us you are running itMost teams that built extraction in-house have never measured gate precision or evidence completeness, because nothing prompted them to. Run the harness against your own pipeline before you decide whether to buy anything. If you score well, you have saved a procurement cycle; if you do not, you now know which two layers are missing.
The five layers, and what each one addsA public corpus is the right way to shortlist and the wrong way to choose. Once the shortlist is down to two, run twenty-five of your own documents through each — labelled by your people, scored on the measures above, on the estate you actually have to defend.