The provenance benchmark

AI is only as good as the facts it can retrieve.

The audit behind that claim, published in full. Two tiers: structural integrity across all 29.8 million facts, and source-verified accuracy on a stratified sample checked back to the filing it cites. Every number carries its method and its exceptions.

The method

Two tiers, because accurate means two things.

Conflating them is how a vendor claims 99% without proving it. We publish both, separately.

Tier A, structural integrity

Does the corpus contradict itself? Checkable from the fact rows alone, so it runs deterministically over every fact and proves the pipeline's invariants hold. A pass here is necessary, not sufficient.

Tier B, source-verified accuracy

Does each value match the filing it cites? That needs the primary source, so it runs on a stratified sample checked back to the SEC XBRL record. This is the number that answers whether the citations hold.

Tier A

Structural integrity, across every fact.

Deterministic and reproducible over the whole corpus of 29,818,448 facts. Same corpus in, same numbers out.

100%

Key integrity

Every fact's join key recomputes from its own content, with 0 mismatches across 29.8M facts.

100%

Citation well-formed

Every one of 27.2M financial facts carries a valid citation in its jurisdiction's scheme.

99.88%

Balance identity

Assets equal liabilities plus equity, bar 465 real exceptions surfaced and listed, not hidden.

99.68%

Overall integrity

Row-level integrity across 168.5M individual check-instances.

Tier B

Source-verified against the filing.

A stratified sample of 479 facts, each checked back to the SEC XBRL record, across US and foreign filers.

Source-verified sample479 facts, Wilson 95% CI
CheckResult95% CI
Value matches the SEC source100.00%99.2 to 100%
Value and the exact cited filing99.79%98.8 to 100%
Value, citation and exact concept99.16%97.9 to 99.7%
Wrong numbers, fabrications00 of 479
Not one figure was wrong. Every one either matched the SEC record exactly, or was run down to a display rule such as native currency versus USD translation. Reproducible with the published harness on a stated sample.

Want the harness and the full write-up?

Request access and we will share the benchmark method, the sample and the re-runnable harness.