We ran the audit, we publish the method, and the harness reports its own exceptions, which is the part that makes the rest believable. Here is exactly what we measure and how.
479 figures re-checked against the regulator's own structured record. Every one matched, at a 95% confidence interval of 99.2 to 100%.
The number is right, and the document we attribute it to is the one the regulator records it in.
Every fact's key recomputes from its own content, with 0 mismatches across 29.8M facts and no sampling.
Assets equal liabilities plus equity. The exceptions are surfaced and listed, not hidden.
Enforced in the product, not promised in a brochure. Each rule exists because the alternative quietly produces a plausible, wrong answer.
The same rules hold when the answer comes from an agent. The model retrieves and explains; it does not invent the numbers, and it cannot pass off a figure it did not fetch.
Valuations and calculations run in a deterministic engine that reproduces the same result byte for byte. The model narrates that result and checks it for sense; it never produces the number itself.
A guardrail checks that each number in an answer traces back to a retrieved cell in the corpus. If it does not resolve, the answer is refused rather than guessed.
Retrieval is scoped to the companies, periods and documents you asked for, and every claim carries the filing, accession and location it came from.
A fixed set of questions with known answers runs on every model and prompt change, so a regression is caught by us, not by you.
Every figure on every surface names the filing behind it. Request access and click one.