The checked-in 2026-04-26 run scored 16/25. The 72% figure is a stable configuration result observed during iteration, not a minimum expected score. Questions whose intended filing is no longer the latest in production can become harder to reproduce without an explicit period or filing pin.
What the evaluation tests
The canary includes numeric extraction, section evidence, and multi-step financial reasoning across real issuer filings. The harness uses structured financial, segment, calculation, filing-section, and earnings-material capabilities. It evaluates whether the agent produces the benchmark answer under that configuration; it does not establish correctness for an arbitrary question, issuer, model, or integration. Answers are assessed first by numeric comparison where a numeric reference exists, then by semantic judgment for free text. A pass requires the configured judge to accept the answer above its threshold. Judge behavior, source revisions, and model behavior can affect the result.How to use the result responsibly
Use it as a regression signal for a bounded workflow. Pin issuer, filing period, and source identifiers in your own evaluation. Retain the API result’s provenance and request metadata. Independently validate calculations, non-GAAP measures, and legal or investment conclusions against the filing. One published gold-answer adjustment forfb_135 accepts both near-equal XBRL-supported geographic-revenue answers after review. The dataset and published ground truth remain the external source: PatronusAI/financebench.
