What we evaluate
We evaluate whether AI supports financial conclusions with source evidence, calculates accurately and identifies missing information.
Task scope
Document-based tasks built around financial statements, notes and professional scenarios. Evaluations use public, licensed or purpose-created documents and client-authorised material.
Evaluation methods
What we measure
- Source-grounded accuracy
- Calculation correctness
- Assumption disclosure
- Handling missing information
Expert review
Finance and accounting professionals with experience interpreting the relevant documents. Reviewers design tasks, refine scoring criteria, assess responses and resolve disagreements.
Evaluation design
We document task selection, data permissions, scoring rules and model settings before evaluation. Pilot annotation checks whether reviewers apply the criteria consistently. Material disagreements receive additional expert review.
Evaluation reports
Reports document model versions, evaluation dates, sample sizes, performance and limitations. We provide scoring criteria and supporting material within the project’s data permissions.
Benchmark scope
Findings apply to the documents, tasks, model versions and conditions tested. Benchmark evaluation supports model development; it does not constitute investment advice or certification for financial decision-making.
Custom benchmarks for your AI systems
Test model performance against your industry’s tasks and standards.
Discuss benchmark developmentView our benchmarks