human data.Get in touch

Our benchmarks

Financial evidence & reasoning

Assess financial calculations, source evidence and the assumptions behind model answers.

Service
Benchmark development
Focus
Financial services
Review
Domain experts
Request benchmark

What we evaluate

We evaluate whether AI supports financial conclusions with source evidence, calculates accurately and identifies missing information.

Task scope

Document-based tasks built around financial statements, notes and professional scenarios. Evaluations use public, licensed or purpose-created documents and client-authorised material.

Evaluation methods

What we measure

  • Source-grounded accuracy
  • Calculation correctness
  • Assumption disclosure
  • Handling missing information

Expert review

Finance and accounting professionals with experience interpreting the relevant documents. Reviewers design tasks, refine scoring criteria, assess responses and resolve disagreements.

Evaluation design

We document task selection, data permissions, scoring rules and model settings before evaluation. Pilot annotation checks whether reviewers apply the criteria consistently. Material disagreements receive additional expert review.

Evaluation reports

Reports document model versions, evaluation dates, sample sizes, performance and limitations. We provide scoring criteria and supporting material within the project’s data permissions.

Benchmark scope

Findings apply to the documents, tasks, model versions and conditions tested. Benchmark evaluation supports model development; it does not constitute investment advice or certification for financial decision-making.

Custom benchmarks for your AI systems

Test model performance against your industry’s tasks and standards.

Discuss benchmark developmentView our benchmarks