human data.Get in touch

Financial services

Financial services AI, tested by Canadian finance professionals.

Open-model benchmarking, acquisition due diligence, financial modelling and fraud investigation, evaluated in English and Canadian French by CPAs, CFAs and bank practitioners.

Open-weight modelsPublished benchmarksExpert adjudication

What we evaluate

Five lines of work, each with a public benchmark to start from and a Canadian layer our experts add.

Open-model benchmarking

Open-weight models run on Canadian analyst tasks: filings, disclosures, comparables and modelling, scored against professional answers in both official languages.

  • Open weights only
  • Published baselines
  • EN · FR

Acquisition and due diligence

Quality of earnings, underwriting, reconciliation and closing tasks built from anonymized Canadian transactions. Experts score method, tie-out and presentation.

  • Quality of earnings
  • Reconciliation
  • Deal documents

Financial modelling

Multi-sheet models, templates and debugging tasks from real financial reports. Rewards check formulas and target cells, not the look of the answer.

  • Multi-sheet workbooks
  • Formula checks
  • Audit trail

Fraud and investigations

Transaction labels with investigator rationale, alert triage and evaluation of how a model explains a suspicious pattern, including false positives and missed patterns.

  • Alert triage
  • Investigator rationale
  • Evidence quality

Bilingual client service

Agents and assistants tested on Canadian products with Quebec clients: the same outcome, in French, with the right terminology and disclosures.

  • Quebec French
  • Disclosures
  • Hand-off

Public benchmarks we build on

We start from benchmarks the field already trusts, restricted to open-weight models you can run yourself, then add the Canadian layer: our products, our rules, our French. Results are shown as published; they are not Human Data runs.

Public benchmark

Finance Agent v2

Entry-level analyst questions on public-company filings and transcripts: earnings analysis, disclosures, adjustments, comparables, precedents and financial modelling, with retrieval and calculation tools.

Published by
Vals AI
Scale
927 expert-reviewed questions · 9 task categories · 75 models
Metric
Accuracy (%) · cost per test (USD)
ModelAccuracyAccuracyCost / test
GLM 5.3 FlashOpen weightsZ.ai · MIT licence57.85%$0.05
MiMo V2.6 ProOpen weightsXiaomi · MIT licence57.34%$0.20
MiMo V2.6 FlashOpen weightsXiaomi · MIT licence56.28%$0.07
GLM 5.3Open weightsZ.ai · open weights55.84%$1.07
DeepSeek V4.1 FlashOpen weightsDeepSeek · MIT licence53.48%$0.21
Kimi K3Open weightsMoonshot AI · modified MIT53.11%$1.91
Qwen 3.8 27BOpen weightsAlibaba · Apache 2.048.55%$0.75
MiniMax-M3Open weightsMiniMax · open weights48.27%$0.32

Best overall on the same leaderboard: Gemini 4 Argon, 65.40% at $4.38 per test. The open-weight leaders sit within eight points at a fraction of the cost.

Source: Vals AI, leaderboard retrieved 6 October 2026. Published results as reported by the source, not Human Data runs. Open-weight status checked against the model publishers in October 2026.

Public benchmark

Westworld Finance Diligence Bench

A complete company-acquisition due-diligence process on anonymized private transactions, run inside desktop environments with trajectories reaching hundreds of steps. Every task was written and reviewed by practising deal professionals.

Published by
Halluminate
Scale
88 tasks · 7 categories (diligence review 22, pitch materials 15, modelling 14, underwriting 13, data preparation 11, quality of earnings 9, closing 4) · 8 models × 3 attempts
Metric
Mean score 0–1 in a native CLI harness / in a computer-use harness
ModelNative CLINative CLIComputer use
Claude Opus 5.5Anthropic0.590.59
Claude Fable 5.1Anthropic0.560.58
GPT-6.1 SolOpenAI0.540.55
GPT-6 AstraOpenAI0.530.54
Gemini 3.8 FlashGoogle0.500.45
DeepSeek V4.1 FlashOpen weightsDeepSeek · MIT licence0.480.40
Grok 4.7xAI0.470.30
Muse Spark 1.3Meta0.450.26

The best configuration clears 0.59 of the available credit. Modelling and closing are the weakest categories for every model.

  • Finance-reasoning failures: wrong analytical method 22.1%, correct work presented wrongly 19.6%, reconciliation and tie-out 13.7%, copied instead of derived 11.8%.
  • Long-horizon failures: premature discovery closure 22.1%, delivery mechanics 14.9%, false verification 14.7%, instruction loss over the horizon 14.5%.

Source: Halluminate, published 2026. Published results as reported by the source, not Human Data runs. Open-weight status checked against the model publishers in October 2026.

Public benchmark

SpreadsheetBench 2

End-to-end business spreadsheet workflows built from real financial reports and corporate filings: financial modelling, templates, debugging and visualization, each averaging 11.8 worksheets.

Published by
Zhu et al., arXiv 2606.29955
Scale
321 tasks (100 financial modelling, 100 debugging, 97 template, 24 visualization) · expert-annotated
Metric
Overall task accuracy (%) · financial-modelling task accuracy (%)
ModelOverallOverallFinancial modelling
GLM-5Open weightsZ.ai · MIT licence17.14%22%
DeepSeek-V3.2Open weightsDeepSeek · MIT licence15.58%7%
Kimi K2.5Open weightsMoonshot AI · modified MIT14.64%15%
Qwen3.5-397B-A17BOpen weightsAlibaba · Apache 2.011.22%10%
MiniMax M2.5Open weightsMiniMax · open weights7.17%8%

Best closed model: Claude Opus 4.6 at 34.89% overall and 34.00% on financial modelling. Debugging stays at 12.00% even for the best model.

Source: Zhu et al., arXiv 2606.29955, June 2026. Published results as reported by the source, not Human Data runs. Open-weight status checked against the model publishers in October 2026.

Public benchmark

τ³-Banking

Text agents resolving banking customer-service tasks over a knowledge base of about 700 documents, with a simulated customer and a policy the agent must follow.

Published by
Sierra
Scale
27 models · banking knowledge domain of τ-bench
Metric
pass@1 (%)
Modelpass@1pass@1
Kimi K3Open weightsMoonshot AI · modified MIT · all tools, extended reasoning37.10%
GLM-5.2Open weightsZ.ai · MIT licence · all tools, extended reasoning37.10%
InklingOpen weightsThinking Machines · Apache 2.0 · all tools, extended reasoning25%
GLM-5Open weightsZ.ai · MIT licence · embedding retrieval, no extended reasoning9.80%
Qwen3.5-397B-A17BOpen weightsAlibaba · Apache 2.0 · embedding retrieval, no extended reasoning9.80%

Best overall: Qwen 3.8 Max, 55.2%. On the original τ²-bench domains (retail, airline, telecom) the open-weight Qwen3.5-397B-A17B leads at 87.9% pass^1.

  • Retrieval set-up changes the result more than model size: the same family scores 37.1% with full tools and 9.8% with embedding-only retrieval.

Source: Sierra, leaderboard retrieved 6 October 2026. Published results as reported by the source, not Human Data runs. Open-weight status checked against the model publishers in October 2026.

Fraud: labelled datasets, then investigator judgment

Public fraud datasets give a label per transaction. Investigations need more: the pattern, the evidence and the reasons a human would escalate or clear. Our experts add that layer and evaluate how a model explains its decision.

590,540card-not-present transactions

IEEE-CIS Fraud Detection

3.5% fraud (20,663 cases), 431 features, six months of activity.

Vesta · Kaggle, 2019
6.3Msimulated mobile-money transactions

PaySim

0.13% fraud (8,213 cases) over a 30-day simulation.

Lopez-Rojas et al., 2016

Where models fail on finance work

Failure analysis from the Westworld Finance Diligence Bench, 88 due-diligence tasks across eight models. These are the categories our rubrics are written to catch.

Finance reasoning failures

  1. Wrong analytical method
    22.1%
  2. Correct work, wrong presentation
    19.6%
  3. Reconciliation and tie-out failure
    13.7%
  4. Copied instead of derived
    11.8%

Long-horizon failures

  1. Premature discovery closure
    22.1%
  2. Delivery mechanics failure
    14.9%
  3. False verification
    14.7%
  4. Instruction loss over the horizon
    14.5%

Source: Halluminate, Westworld Finance Diligence Bench. Share of failures by category across all runs.

RL environments in Canada

RL environments for financial services

Simulated Canadian banking tasks where a model acts and an expert-written reward scores what it did. The same environments train a model and evaluate the one you already run.

Tasks

  • Registered-account transfers: TFSA, RRSP, FHSA, RESP
  • Mortgage renewal and rate-hold decisions
  • Card disputes and chargebacks
  • KYC onboarding and beneficial-ownership checks
  • Bill payments and payee matching for credit-union members
  • Fraud-alert triage with investigator escalation

Rewards

  • Exact checks: account type, amount, payee, date
  • Consent captured before any money moves
  • Disclosures delivered in the client’s language
  • Policy limits respected; out-of-scope requests handed to a person
  • Expert rubric for explanation quality, calibrated before use
Testing agents that act for clientsBanking

Move $2,000 into my TFSA before Friday.

Done. $2,000 moved.

MeantTFSA
DidRRSP
Flagged by Human Data evaluators: wrong registered account.Contribution room and tax treatment differ. The move cannot be undone without a withdrawal.
1 / 3

Credit unions and fintechs get the same environments with their own products, limits and vendor interfaces. Environments ship as Python packages for OpenEnv, Verifiers, SkyRL or NeMo Gym, or run in our evaluation harness.

How we build environments

Our benchmark

Financial evidence & reasoning

We evaluate whether AI supports financial conclusions with source evidence, calculates accurately and identifies missing information.

  • Source-grounded accuracy
  • Calculation correctness
  • Assumption disclosure
  • Handling missing information

Canadian finance expertise.
Measured, not assumed.

Tell us what your model needs to do and which open models you are weighing.