Open-model benchmarking
Open-weight models run on Canadian analyst tasks: filings, disclosures, comparables and modelling, scored against professional answers in both official languages.
- Open weights only
- Published baselines
- EN · FR
Financial services
Open-model benchmarking, acquisition due diligence, financial modelling and fraud investigation, evaluated in English and Canadian French by CPAs, CFAs and bank practitioners.
Five lines of work, each with a public benchmark to start from and a Canadian layer our experts add.
Open-weight models run on Canadian analyst tasks: filings, disclosures, comparables and modelling, scored against professional answers in both official languages.
Quality of earnings, underwriting, reconciliation and closing tasks built from anonymized Canadian transactions. Experts score method, tie-out and presentation.
Multi-sheet models, templates and debugging tasks from real financial reports. Rewards check formulas and target cells, not the look of the answer.
Transaction labels with investigator rationale, alert triage and evaluation of how a model explains a suspicious pattern, including false positives and missed patterns.
Agents and assistants tested on Canadian products with Quebec clients: the same outcome, in French, with the right terminology and disclosures.
We start from benchmarks the field already trusts, restricted to open-weight models you can run yourself, then add the Canadian layer: our products, our rules, our French. Results are shown as published; they are not Human Data runs.
Entry-level analyst questions on public-company filings and transcripts: earnings analysis, disclosures, adjustments, comparables, precedents and financial modelling, with retrieval and calculation tools.
| Model | Accuracy | Accuracy | Cost / test |
|---|---|---|---|
| GLM 5.3 FlashOpen weightsZ.ai · MIT licence | 57.85% | $0.05 | |
| MiMo V2.6 ProOpen weightsXiaomi · MIT licence | 57.34% | $0.20 | |
| MiMo V2.6 FlashOpen weightsXiaomi · MIT licence | 56.28% | $0.07 | |
| GLM 5.3Open weightsZ.ai · open weights | 55.84% | $1.07 | |
| DeepSeek V4.1 FlashOpen weightsDeepSeek · MIT licence | 53.48% | $0.21 | |
| Kimi K3Open weightsMoonshot AI · modified MIT | 53.11% | $1.91 | |
| Qwen 3.8 27BOpen weightsAlibaba · Apache 2.0 | 48.55% | $0.75 | |
| MiniMax-M3Open weightsMiniMax · open weights | 48.27% | $0.32 |
A complete company-acquisition due-diligence process on anonymized private transactions, run inside desktop environments with trajectories reaching hundreds of steps. Every task was written and reviewed by practising deal professionals.
| Model | Native CLI | Native CLI | Computer use |
|---|---|---|---|
| Claude Opus 5.5Anthropic | 0.59 | 0.59 | |
| Claude Fable 5.1Anthropic | 0.56 | 0.58 | |
| GPT-6.1 SolOpenAI | 0.54 | 0.55 | |
| GPT-6 AstraOpenAI | 0.53 | 0.54 | |
| Gemini 3.8 FlashGoogle | 0.50 | 0.45 | |
| DeepSeek V4.1 FlashOpen weightsDeepSeek · MIT licence | 0.48 | 0.40 | |
| Grok 4.7xAI | 0.47 | 0.30 | |
| Muse Spark 1.3Meta | 0.45 | 0.26 |
End-to-end business spreadsheet workflows built from real financial reports and corporate filings: financial modelling, templates, debugging and visualization, each averaging 11.8 worksheets.
| Model | Overall | Overall | Financial modelling |
|---|---|---|---|
| GLM-5Open weightsZ.ai · MIT licence | 17.14% | 22% | |
| DeepSeek-V3.2Open weightsDeepSeek · MIT licence | 15.58% | 7% | |
| Kimi K2.5Open weightsMoonshot AI · modified MIT | 14.64% | 15% | |
| Qwen3.5-397B-A17BOpen weightsAlibaba · Apache 2.0 | 11.22% | 10% | |
| MiniMax M2.5Open weightsMiniMax · open weights | 7.17% | 8% |
Text agents resolving banking customer-service tasks over a knowledge base of about 700 documents, with a simulated customer and a policy the agent must follow.
| Model | pass@1 | pass@1 |
|---|---|---|
| Kimi K3Open weightsMoonshot AI · modified MIT · all tools, extended reasoning | 37.10% | |
| GLM-5.2Open weightsZ.ai · MIT licence · all tools, extended reasoning | 37.10% | |
| InklingOpen weightsThinking Machines · Apache 2.0 · all tools, extended reasoning | 25% | |
| GLM-5Open weightsZ.ai · MIT licence · embedding retrieval, no extended reasoning | 9.80% | |
| Qwen3.5-397B-A17BOpen weightsAlibaba · Apache 2.0 · embedding retrieval, no extended reasoning | 9.80% |
Public fraud datasets give a label per transaction. Investigations need more: the pattern, the evidence and the reasons a human would escalate or clear. Our experts add that layer and evaluate how a model explains its decision.
3.5% fraud (20,663 cases), 431 features, six months of activity.
Vesta · Kaggle, 20190.13% fraud (8,213 cases) over a 30-day simulation.
Lopez-Rojas et al., 2016492 frauds (0.172%), 28 anonymized components plus time and amount.
Université Libre de Bruxelles · WorldlineFailure analysis from the Westworld Finance Diligence Bench, 88 due-diligence tasks across eight models. These are the categories our rubrics are written to catch.
Source: Halluminate, Westworld Finance Diligence Bench. Share of failures by category across all runs.
RL environments in Canada
Simulated Canadian banking tasks where a model acts and an expert-written reward scores what it did. The same environments train a model and evaluate the one you already run.
Move $2,000 into my TFSA before Friday.
Done. $2,000 moved.
Credit unions and fintechs get the same environments with their own products, limits and vendor interfaces. Environments ship as Python packages for OpenEnv, Verifiers, SkyRL or NeMo Gym, or run in our evaluation harness.
How we build environmentsOur benchmark
We evaluate whether AI supports financial conclusions with source evidence, calculates accurately and identifies missing information.
Tell us what your model needs to do and which open models you are weighing.