human data.Get in touch

Insurance

Insurance AI, from claimsdocuments to driving footage

Canadian insurance professionals label documents, footage and decisions, and test the models that act on them, in English and French.

What we deliver

Four lines of insurance work, each with a public benchmark or dataset to start from and a Canadian layer our experts add.

Claims documents, EN / FR

Policy wording, claim forms, estimates and correspondence labelled by field and linked to source passages, in both languages. Evaluations check extraction, coverage interpretation and escalation when information is missing.

  • Field extraction
  • Coverage wording
  • Bilingual

Driving and telematics footage

Road footage and in-cab video labelled for objects, events and driver behaviour, from a safe stop to distraction. Used for rewards programmes, claims evidence and behaviour models.

  • Object tracking
  • Driver behaviour
  • Event timestamps

Underwriting and coverage decisions

Appetite, qualification, limits and product recommendation tasks over guidelines and application data. Experts score the decision and the reasoning.

  • Appetite
  • Classification
  • Guidelines

Quebec insurance knowledge

Questions and scenarios under Quebec civil law and the certification material for insurance representatives, in French. The gap between English and French performance is measured, not assumed.

  • Civil law
  • Certification material
  • French first

Claims documents

One record from two languages

English and French insurance document fields are located, extracted and aligned to one structured record. Experts link every extracted value to its source passage and flag what the document does not support.

Insurance document processing

Driving footage

Behaviour, not just objects

Rewards programmes and claims evidence depend on what a driver did, not only what the camera saw. Annotators label events, interactions and outcomes with timestamps, and evaluate model descriptions against the footage.

Video annotation

Public benchmarks we build on

Published results for open-weight models you can run yourself. Our Canadian work adds Quebec civil law, bilingual documents and the products Canadian insurers actually sell.

Public benchmark

AEPC-QA · Quebec insurance certification

807 multiple-choice questions in French, digitized from the manuals used to certify insurance representatives in Quebec, under Quebec civil law.

Published by
Beauchemin & Khoury, Université Laval, arXiv 2603.07825
Scale
807 questions · 51 models · closed-book and retrieval-augmented
Metric
Accuracy (%), closed-book / with retrieval (random: 25%)
ModelClosed-bookClosed-bookWith retrieval
DeepSeek-R1 (deepseek-reasoner)Open weightsDeepSeek · MIT licence36.30%71.77%
R1-1776Open weightsPerplexity · MIT licence49.38%69.09%
DeepSeek-V3 (deepseek-chat)Open weightsDeepSeek · MIT licence58.23%67.78%
Qwen3-30B-A3BOpen weightsAlibaba · Apache 2.054.77%37.78%
QwQ-32BOpen weightsAlibaba · Apache 2.051.36%27.49%
Llama-3.3-70B-InstructOpen weightsMeta · Llama licence61.48%0.99%
Llama-3.3-Nemotron-Super-49BOpen weightsNVIDIA · open weights47.20%3.09%
Granite-3.2-8B-InstructOpen weightsIBM · Apache 2.030%30.86%

Best overall: o3 at 76.13% closed-book and 78.68% with retrieval. Retrieval lifts DeepSeek-R1 by 35 points and collapses Llama-3.3-70B to 0.99%, below random: the model stops following the answer format.

Source: Beauchemin & Khoury, Université Laval, arXiv 2603.07825, March 2026. Published results as reported by the source, not Human Data runs. Open-weight status checked against the model publishers in October 2026.

Public benchmark

Underwrite

Commercial property and casualty underwriting: appetite, qualification, limits and deductibles, product recommendation, business classification and small-business eligibility, using SQL and guideline tools over a database and proprietary business rules.

Published by
Snorkel AI, arXiv 2602.00456
Scale
300 multi-turn tasks from 3,000 synthesized applicant profiles · 13 models
Metric
Accuracy (%), LLM judge with over 95% agreement with expert annotations
ModelAccuracyAccuracy
DeepSeek V3.1Open weightsDeepSeek · MIT licence73.70%
Qwen3 Coder 480BOpen weightsAlibaba · Apache 2.073.30%
Kimi K2 InstructOpen weightsMoonshot AI · modified MIT56.70%
Qwen3 235BOpen weightsAlibaba · Apache 2.030%
gpt-oss-120bOpen weightsOpenAI · Apache 2.030%

Best overall: Claude Sonnet 4.5, 90.3%. Smaller models hallucinated insurance products absent from the guidelines in 58–66% of failures; pass^k over four attempts drops about 20 points.

Source: Snorkel AI, arXiv 2602.00456, January 2026. Published results as reported by the source, not Human Data runs. Open-weight status checked against the model publishers in October 2026.

Footage and vision datasets

Public driving and insurance-vision datasets set the baseline for object and behaviour labels. Our footage work adds Canadian roads, seasons and signage, and the behaviour labels an adjuster needs.

22,424in-cab images

State Farm Distracted Driver Detection

10 classes, from safe driving to texting, phone use, reaching behind, drinking and talking to a passenger.

State Farm · Kaggle, 2016
100,000driving video clips of about 40 seconds

BDD100K

10 perception tasks: object boxes, drivable area, lane markings, segmentation and more; crowdsourced from over 50,000 rides.

UC Berkeley
4insurance lines: auto, property, health, agricultural

INS-MMBench

22 fundamental, 12 meta and 5 scenario tasks for vision-language models; 11 models evaluated.

ICCV 2025

Industry benchmarks in progress

2tracks: underwriting, claims and coverage

InsureBench

Document-grounded cases written with practising underwriters and claims handlers, scored pass@1. Leaderboard opening 2026.

Huzzle Labs
629prompt-injection test cases over 97 tasks

AgentDojo

Four suites including banking. Published open-weight result: Llama 3 70B completes 34.02% of tasks, 18.28% under attack; 25.60% of injections succeed.

ETH Zurich · NeurIPS 2024

RL environments in Canada

RL environments for insurance

Simulated claims, renewal and underwriting tasks where a model acts on a policyholder’s behalf and an expert-written reward scores what it did. Built on Quebec and common-law policy wording, in English and French.

Tasks

  • First notice of loss and claim intake
  • Coverage questions against policy wording
  • Renewals with unchanged coverage
  • Driver and vehicle changes
  • Underwriting triage: appetite, classification, limits

Rewards

  • Right policy, right line of business, right deductible
  • Coverage changes disclosed, never silent
  • Exclusions and conditions applied as written
  • Escalation to an adjuster or underwriter when required
  • French parity for Quebec policyholders
Testing agents that act for policyholdersProperty claims

File a claim for the hail damage from last Thursday.

Your claim is open.

MeantHail damage to the roof, property policy
DidOpened an auto claim
Flagged by Human Data evaluators: wrong policy.The client holds both policies. The claim now sits in the wrong queue with the wrong deductible.
1 / 3

Canadian insurance expertise.
Documents, footage, decisions.

Tell us which line of business and which models you are evaluating.