human data.Get in touch

Our benchmarks

Bilingual professional reasoning

Evaluate meaning, terminology and evidence across Canadian English and French.

Service
Benchmark development
Focus
Language & context
Review
Domain experts
Request benchmark

What we evaluate

We evaluate whether AI preserves professional meaning across Canadian English and French, including the evidence, qualifications and uncertainty in the original source.

Task scope

Paired English and Canadian French tasks in professional communication, with explicit context and source material. Tasks use public, licensed or purpose-created content.

Evaluation methods

What we measure

  • Meaning preservation
  • Terminology and register
  • Evidence consistency
  • Appropriate uncertainty

Expert review

Bilingual editors, translators and professionals with relevant subject expertise. Reviewers design tasks, refine scoring criteria, assess responses and resolve disagreements.

Evaluation design

We document task selection, data permissions, scoring rules and model settings before evaluation. Pilot annotation checks whether reviewers apply the criteria consistently. Material disagreements receive additional expert review.

Evaluation reports

Reports document model versions, evaluation dates, sample sizes, performance and limitations. We provide scoring criteria and supporting material within the project’s data permissions.

Benchmark scope

Findings apply to the professional tasks and language varieties included in the evaluation. Coverage is documented so teams can identify where additional language testing is needed.

Custom benchmarks for your AI systems

Test model performance against your industry’s tasks and standards.

Discuss benchmark developmentView our benchmarks