What we evaluate
We evaluate whether AI preserves professional meaning across Canadian English and French, including the evidence, qualifications and uncertainty in the original source.
Task scope
Paired English and Canadian French tasks in professional communication, with explicit context and source material. Tasks use public, licensed or purpose-created content.
Evaluation methods
What we measure
- Meaning preservation
- Terminology and register
- Evidence consistency
- Appropriate uncertainty
Expert review
Bilingual editors, translators and professionals with relevant subject expertise. Reviewers design tasks, refine scoring criteria, assess responses and resolve disagreements.
Evaluation design
We document task selection, data permissions, scoring rules and model settings before evaluation. Pilot annotation checks whether reviewers apply the criteria consistently. Material disagreements receive additional expert review.
Evaluation reports
Reports document model versions, evaluation dates, sample sizes, performance and limitations. We provide scoring criteria and supporting material within the project’s data permissions.
Benchmark scope
Findings apply to the professional tasks and language varieties included in the evaluation. Coverage is documented so teams can identify where additional language testing is needed.
Custom benchmarks for your AI systems
Test model performance against your industry’s tasks and standards.
Discuss benchmark developmentView our benchmarks