human data.Get in touch

RL environments · Canada

RL environments for Canadian domains

Simulated banking, insurance, public-service and bilingual tasks, with rewards written and verified by Canadian professionals. Built to run in your trainer, or in ours for evaluation.

0/ 312 concurrent environments
Each cell is one sandboxed environment. Dots are sessions in progress. Every scored episode feeds training, or, for evaluation, a report on the model you have.
Expert-written tasksVerifiable rewardsFramework-agnostic delivery

What an RL environment is

A task, the tools a model may use, and a verifier that scores what the model did. The model attempts the task many times. Each attempt is scored. The scores train the next version of the model, or, for evaluation, describe the version you have.

Human Data writes and verifies the nine components on the left: realistic Canadian tasks, the tools and state that make them real, and the rewards that decide what counts as correct. The loop on the right runs in your training stack. For evaluation, it runs in ours.

Where we sit in the pipeline

Every RL training system has the same five stages. Frameworks differ in how many they cover. Human Data supplies the three stages that need domain experts and hands off to whichever rollout and trainer you run. For evaluation engagements we also run the rollouts.

Human Data · environment delivery
Human Data · evaluation runs
OpenEnvMeta · PyTorch
ORSOpen Reward Standard
NeMo GymNVIDIA
Verifiers + Prime RLPrime Intellect
SkyRLNovaSky
GEMGeneral Experience Maker
  1. TasksWhat to solve
  2. HarnessHow the model interacts
  3. RewardHow good was it
  4. RolloutCollect episodes
  5. TrainerUpdate weights

Framework coverage as mapped in the Hugging Face guide to RL environments (May 2026). Select a chip to isolate one row. Environments are delivered as standard Python packages and can target the OpenEnv, Verifiers, SkyRL or NeMo Gym interfaces.

Environments for Canadian domains

Each environment is built from real Canadian products, rules and documents, in English and Canadian French.

Financial services

Mortgage renewals, registered-account transfers (TFSA, RRSP, FHSA), disputes and KYC onboarding against Canadian product rules. Rewards check the account, the amount, the consent and the disclosure.

  • Banking products
  • Consent and disclosure
  • EN · FR
Explore financial services

Insurance

Claims intake, coverage questions, renewals and underwriting triage built on Quebec and common-law policy wording. Rewards check coverage logic, exclusions and escalation.

  • Claims
  • Underwriting
  • Quebec civil law
Explore insurance

Public sector

Benefit applications, permit workflows and correspondence over public-service documents in both official languages. Rewards check eligibility rules, language requirements and hand-off to a person.

  • Benefits and permits
  • Official languages
  • Human hand-off
Explore public sector

Fintechs and credit unions

Payments, bill pay, card disputes and member service for smaller institutions, with the vendor interfaces they actually use. Rewards check payee, amount, timing and member consent.

  • Payments
  • Member service
  • Vendor APIs
Explore fintechs and credit unions

Bilingual service

Every environment runs in English and Canadian French with Quebec terminology and register. Rewards check meaning, terminology and that French clients are served in French.

  • Canadian French
  • Terminology
  • Language rights
Explore bilingual service

Testing agents that act for people

An agent that says “done” is not evidence that the right thing happened. Our environments record every action and compare it with the instruction, the policy and the client’s situation. Canadian evaluators adjudicate the gap, in English and French.

Example episodeBanking

Move $2,000 into my TFSA before Friday.

Done. $2,000 moved.

MeantTFSA
DidRRSP
Flagged by Human Data evaluators: wrong registered account.Contribution room and tax treatment differ. The move cannot be undone without a withdrawal.
1 / 5
  • Reliability across repeated attempts (pass^k), not a single lucky run
  • Policy compliance and consent before money moves
  • Correct hand-off to a person when the task leaves scope
  • Resistance to instructions hidden in documents and tool results
  • French parity: the same outcome for a French-speaking client

Public agent benchmarks we build on

Published results for open-weight models you can run yourself. Our Canadian environments add the products, rules and French that these benchmarks do not cover.

Public benchmark

τ³-Banking

Text agents resolving banking customer-service tasks over a knowledge base of about 700 documents, with a simulated customer and a policy the agent must follow.

Published by
Sierra
Scale
27 models · banking knowledge domain of τ-bench
Metric
pass@1 (%)
Modelpass@1pass@1
Kimi K3Open weightsMoonshot AI · modified MIT · all tools, extended reasoning37.10%
GLM-5.2Open weightsZ.ai · MIT licence · all tools, extended reasoning37.10%
InklingOpen weightsThinking Machines · Apache 2.0 · all tools, extended reasoning25%
GLM-5Open weightsZ.ai · MIT licence · embedding retrieval, no extended reasoning9.80%
Qwen3.5-397B-A17BOpen weightsAlibaba · Apache 2.0 · embedding retrieval, no extended reasoning9.80%

Best overall: Qwen 3.8 Max, 55.2%. On the original τ²-bench domains (retail, airline, telecom) the open-weight Qwen3.5-397B-A17B leads at 87.9% pass^1.

  • Retrieval set-up changes the result more than model size: the same family scores 37.1% with full tools and 9.8% with embedding-only retrieval.

Source: Sierra, leaderboard retrieved 6 October 2026. Published results as reported by the source, not Human Data runs. Open-weight status checked against the model publishers in October 2026.

629prompt-injection test cases over 97 tasks

AgentDojo

Four suites including banking. Published open-weight result: Llama 3 70B completes 34.02% of tasks, 18.28% under attack; 25.60% of injections succeed.

ETH Zurich · NeurIPS 2024
2tracks: underwriting, claims and coverage

InsureBench

Document-grounded cases written with practising underwriters and claims handlers, scored pass@1. Leaderboard opening 2026.

Huzzle Labs

How we build one

From a professional’s description of the work to a package your trainer can run.

01

Source the task

Practising Canadian professionals write scenarios from real work: the products, the documents, the edge cases and the French version. Each task has a gold outcome.

02

Build the harness

Tools, state and documents become a sandbox: a mock core-banking interface, a policy database, a case file. Nothing touches a production system.

03

Write the reward

Deterministic checks where the answer is exact (account, amount, date). Expert rubrics where judgment is needed. Every rubric is calibrated on a shared sample before use.

04

Verify and deliver

Two experts review each task; disagreements go to a senior adjudicator. Environments ship as Python packages in the interface your trainer expects.

Environment package
Tasks, tools, state and reward functions as a versioned Python package with a container image.
Task set
Scenarios with gold outcomes, in English and Canadian French, with the professional’s rationale.
Rubrics and calibration
Scoring rubrics, calibration results and reviewer agreement for every judgment-based reward.
Evaluation report
For evaluation engagements: pass^k by task family, failure taxonomy and corrected trajectories.

Scale is a property of the infrastructure

Environments run as containers. The task and the reward are what need experts; running thousands of sessions is a scheduling problem the ecosystem has already solved.

16,384concurrent environments on a 96-core multi-node cluster
2,048on an 8-core workstation running Docker
128on a 2-core hosted Space

All at 95% or higher session success. Source: OpenEnv scaling benchmark, cited in the Hugging Face guide to RL environments (May 2026).

Canadian tasks.
Verifiable rewards.

Tell us the domain, the trainer you run and the behaviour you need to measure.