Simulated banking, insurance, public-service and bilingual tasks, with rewards written and verified by Canadian professionals. Built to run in your trainer, or in ours for evaluation.
Each cell is one sandboxed environment. Dots are sessions in progress. Every scored episode feeds training, or, for evaluation, a report on the model you have.
A task, the tools a model may use, and a verifier that scores what the model did. The model attempts the task many times. Each attempt is scored. The scores train the next version of the model, or, for evaluation, describe the version you have.
RL environmentdocker · in-processHuman Data builds and verifies
Tasksdataset
Prompt templatetask → prompt
Initial statereset()
Tools / harnesstool definitions + wiring
Stateworld state
Execution backendwhere tools run
Observationnext-turn input
Rewardrubric / verify
Donetermination
Episode controlactionobservation
Training loopagent · collector · trainerYour stack, or ours for evaluation
LLM agentpolicy · token generation
Rolloutcollect episode
Trainerloss + gradient
Weight updateapply update
policy update
Human Data writes and verifies the nine components on the left: realistic Canadian tasks, the tools and state that make them real, and the rewards that decide what counts as correct. The loop on the right runs in your training stack. For evaluation, it runs in ours.
Where we sit in the pipeline
Every RL training system has the same five stages. Frameworks differ in how many they cover. Human Data supplies the three stages that need domain experts and hands off to whichever rollout and trainer you run. For evaluation engagements we also run the rollouts.
Human Data · environment delivery
Human Data · evaluation runs
OpenEnvMeta · PyTorch
ORSOpen Reward Standard
NeMo GymNVIDIA
Verifiers + Prime RLPrime Intellect
SkyRLNovaSky
GEMGeneral Experience Maker
TasksWhat to solve
HarnessHow the model interacts
RewardHow good was it
RolloutCollect episodes
TrainerUpdate weights
RL environment
Framework coverage as mapped in the Hugging Face guide to RL environments (May 2026). Select a chip to isolate one row. Environments are delivered as standard Python packages and can target the OpenEnv, Verifiers, SkyRL or NeMo Gym interfaces.
Environments for Canadian domains
Each environment is built from real Canadian products, rules and documents, in English and Canadian French.
Financial services
Mortgage renewals, registered-account transfers (TFSA, RRSP, FHSA), disputes and KYC onboarding against Canadian product rules. Rewards check the account, the amount, the consent and the disclosure.
Claims intake, coverage questions, renewals and underwriting triage built on Quebec and common-law policy wording. Rewards check coverage logic, exclusions and escalation.
Benefit applications, permit workflows and correspondence over public-service documents in both official languages. Rewards check eligibility rules, language requirements and hand-off to a person.
Payments, bill pay, card disputes and member service for smaller institutions, with the vendor interfaces they actually use. Rewards check payee, amount, timing and member consent.
Every environment runs in English and Canadian French with Quebec terminology and register. Rewards check meaning, terminology and that French clients are served in French.
An agent that says “done” is not evidence that the right thing happened. Our environments record every action and compare it with the instruction, the policy and the client’s situation. Canadian evaluators adjudicate the gap, in English and French.
Example episodeBanking
Move $2,000 into my TFSA before Friday.
Done. $2,000 moved.
MeantTFSA
DidRRSP
Flagged by Human Data evaluators: wrong registered account.Contribution room and tax treatment differ. The move cannot be undone without a withdrawal.
1 / 5
Reliability across repeated attempts (pass^k), not a single lucky run
Policy compliance and consent before money moves
Correct hand-off to a person when the task leaves scope
Resistance to instructions hidden in documents and tool results
French parity: the same outcome for a French-speaking client
Public agent benchmarks we build on
Published results for open-weight models you can run yourself. Our Canadian environments add the products, rules and French that these benchmarks do not cover.
Public benchmark
τ³-Banking
Text agents resolving banking customer-service tasks over a knowledge base of about 700 documents, with a simulated customer and a policy the agent must follow.
Published by
Sierra
Scale
27 models · banking knowledge domain of τ-bench
Metric
pass@1 (%)
Model
pass@1
pass@1
Kimi K3Open weightsMoonshot AI · modified MIT · all tools, extended reasoning
37.10%
GLM-5.2Open weightsZ.ai · MIT licence · all tools, extended reasoning
From a professional’s description of the work to a package your trainer can run.
01
Source the task
Practising Canadian professionals write scenarios from real work: the products, the documents, the edge cases and the French version. Each task has a gold outcome.
02
Build the harness
Tools, state and documents become a sandbox: a mock core-banking interface, a policy database, a case file. Nothing touches a production system.
03
Write the reward
Deterministic checks where the answer is exact (account, amount, date). Expert rubrics where judgment is needed. Every rubric is calibrated on a shared sample before use.
04
Verify and deliver
Two experts review each task; disagreements go to a senior adjudicator. Environments ship as Python packages in the interface your trainer expects.
Environment package
Tasks, tools, state and reward functions as a versioned Python package with a container image.
Task set
Scenarios with gold outcomes, in English and Canadian French, with the professional’s rationale.
Rubrics and calibration
Scoring rubrics, calibration results and reviewer agreement for every judgment-based reward.
Evaluation report
For evaluation engagements: pass^k by task family, failure taxonomy and corrected trajectories.
Scale is a property of the infrastructure
Environments run as containers. The task and the reward are what need experts; running thousands of sessions is a scheduling problem the ecosystem has already solved.
16,384concurrent environments on a 96-core multi-node cluster