What we evaluate
We evaluate AI reasoning in maintenance and automation tasks, including constraint adherence, failure recognition and decisions that require human input.
Task scope
Text-based maintenance and automation scenarios with documented operating constraints, task requirements and escalation criteria.
Evaluation methods
What we measure
- Constraint adherence
- Task completeness
- Failure recognition
- Appropriate escalation
Expert review
Engineers, automation specialists and experienced technical reviewers. Reviewers design tasks, refine scoring criteria, assess responses and resolve disagreements.
Evaluation design
We document task selection, data permissions, scoring rules and model settings before evaluation. Pilot annotation checks whether reviewers apply the criteria consistently. Material disagreements receive additional expert review.
Evaluation reports
Reports document model versions, evaluation dates, sample sizes, performance and limitations. We provide scoring criteria and supporting material within the project’s data permissions.
Benchmark scope
This benchmark evaluates reasoning in documented scenarios. Physical system validation and domain-specific safety assurance require separate testing.
Custom benchmarks for your AI systems
Test model performance against your industry’s tasks and standards.
Discuss benchmark developmentView our benchmarks