Task Catalog
Task & Benchmark Catalog
We build and staff pipelines across the task formats most relevant to modern model training and evaluation. Don't see your exact format below — we design custom pipelines too.
Supported formats
Task types we staff and scale.
HLE-style expert reasoning
Graduate/expert-level questions requiring deep domain reasoning, not surface pattern matching.
BrowseComp-style web research
Multi-step web research tasks requiring information synthesis and verification.
DeepSWE
Extended software engineering tasks requiring multi-file reasoning and iterative debugging.
GDPVal
Real-world, professional-grade task evaluation across knowledge-work domains.
SciCode
Scientific coding problems requiring both domain knowledge and implementation skill.
SWE-bench
Real GitHub issue resolution tasks.
Terminal-Bench
Command-line and terminal-based task execution.
VQA-style tasks
Visual question answering requiring multimodal reasoning.
Our process
How we build task pipelines
Scope
Define task format, difficulty distribution, and domain coverage with your team.
Source
Recruit and vet experts matched to the domain.
Produce
Generate, collect, or review tasks against your specification.
Grade & Verify
Apply rubric-based scoring, gold-answer checks, and inter-rater agreement analysis.
Deliver
Structured datasets in your preferred schema, with full documentation and QA reports.
Need a custom task format?
We design bespoke pipelines for specialized domains, custom schemas, and internal evaluation needs that off-the-shelf benchmarks can't cover.