Task Catalog

Task & Benchmark Catalog

We build and staff pipelines across the task formats most relevant to modern model training and evaluation. Don't see your exact format below — we design custom pipelines too.

Supported formats

Task types we staff and scale.

HLE-style expert reasoning

Graduate/expert-level questions requiring deep domain reasoning, not surface pattern matching.

Frontier reasoning model evaluation

BrowseComp-style web research

Multi-step web research tasks requiring information synthesis and verification.

Agentic and retrieval-augmented model evaluation

DeepSWE

Extended software engineering tasks requiring multi-file reasoning and iterative debugging.

Coding agent training/evaluation

GDPVal

Real-world, professional-grade task evaluation across knowledge-work domains.

General capability benchmarking

SciCode

Scientific coding problems requiring both domain knowledge and implementation skill.

Scientific computing model evaluation

SWE-bench

Real GitHub issue resolution tasks.

Coding agent benchmarking

Terminal-Bench

Command-line and terminal-based task execution.

Agentic tool-use evaluation

VQA-style tasks

Visual question answering requiring multimodal reasoning.

Multimodal model training/evaluation

Our process

How we build task pipelines

01

Scope

Define task format, difficulty distribution, and domain coverage with your team.

02

Source

Recruit and vet experts matched to the domain.

03

Produce

Generate, collect, or review tasks against your specification.

04

Grade & Verify

Apply rubric-based scoring, gold-answer checks, and inter-rater agreement analysis.

05

Deliver

Structured datasets in your preferred schema, with full documentation and QA reports.

Need a custom task format?

We design bespoke pipelines for specialized domains, custom schemas, and internal evaluation needs that off-the-shelf benchmarks can't cover.

Discuss Your Requirements