Task & Benchmark Catalog

Expert-built tasks. Evaluation-ready pipelines.

We build and staff task pipelines for modern model training and evaluation, including custom formats.

Illustrative workflow
Inputs
Expert questionprompt.md
Reference answergold.json
Grading rubricrubric.yaml
Expert
pipeline
Structured output
{
  "task_id": "t_0142",
  "domain": "physics",
  "score": 0.92,
  "verified": true
}
Supported formats

Task types we staff and scale

Eight established formats, each shaped around the reasoning and verification your evaluation requires.

resolved answer
Illustrative

HLE-style expert reasoning

Graduate and expert-level questions requiring deep domain reasoning, not surface pattern matching.

Frontier reasoning evaluation
synthesisverified ✓
Illustrative

BrowseComp-style web research

Multi-step web research tasks requiring information synthesis and verification.

Agentic & retrieval evaluation
solver.tsutils.ts- return cache[key]+ return cache.get(key)
Illustrative

DeepSWE

Extended software engineering tasks requiring multi-file reasoning and iterative debugging.

Coding agent training
EVALUATION RUBRIC
Illustrative

GDPVal

Real-world, professional-grade task evaluation across knowledge-work domains.

General capability benchmarking
∂u/∂t = α∇²usolve(grid, dt)return field
Illustrative

SciCode

Scientific coding problems requiring both domain knowledge and implementation skill.

Scientific computing evaluation
ISSUE #42PATCH+ 4 linesTESTSPASS ✓
Illustrative

SWE-bench

Real GitHub issue resolution tasks that connect problems, patches, and test outcomes.

Coding agent benchmarking
$ run task --schema eval.jsonvalidating records...✓ output/results.jsonl
Illustrative

Terminal-Bench

Command-line and terminal-based task execution in reproducible environments.

Agentic tool-use evaluation
What is shown?A: instrument
Illustrative

VQA-style tasks

Visual question answering requiring grounded multimodal reasoning.

Multimodal training & evaluation
Our process

One connected production system

Five coordinated stages turn a task specification into a validated, documented dataset.

Illustrative stage view
Task brief
Format
Difficulty
Domain
Specification aligned
Stage 01 / 05

Scope

Define format, difficulty distribution, and domain coverage with your team.

Need a custom task format?

We design bespoke pipelines for specialized domains, custom schemas, and internal evaluation needs that off-the-shelf benchmarks can't cover.