We turn expert work into training signal.

The next gains won’t come from more data. They’ll come from harder, verifiable work.

120+Expert domains
99.2%Acceptance target
72hrsPilot turnaround
1:1Research partnership

Problem

Models plateau on outputs. They improve on reasoning.

Today’s models answer well and work poorly. Real work is multi-step. It has state, tools, constraints and consequences. Most training data captures none of that.

The signal that matters is how experts decide, what they check, and where they stop. It isn’t on the internet. It has to be built.

What we build

One partner for the whole learning loop.

From the first expert demonstration to the millionth agent rollout.

Expert training data

Demonstrations, reasoning traces and preference data from people who do the work. Built to a specification and reviewed before delivery.

SFT & preference dataExpert trajectories

RL environments

Stateful worlds with real tools, hard tasks and verifiable rewards. Resettable, versioned and ready for agents.

Multi-step workflowsExecutable verifiers

Evaluation

Held-out tasks, calibrated rubrics and failure analysis. Know what improved, what didn’t, and why.

Capability evaluationsReward integrity
Quality is a process, not a checkpoint.See how we work

Environment library

Real work. Repeatable worlds.

Six environments built around the complexity of professional work. Each one has state, tools, a task and a verifier.

Have a workflow in mind?

SWE / 001 · Sandbox

From a failing test to a verified fix.

Verifiable

Fix a duplicate payment caused by retrying a timed-out request. Keep existing behaviour across the repository intact.

TerminalCode editorTest runner
payments / retry.tsread-only
01async function processPayment(request) {
02 const key = request.idempotencyKey;
03 const existing = await ledger.find(key);
04
05 if (existing) return existing.result;
06
07 return ledger.transaction(async (tx) => {
08 const result = await charge(request);
09 await tx.record(key, result);
10 return result;
11 });
12}
RolloutReward —
Reproduced the duplicate-charge failurePending
Applied idempotency guardPending
24 / 24 regression tests passedPending
Deterministic verification

Custom environments are scoped to your model, tools and evaluation needs.

How we work

More signal. Less guesswork.

Every delivery starts with a definition of success and ends with a measurement against it.

01

Specify

Start with the gap.

Agree on the capability, the task distribution and the acceptance criteria before anything is produced.

02

Construct

Make the work real.

Build tasks, environments and reference solutions with the context an agent actually needs.

03

Challenge

Test the signal.

Review expert work, audit verifiers and hunt for shortcuts that reward the wrong behaviour.

04

Refine

Close the loop.

Inspect model attempts, evaluate on held-out tasks and target the failures that persist.

Expert-reviewed Verifier-tested Version-controlled Evaluation-led

Questions

Built for hard problems.

Different models. Different frontiers. This is how we usually start.

A focused expert dataset, a custom reinforcement learning environment, or a program that spans demonstrations, rewards and evaluation. Start with the capability you want to improve. We shape the task distribution and the delivery around it.

What should your model learn next?

Bring the capability gap. We’ll build the work that closes it.

Pilot brief

Start with a focused pilot.

Tell us what your model needs to learn. We’ll turn it into a brief you can build on.

We reply to the email you give us. No newsletter, no sharing.

Environment specification

Software engineering

The structure behind every verifiable task.

ID
SWE / 001
Objective
Fix a duplicate payment caused by retrying a timed-out request. Keep existing behaviour across the repository intact.
Tools
Terminal · Code editor · Test runner
Verification
Deterministic verification
Acceptance criteria
  • Reproduced the duplicate-charge failure
  • Applied idempotency guard
  • 24 / 24 regression tests passed
Episode library
2,400 episodes

No proprietary or customer data.