Expert training data
Demonstrations, reasoning traces and preference data from people who do the work. Built to a specification and reviewed before delivery.
The next gains won’t come from more data. They’ll come from harder, verifiable work.
Problem
Today’s models answer well and work poorly. Real work is multi-step. It has state, tools, constraints and consequences. Most training data captures none of that.
The signal that matters is how experts decide, what they check, and where they stop. It isn’t on the internet. It has to be built.
What we build
From the first expert demonstration to the millionth agent rollout.
Demonstrations, reasoning traces and preference data from people who do the work. Built to a specification and reviewed before delivery.
Stateful worlds with real tools, hard tasks and verifiable rewards. Resettable, versioned and ready for agents.
Held-out tasks, calibrated rubrics and failure analysis. Know what improved, what didn’t, and why.
Environment library
Six environments built around the complexity of professional work. Each one has state, tools, a task and a verifier.
Have a workflow in mind?
SWE / 001 · Sandbox
Fix a duplicate payment caused by retrying a timed-out request. Keep existing behaviour across the repository intact.
async function processPayment(request) { const key = request.idempotencyKey; const existing = await ledger.find(key); if (existing) return existing.result; return ledger.transaction(async (tx) => { const result = await charge(request); await tx.record(key, result); return result; });}Custom environments are scoped to your model, tools and evaluation needs.
How we work
Every delivery starts with a definition of success and ends with a measurement against it.
Agree on the capability, the task distribution and the acceptance criteria before anything is produced.
Build tasks, environments and reference solutions with the context an agent actually needs.
Review expert work, audit verifiers and hunt for shortcuts that reward the wrong behaviour.
Inspect model attempts, evaluate on held-out tasks and target the failures that persist.
Questions
Different models. Different frontiers. This is how we usually start.
A focused expert dataset, a custom reinforcement learning environment, or a program that spans demonstrations, rewards and evaluation. Start with the capability you want to improve. We shape the task distribution and the delivery around it.
We set acceptance criteria before production, review expert work, test verifiers against shortcuts and keep evaluation separate from training. The goal is measurable improvement on held-out tasks, not a larger pile of labels.
Yes. A custom program starts with permitted uses, access boundaries and delivery requirements. Environments can mirror your tool interfaces and workflows, with resettable state and versioned tasks.
A capability brief, a small reviewed task set, working verification and an evaluation plan. Use the brief builder on this page to send us a short brief. You keep a copy, and we reply by email.
Bring the capability gap. We’ll build the work that closes it.