How can verifiers be bypassed? Explore our adversarial tests →

Products

How does someone learn to do a job? Reading the manual is only the beginning.

They make a mistake, someone points it out, and they try again.

And the person spotting the mistake knows things that never made it into the manual.

Our products share one purpose:

Turn that unwritten judgment into something a machine can be tested against.

01

Simulation environments

Practicing a tennis swing ten thousand times does not mean you can play a match. The real test happens on the court.

We build stateful work scenarios where software, data, tasks, and exceptions operate together. Every run can be reset, replayed, and inspected. Agents do the work, and every action changes what happens next.

02

Verifiers

A balanced ledger can still be wrong. Move a discrepancy into suspense, and the trial balance returns to zero.

Alongside the final state, we check invariants: constraints that must hold throughout the trajectory. Tasks in the same domain can share business rules, with evidence used to verify the quality of completion.

03

Adversarial testing

Everyone says their grader is rigorous. Few tell you how many times it has been bypassed.

We attack verifiers with known degenerate strategies and preserve the original gaps, the checks we add, and the results of running again. Executable cases and checking code are published together so every number can be questioned.

04

Expert trajectories

A senior accountant and a newcomer may reach the same reconciliation result through very different paths. The process carries the expertise.

Practitioners work through an environment while we record each action, the basis for key decisions, and the moments they stop to reconsider. Trajectories retain rework; professional review determines whether a record qualifies for training data.

05

Process supervision data

A teacher marking an essay does more than write a score at the end. They underline a sentence in the third paragraph.

We check individual steps: whether an action is valid, evidence is sufficient, and constraints remain intact. Process judgments and final scores are stored separately, providing traceable supervision signals for process reward models.

06

Preference data

Five experienced people may suggest five different revisions to the same proposal. Their disagreements can carry more information than their agreement.

Our workflows separate independent work from independent review and identify disagreements explicitly. We preserve the evidence behind each judgment before domain reviewers adjudicate, capturing what makes professional work difficult.

07

Evaluation sets

An exam loses its usefulness when the questions have been public for too long. Refreshing the tasks is part of evaluation.

Evaluation instances are separated from training data. Task distributions, verifiers, and contamination checks are versioned. Quarterly refreshes are part of the delivery plan, with a checking report accompanying each update.

08

Enterprise software beyond the US

Knowing how to work in one application does not mean you can walk into another company and finish the job.

Our environment roadmap includes Feishu, DingTalk, WeCom, Kingdee, Yonyou, 1688, Shopee, and Lazada. Starting with local business practices and workflows, we are building training scenarios across Chinese and other non-US software ecosystems.

09

Ready-to-use environments

A library takes generations to fill, but you can walk in and start reading tomorrow.

Initial tasks in education, marketing, office collaboration, mathematics, programming, and customer support are runnable today. Six environment families and eighteen scenario templates share state, action, reward, and replay interfaces for training and evaluation workflows.

Passing tests ≠ Doing the work.

Help agents learn to do real work correctly.