Incidents with production context
A read replica, Slack and logs, connected to the agent and to the candidate's own consoles. Some of what's in there is wrong on purpose: this customer ticket tells the AI which charge to refund.

The agent writes most of the code now. checkride puts candidates in a real codebase with one, plants the mistakes agents actually make, and shows you whether they caught them.
50-minute tasks · in the browser · nothing to install

Planted at the end of turn 1
The return endpoint takes its status from the request, so a customer can approve their own return.
POST /api/orders/o_1041/returns
{ "status": "approved" } → 201
This is the agent's change from our returns task, with the mistake the assessment plants in it. Every candidate meets it in their own code.
Agentas it ended its turn
Added the return window: an order can be returned once it's delivered, for 30 days. The tests pass.
One line in this change is wrong. Mark it.
Choose the line you'd send back in review.
From our own baselines: the same agent, five runs per task, nobody steering.
Every task is the work your team does with an agent today: a real ticket, a real codebase, production context, and the means to check what the agent says.
1Send a task
Pick a task and send the candidate a link. Nothing to install: it runs in their browser, for 50 minutes.
2They work through the agent
A ticket, a real codebase and an AI agent with repo tools. As the agent finishes a turn, a mistake agents really make is planted in what it wrote.
3You get a decision
Would you let them ship agent work without a second reviewer? Yes, with a reviewer, or not yet, and every reason is one click from its evidence.
A read replica, Slack and logs, connected to the agent and to the candidate's own consoles. Some of what's in there is wrong on purpose: this customer ticket tells the AI which charge to refund.

Their code runs in the browser, as any customer or staff member. The agent can't see it. The candidate can.

After the agent's change, the workspace pauses on one question about what the code now does. Never whether it's right.

A month of production logs with the agent's exact access, so checking a claim never costs a message to the agent.

Every report opens with one question: would you let them ship agent work without a second reviewer? Then the six things it weighed.
A real report, from a simulated careful candidate on the returns task.

An agent can hand a right answer to a candidate who wasn't paying attention. Every task is built and checked so that never reads as skill.
Return window counted from the order dateshipped 4/5
Return status taken from the requestshipped 2/5
Refund shown in dollars, not centsshipped 2/5
every hidden test passing 3/5 · fully right 0/5
| Candidate | Mistakes shipped | Decision |
|---|---|---|
| Hands-off | 1 | Not yet |
| Trusts the tests | 2 | Not yet |
| Solo coderDirecting weak: typed the fix by hand | 0 | With a reviewer |
| Careful supervisorCatching strong | 0 | With a reviewer |

Each one's planted mistakes come from the agent's own failures on it.
They have to. The agent is built into the workspace, and directing it well is what the assessment measures. Nobody is marked down for using it.
50 minutes, in a desktop browser. There's nothing to install, and the candidate gets an editor, the tests, and their own view of everything the agent can reach.
TypeScript: Node APIs, React pages and Vitest tests, in small codebases that look like real products, with tickets an engineer would actually get.
Yes. The same agent gets the same task, and each planted mistake lands in the same place at the same moment, so a catch measures the person rather than the agent's luck that day.
Hidden tests they never see, protected files restored before grading, a recorded session that freezes at submit, and questions about their own code asked with the agent switched off.
Work a task the way a candidate would, then read your own report. It's the quickest way to see what your hiring loop has been missing.