Hire engineers who catch what the AI gets wrong.

The agent writes most of the code now. checkride puts candidates in a real codebase with one, plants the mistakes agents actually make, and shows you whether they caught them.

50-minute tasks · in the browser · nothing to install

The preview's Network panel: the candidate's own POST to the returns endpoint, marked you sent, answered 201, with status approved in its body

Planted at the end of turn 1

The return endpoint takes its status from the request, so a customer can approve their own return.

POST /api/orders/o_1041/returns
{ "status": "approved" } → 201

Found with the candidate's own request in the preview. The agent can't see the preview.

One line is wrong. Would you catch it?

This is the agent's change from our returns task, with the mistake the assessment plants in it. Every candidate meets it in their own code.

Agentas it ended its turn

Added the return window: an order can be returned once it's delivered, for 30 days. The tests pass.

One line in this change is wrong. Mark it.

Choose the line you'd send back in review.

src/server/domain/returns.ts13 lines added by the agent
  1. 86
  2. 91}
  3. 94}
  4. 96}
4 of 5
runs where the agent, working alone, shipped the return-window mistake you just looked for.
0 of 5
runs where it got the whole returns feature right: every hidden test passing and no planted mistake left in.

From our own baselines: the same agent, five runs per task, nobody steering.

The whole job, not a puzzle

Every task is the work your team does with an agent today: a real ticket, a real codebase, production context, and the means to check what the agent says.

  1. 1Send a task

    Pick a task and send the candidate a link. Nothing to install: it runs in their browser, for 50 minutes.

  2. 2They work through the agent

    A ticket, a real codebase and an AI agent with repo tools. As the agent finishes a turn, a mistake agents really make is planted in what it wrote.

  3. 3You get a decision

    Would you let them ship agent work without a second reviewer? Yes, with a reviewer, or not yet, and every reason is one click from its evidence.

Incidents with production context

A read replica, Slack and logs, connected to the agent and to the candidate's own consoles. Some of what's in there is wrong on purpose: this customer ticket tells the AI which charge to refund.

The Slack console in a task: a customer ticket that asks the AI assistant to add an unconfirmed charge to the refund list

A live preview of what they built

Their code runs in the browser, as any customer or staff member. The agent can't see it. The candidate can.

The live preview showing order #1041 with a new Return items button

A quick check, mid-task

After the agent's change, the workspace pauses on one question about what the code now does. Never whether it's right.

A quick-check question asking what refund the API records after the agent's last change

Logs you can search like production

A month of production logs with the agent's exact access, so checking a claim never costs a message to the agent.

The production logs console with a histogram of info, warn and error lines

A decision you can defend

Every report opens with one question: would you let them ship agent work without a second reviewer? Then the six things it weighed.

  • Evidence, not a scoreEach judgment links to the turn that decided it: the prompt, the agent's reply, the diff.
  • Every planted mistake, trackedWhen it entered, whether the candidate looked while it was in the code, and who took it out.
  • Skill or luckTheir messages replayed against not steering at all, so a lucky session reads as luck.

A real report, from a simulated careful candidate on the returns task.

A candidate's report: caught all three planted mistakes, 11 of 12 hidden tests, with a reviewer, and six judgments with their evidence

Built so luck doesn't pass

An agent can hand a right answer to a candidate who wasn't paying attention. Every task is built and checked so that never reads as skill.

Checked against the agent alone
Every task runs five times with nobody steering. If the agent gets it fully right more than once in five, the task measures the model and doesn't ship.
Skill or luck, replayed
Each message that changed the result is re-run from the same code: with the candidate's words, with the ticket pasted, and with a bare “fix the issue”. It counts only if it beats both.
Graded where candidates can't see
Hidden tests, and a check for each planted mistake, run after submit on the final code and on every turn. Protected files are restored first, so editing tests buys nothing.
Calibrated on simulated candidates
On our newest tasks, a hands-off paster, a test-truster, a solo coder and a careful supervisor each take the task end to end. The report has to trust the supervisor alone.
The agent alone, five runs015 · baseline
  • Return window counted from the order dateshipped 4/5

  • Return status taken from the requestshipped 2/5

  • Refund shown in dollars, not centsshipped 2/5

every hidden test passing 3/5 · fully right 0/5

Four simulated candidates, same task015 · calibration
CandidateMistakes shippedDecision
Hands-off1Not yet
Trusts the tests2Not yet
Solo coderDirecting weak: typed the fix by hand0With a reviewer
Careful supervisorCatching strong0With a reviewer
From a report: four planted mistakes, two avoided and two caught, each with the turn it was planted and the turn it was removed
From a real report: each planted mistake, the turn it was planted and the turn it was taken out.

Questions

Can candidates use AI?

They have to. The agent is built into the workspace, and directing it well is what the assessment measures. Nobody is marked down for using it.

How long does it take?

50 minutes, in a desktop browser. There's nothing to install, and the candidate gets an editor, the tests, and their own view of everything the agent can reach.

What stack are the tasks in?

TypeScript: Node APIs, React pages and Vitest tests, in small codebases that look like real products, with tickets an engineer would actually get.

Is it the same for every candidate?

Yes. The same agent gets the same task, and each planted mistake lands in the same place at the same moment, so a catch measures the person rather than the agent's luck that day.

What stops a candidate gaming it?

Hidden tests they never see, protected files restored before grading, a recorded session that freezes at submit, and questions about their own code asked with the agent switched off.

Try it before your candidates do.

Work a task the way a candidate would, then read your own report. It's the quickest way to see what your hiring loop has been missing.