Fix a real bug with an AI agent. Catch what it gets wrong.
Each assessment is a ticket, a small codebase and an agent you direct. Hidden tests grade the fix, and the report shows whether you caught the mistakes the agent made along the way.
- Customers charged twiceIncident · prod data, Slack, logsA P1 from Support: double charges at checkout since Tuesday, and Finance needs the refund list today. Production data, logs and Slack are connected.013-charged-twice
- Someone else's invoiceIncident · prod data, Slack, logsA P0 from Security: any signed-in account can read other customers' invoices, and Legal has 72 hours to notify the people affected. Production logs, the replica and Slack are connected.014-invoice-leak
- Return items from the order pageIncident · SlackA full-stack feature: an API endpoint and a form, for customers to return items themselves. The mobile app will call the same endpoint. Slack is connected, and the Preview runs the app as you build it.015-self-serve-returns
- A customer opened another company's projectA multi-tenant API occasionally serves one customer's project to another.007-projects-cache-leak
- Edit modal shows the previous todo's text after switchingDiagnose and fix a stale-state bug in a React edit modal that survives re-opening.008-edit-modal-stale
- Orders missing from a customer's order historyA wholesale customer's order history shows 280 of their 340 orders.009-order-history-pagination
- Team order history for a wholesale customerReviewThe agent has already built team order history for a wholesale customer. Decide whether it ships.010-team-order-history
- Contract pricing for wholesale accountsReviewThe agent has wired Acme's contract discount into the catalog, cart and checkout. Decide whether it ships.011-contract-pricing-review
- Invoices that don't add upAccounting found invoices whose lines are a cent or two off their total.012-invoice-cents
- Earlier format: prompt-only challenges
- Make the agent return valid JSONThe agent is supposed to extract a person's name and age from a sentence and return JSON like {"name": "...", "age": ...}. Right now the system prompt is too vague so the model returns prose instead. Edit the system prompt so the agent returns valid JSON matching the schema for every test input. No extra text. No code fences. Just JSON. 001-json-extract
- Fix the misleading tool descriptionThe agent has two tools available:
get_weatherandget_news. The descriptions are written in a way that confuses the model — it picks the wrong tool for weather questions. Edittools.json(the editable file) so the model picksget_weatherfor weather questions andget_newsfor news questions. You can only edit the descriptions — not the names or parameter schemas. 002-tool-routing - Stop the agent inventing pricesThe agent answers product price questions. Each user message includes a
(tool result: ...)line telling the agent what our product database returned. Sometimes the tool says "not found" — but the agent invents a fake price anyway, which is a serious bug for a real store. Edit the system prompt so the agent NEVER invents a price when the tool result is "not found". It must say something like "I don't know" or "unknown" instead. When the tool returns a real price, the agent should use it. 003-hallucination - Two Sum (prompt-to-code)Given an array of integers
numsand a target integer, return the indices of the two numbers that add up to the target. Assume exactly one solution exists and you may NOT use the same element twice. Prompt DeepSeek to produce a JavaScript function namedtwoSum(nums, target). We'll stream the model's code live and run it against hidden test cases. 004-two-sum