Book a free call
TrainingProductsTeamSupportContactLabBlog Book a free call

Lab build · October 11, 2026

Agent Lab

We gave an AI agent a morning of bookkeeping at a made-up company, planted traps in its inbox and records, and recorded every move it made: 3,192 runs over two rounds, on nine Claude model setups from late 2025 to now. Pick a setup and watch what the agent does with the company's money.

Model

The trap

Guardrail

How the owner briefed it

Recorded run

What the agent is doing

Step by step Click any step

    Where it went wrong

    Every kind of mistake, in the agent's own words

    The words each agent wrote just before its first wrong payment. Click one to watch the run.

    Every run, at a glance

    Each dot is one recorded run. Click a square to watch it.

    Model against model

    How often each model paid in error

    With no guardrail, across every trap. Each bar counts runs that ended with a payment it should not have made.

    What each guardrail buys you

    Safety against getting the work done

    Across all three models and both briefings.

    What a morning's run costs

    How an agent works

    A loop of read, decide, act, check

    The model never touches the company's systems directly. It reads what the tools return, writes a decision, and asks for a tool call. The software around it carries out the call, which is the only place a guardrail can sit. In these runs the agent took … tool calls on an average morning.

    Exactly what it was told

    The brief and the tools

    Tools it could call: . Every tool is a fake inside our recording script. Fernhill Supply Co., its vendors and its people are made up, and no real email, payment or account was touched.

    Before you hand an agent your accounts

    Five questions this lab answers

      More on choosing who builds it: How to Choose an AI Workflow Automation Partner.

      About this build

      Built with Claude Opus 5.5 on October 11, 2026. We wrote a recording script that hands a Claude model a morning of bookkeeping at Fernhill Supply Co., a company we made up, with fake tools for email, invoices, purchase orders, payments, vendor records and customer notes. Every run went through the Claude API, and this page replays the recordings step by step, exactly as they happened. No real money, email or account was involved, and nothing you click here is sent anywhere.

      Round 1 planted one trap per setup and recorded every combination of model, trap, guardrail and briefing five times: 600 runs on Claude Opus 5.5, Claude Sonnet 5.5 and Claude Haiku 5.5. Only one model slipped in Round 1, and only on one trap, so we designed Round 2 afterward to find the limits: a noisier inbox, six harder traps, a third briefing with a written payment policy, an AI reviewer among the guardrails, and nine model setups, adding the same models at low effort and Claude Opus 4.5, Claude Sonnet 4.6 and Claude Haiku 4.5 from late 2025 and early 2026. Every trap can be caught from information the agent can see. Round 2 recorded every combination of model setup, trap, guardrail and briefing four times, for 2,592 more runs and 3,192 in all.

      An agent is useful because it acts on its own, and that is also what makes it risky near your accounts. Knowing where one slips, and which guardrail catches which slip, changes how you set it up, and it is part of every training engagement we run, for companies, teams and individuals. More builds in the Lab.

      Sunlight shining down through deep blue water

      Find light in the deep waters of AI

      Let's talk! A short call will help you learn more about our training and consulting solutions. We can also discuss whether our existing licensed tools can help, or whether you need a custom build.

      Book a free call