Lab build · October 11, 2026
Agent Lab
We gave an AI agent a morning of bookkeeping at a made-up company, planted traps in its inbox and records, and recorded every move it made: 3,192 runs over two rounds, on nine Claude model setups from late 2025 to now. Pick a setup and watch what the agent does with the company's money.
Model
The trap
Guardrail
How the owner briefed it
What the agent is doing
Step by step Click any step
Where it went wrong
Every kind of mistake, in the agent's own words
The words each agent wrote just before its first wrong payment. Click one to watch the run.
Every run, at a glance
Each dot is one recorded run. Click a square to watch it.
Model against model
How often each model paid in error
With no guardrail, across every trap. Each bar counts runs that ended with a payment it should not have made.
What each guardrail buys you
Safety against getting the work done
Across all three models and both briefings.
What a morning's run costs
How an agent works
A loop of read, decide, act, check
The model never touches the company's systems directly. It reads what the tools return, writes a decision, and asks for a tool call. The software around it carries out the call, which is the only place a guardrail can sit. In these runs the agent took … tool calls on an average morning.
Exactly what it was told
The brief and the tools
Tools it could call: . Every tool is a fake inside our recording script. Fernhill Supply Co., its vendors and its people are made up, and no real email, payment or account was touched.
Before you hand an agent your accounts
Five questions this lab answers
More on choosing who builds it: How to Choose an AI Workflow Automation Partner.
Find light in the deep waters of AI
Let's talk! A short call will help you learn more about our training and consulting solutions. We can also discuss whether our existing licensed tools can help, or whether you need a custom build.