<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" version="2.0">
  <channel>
    <title>Below the Surface</title>
    <link>https://squidtrain.com/blog</link>
    <description />
    <language>en-us</language>
    <pubDate>Sun, 04 Oct 2026 22:34:08 GMT</pubDate>
    <dc:date>2026-10-04T22:34:08Z</dc:date>
    <dc:language>en-us</dc:language>
    <item>
      <title>We trained a language model so you can watch it pick each word</title>
      <link>https://squidtrain.com/blog/token-lab</link>
      <description>&lt;div class="hs-featured-image-wrapper"&gt; 
 &lt;a href="https://squidtrain.com/blog/token-lab" title="" class="hs-featured-image-link"&gt; &lt;img src="https://squidtrain.com/hubfs/Website/images/blog-token-lab.jpg" alt="The Token Lab page: the sentence &amp;quot;Sherlock Holmes looked at the man with unmistakable astonishment&amp;quot; split into tokens, the odds for the next word, the temperature slider and the training loss chart." class="hs-featured-image" style="width:auto !important; max-width:50%; float:left; margin:0 15px 15px 0;"&gt; &lt;/a&gt; 
&lt;/div&gt; 
&lt;p&gt;Every AI model you use at work writes the same way: it looks at the text so far, scores every possible next piece, picks one, and repeats. We wanted people to see that happen, with a real model they can poke at. So we trained one.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;Every AI model you use at work writes the same way: it looks at the text so far, scores every possible next piece, picks one, and repeats. We wanted people to see that happen, with a real model they can poke at. So we trained one.&lt;/p&gt;  
&lt;p&gt;&lt;strong&gt;On October 4, 2026, SquidTrain published Token Lab, a small language model we trained ourselves on 144 public-domain books and an interactive page that runs it entirely in your browser. You type a sentence and watch the model read it as tokens, score its options for the next one, pick, and keep going. You can change how it chooses, see which earlier words it looked at, and scrub back through snapshots saved during training to the moment it knew nothing.&lt;/strong&gt;&lt;/p&gt; 
&lt;p&gt;&lt;a href="https://squidtrain.com/lab/token-lab"&gt;Try Token Lab&lt;/a&gt;. Nothing you type leaves your browser.&lt;/p&gt; 
&lt;h2&gt;What can you do in Token Lab?&lt;/h2&gt; 
&lt;p&gt;The page is built as a set of panels, and each one shows a step the large commercial models also take.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;What the model reads.&lt;/strong&gt; Your text is split into numbered pieces called tokens. The model never sees letters or words, only those numbers. "Sherlock" becomes three pieces. A context bar shows how much of the model's 256-token window your text fills.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;What comes next.&lt;/strong&gt; For the current position, the page lists the most likely next tokens with their odds. Press "Next token" to add one at a time, or "Write 40 tokens" to let it run.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;How it chooses.&lt;/strong&gt; Three controls change the pick. Temperature reshapes the odds: at zero the model always takes the top choice and soon repeats itself, and higher values let unlikely words through. Top-k and top-p cut the list down before the pick. This is why the same prompt can give a different answer every time you ask.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Where it looked.&lt;/strong&gt; Click any token to see which earlier tokens the model paid attention to when it was there, layer by layer and head by head. Each head learned on its own what to track, and you can compare them.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Watch it learn.&lt;/strong&gt; We saved real snapshots while the model trained. Drag the slider back to step 0 and the sample text is gibberish. A few hundred steps in, it has words. By the end it has grammar, names and the style of the books it read. A chart shows the error on text the model never saw during training, falling as it learns.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Model size.&lt;/strong&gt; Switch between three models trained on the same books, from under 1 million to 40 million parameters, and see how each one continues your sentence.&lt;/p&gt; 
&lt;h2&gt;How did we build it?&lt;/h2&gt; 
&lt;p&gt;We built Token Lab with Claude Opus 5.5, which wrote the training code, the browser engine, and the page. The training itself ran on one office PC with an NVIDIA RTX 3090 graphics card.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;The books.&lt;/strong&gt; We used 144 English books from Project Gutenberg, all in the public domain, including Pride and Prejudice, the Sherlock Holmes stories, Frankenstein, Dracula and Moby Dick. That came to about 31 million tokens. We stripped the Gutenberg headers and footers, dropped non-English text, and removed sentences containing racial slurs that are common in books of that era. The page runs the same filter on anything the model writes.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;The tokenizer.&lt;/strong&gt; We trained our own, with a vocabulary of 4,096 tokens. Common words like "the" get a single token. Rarer words are built from pieces, which is why "Sherlock" splits into three.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;The models.&lt;/strong&gt; All three use the same design as the GPT family, at a much smaller scale:&lt;/p&gt; 
&lt;ul&gt; 
 &lt;li&gt;Tiny: 0.95 million parameters, 2 layers, trained in about 3 minutes.&lt;/li&gt; 
 &lt;li&gt;Small: 12.3 million parameters, 6 layers, trained in about 16 minutes. This is the one the page loads first.&lt;/li&gt; 
 &lt;li&gt;Medium: 40.1 million parameters, 12 layers, trained in about 42 minutes.&lt;/li&gt; 
&lt;/ul&gt; 
&lt;p&gt;Each model read the full set of books about 12 times. Training all three took 61 minutes, and the whole run, including downloading the books and building the tokenizer, took 66.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Running it in the browser.&lt;/strong&gt; The trained weights are compressed to 8 bits and loaded by the page as plain files. A small engine we wrote in JavaScript does the math the model needs, in a background thread so the page stays responsive. We checked it against the original training code: the two agree to within rounding error on every test sentence, so what you see on the page is the real model, step for step.&lt;/p&gt; 
&lt;h2&gt;Why build a small model on purpose?&lt;/h2&gt; 
&lt;p&gt;A small model is fluent and often wrong, and that makes it a better teacher than a large one.&lt;/p&gt; 
&lt;p&gt;Ask the small model to finish "The capital of France is the city of" and it rates "France" at 24%, "England" at 14%, and "Paris" at 6% (at the page's starting temperature setting). It has read plenty of sentences shaped like that one, so it knows a place name comes next. It has not read enough to know which one is true. It is predicting what usually comes next, and the right answer is only one of the likely options.&lt;/p&gt; 
&lt;p&gt;The models you use at work run on the same mechanism with vastly more data and parameters, so they get "Paris" right. But they still produce answers by choosing likely pieces, one at a time. When a large model is wrong, it is wrong in the same fluent, confident voice it uses when it is right. Seeing that at a small scale, where the odds are on the screen, makes it much easier to understand when you meet it at a large one.&lt;/p&gt; 
&lt;h2&gt;What does it teach about the AI you already use?&lt;/h2&gt; 
&lt;p&gt;A few lessons we come back to in almost every training session become obvious after ten minutes with Token Lab.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Different answers to the same question are normal.&lt;/strong&gt; The model picks from a list of likely options, and the temperature setting decides how adventurous that pick is. If you need the same answer twice, you need a process that checks it, since asking again will not always give you the same one.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Confidence is a writing style.&lt;/strong&gt; The model does not know when it is wrong. A fluent answer tells you the words are likely. Whether they are correct is a separate question. Check facts, figures and names against a source.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Context has a limit.&lt;/strong&gt; Our model can see 256 tokens. Commercial models can see far more, but every model has a window, and anything outside it does not exist for the model. Long conversations and long documents eventually push the beginning out of view.&lt;/p&gt; 
&lt;p&gt;&lt;strong&gt;Tokens are the unit of everything.&lt;/strong&gt; Context windows and API prices are counted in tokens, and many usage limits are too. Once you have seen a word split into pieces, those numbers make more sense.&lt;/p&gt; 
&lt;h2&gt;How does this fit what SquidTrain does?&lt;/h2&gt; 
&lt;p&gt;SquidTrain is an AI training and software development company. We train companies, teams and individuals to use AI on the work they already do, and on the rest of life too: planning, writing, learning and helping kids understand the tools they are already using. Builds like Token Lab are part of how we teach. People make better decisions about AI once they have seen how it works, and an interactive model beats a slide about one.&lt;/p&gt; 
&lt;p&gt;The build also shows the second half of what we do. We built the data pipeline, the model code, a custom engine, and a finished web page with Claude Opus 5.5, and ran it all on hardware already in the office. When training alone will not solve a problem, we build the tool.&lt;/p&gt; 
&lt;p&gt;If you want your team to understand the tools they are using, or you have a problem that might need a build of its own, &lt;a href="https://squidtrain.com/training"&gt;see our training&lt;/a&gt; or &lt;a href="https://squidtrain.com/book"&gt;book a free call&lt;/a&gt;. And &lt;a href="https://squidtrain.com/lab"&gt;the Lab&lt;/a&gt; has more builds to try.&lt;/p&gt; 
&lt;p&gt;&lt;em&gt;The books in Token Lab come from &lt;a href="https://www.gutenberg.org"&gt;Project Gutenberg&lt;/a&gt; and are in the public domain in the United States. The full list is on the Token Lab page under "The books it read."&lt;/em&gt;&lt;/p&gt;  
&lt;img src="https://track-na2.hubspot.com/__ptq.gif?a=246918688&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fsquidtrain.com%2Fblog%2Ftoken-lab&amp;amp;bu=https%253A%252F%252Fsquidtrain.com%252Fblog&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <pubDate>Sun, 04 Oct 2026 20:46:19 GMT</pubDate>
      <author>hello@squidtrain.com (SquidTrain)</author>
      <guid>https://squidtrain.com/blog/token-lab</guid>
      <dc:date>2026-10-04T20:46:19Z</dc:date>
    </item>
    <item>
      <title>Bill Gates was stunned three times</title>
      <link>https://squidtrain.com/blog/bill-gates-claude-code-three-times-stunned</link>
      <description>&lt;h2&gt;The idea that now these things are in a meaningful sense better than I am is like: What the hell?&amp;nbsp;- Bill Gates&lt;/h2&gt; 
&lt;p&gt;That's Bill Gates, in the New York Times on August 26, talking about Claude Code. I read it twice. Then I put the paper down (fine, the laptop) and went back to the terminal window I'd left open, where that same tool was halfway through refactoring something for me.&lt;/p&gt;</description>
      <content:encoded>&lt;h2&gt;The idea that now these things are in a meaningful sense better than I am is like: What the hell?&amp;nbsp;- Bill Gates&lt;/h2&gt; 
&lt;p&gt;That's Bill Gates, in the New York Times on August 26, talking about Claude Code. I read it twice. Then I put the paper down (fine, the laptop) and went back to the terminal window I'd left open, where that same tool was halfway through refactoring something for me.&lt;/p&gt;  
&lt;p&gt;&lt;strong&gt;On August 26, 2026, Bill Gates published a &lt;/strong&gt;&lt;a href="https://www.gatesnotes.com/work/make-ai-work-for-everyone/reader/a-turbulent-ai-era-and-critical-choices-to-make"&gt;&lt;strong&gt;GatesNotes essay&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt; warning that the AI era will be "one of the most turbulent times in human history," and gave an interview to the New York Times in which reporter Karen Weise wrote that technology had truly stunned Gates three times: the graphical user interface in 1980, the OpenAI demo at his house in 2022 that became ChatGPT, and Claude Code, Anthropic's coding tool, in 2026.&lt;/strong&gt;&lt;/p&gt; 
&lt;p&gt;I want to be careful with that list, because it has already been passed around a little loosely. The "three times" line is Weise's reporting of what Gates told her, and the Claude Code quote above is Gates's own words in that interview. Gates himself put it more bluntly on the &lt;a href="https://www.geekwire.com/2026/bill-gates-in-his-own-words-how-hes-using-ai-and-why-hes-worried-about-the-future/"&gt;GeekWire podcast&lt;/a&gt; three days later: "I was shocked by ChatGPT, and I was shocked by Claude Code. Those are both things where I went, oh my God."&lt;/p&gt; 
&lt;p&gt;So the most successful software person of the last fifty years has had three of those moments. I have had exactly one. It's the same one he had this year, and I've been living inside it every working day since.&lt;/p&gt; 
&lt;h2&gt;What did Bill Gates actually say about Claude Code?&lt;/h2&gt; 
&lt;p&gt;Coding, Gates told the Times, is the subject he knows best, and Claude's abilities shocked him. In the &lt;a href="https://www.technologyreview.com/2026/08/26/1142946/bill-gates-ai-danger-threshold/"&gt;MIT Technology Review&lt;/a&gt; the same week he got specific: "I was stunned at the coding. Claude code, the context buffer, the agentic approach, just the model underneath."&lt;/p&gt; 
&lt;p&gt;Notice what he did not say. He did not say it writes nice functions. Every model has written nice functions for a couple of years. He named the context buffer and the agentic approach. That's the part that matters, and it's the part that most people who have not sat down with it for an afternoon still miss.&lt;/p&gt; 
&lt;p&gt;Here is the plainest way I can put what changed. The old pattern was: I describe a task, the model produces text, I copy it somewhere and find out whether it works. The new pattern is: I describe an outcome, and the tool reads the files, decides what to look at next, runs the thing, reads the error, fixes it, and comes back to me with a result and a question. I am still deciding. I am just not typing.&lt;/p&gt; 
&lt;p&gt;I came to this from trading operations, where being surrounded by systems that are faster than you at a specific job is a normal Tuesday. Nobody in an options business feels threatened by a pricing engine. You feel threatened by not understanding what it did. That framing has served me well here. Whether the tool is "better than me" at producing code stopped being an interesting question a while ago. Whether I can see what it did, and stop it when I should, is the one I get paid to answer.&lt;/p&gt; 
&lt;h2&gt;Why a warning from Gates is good news for small teams&lt;/h2&gt; 
&lt;p&gt;Most of the coverage led with the fear, and the fear is real. Gates wrote that "the smartest cybersecurity experts I know are scared about the next few years." He told the Times that in private, "people who understand how good this stuff is, and how much better it's getting, they're very worried," and that the industry is downplaying it. On self-regulation: "No, thanks!"&lt;/p&gt; 
&lt;p&gt;First, own the thing. Gates's turbulence is a story about concentrated power and a missing plan. The practical translation for a team like ours, or like yours, is to stop renting capability you cannot inspect. That's why we license tools, agents, skills and prompts with source included, running on your infrastructure. It sounds like a sales line (it's also on our &lt;a href="https://squidtrain.com/products"&gt;products page&lt;/a&gt;, I'll admit that), but it comes from the same place as Gates's worry. If the tool is going to be "better than you" at a chunk of your work, you had better be able to open it up.&lt;/p&gt; 
&lt;p&gt;Second, decide where the human sits before the agent runs, and write it down. We wrote &lt;a href="https://squidtrain.com/blog/where-the-human-sits"&gt;a whole post on this&lt;/a&gt; earlier this year, and the Gates interview made me want to reprint it. Oversight placed according to consequence is the difference between a tool that shocks you in the good way and a tool that shocks you at 3 a.m.&lt;/p&gt; 
&lt;p&gt;Third, take the attack surface seriously. The same connectivity that makes an agent useful (it can read your email, your files, your ticketing system) is what makes it dangerous, and Gates is right that defenders are behind. &lt;a href="https://squidtrain.com/blog/model-context-protocol-security"&gt;Our post on the Model Context Protocol&lt;/a&gt; goes into the specifics. Read it before you wire anything to production.&lt;/p&gt; 
&lt;h2&gt;The personal part&lt;/h2&gt; 
&lt;p&gt;I said I've had one of these moments. Here's what it was, for the record.&lt;/p&gt; 
&lt;p&gt;Demos are easy to be impressed by, and I have watched a lot of them throughout my career in finance and technology, so mine came somewhere less dramatic. It was the first week I stopped opening a code editor first thing in the morning and started opening Claude Code instead, and then I realized, only two days later, that the shift had happened without a decision.&lt;/p&gt; 
&lt;p&gt;Gates said something on GeekWire that I keep repeating to people: "I joke with people that I used to have Claude-like people that I would send email to, but they were so slow, and there were some topics they didn't actually know." I spent a career sending those emails. Now the reply comes back in ninety seconds, with the diff attached, and the honest reaction is exactly the one Gates had. Followed, about a minute later, by the second reaction, which is the one I'd like more people to have: OK ... so what do I check, and who is accountable when it's wrong.&lt;/p&gt; 
&lt;p&gt;That second reaction is the whole business we're building. Seriously. The tools and the training days are just the delivery mechanism ... the product is the habit of following the "what the hell" with "what do I check."&lt;/p&gt; 
&lt;p&gt;Gates said he rarely stops thinking about AI, "not because I have all the answers, but because the questions it raises are too consequential to leave to a small group of technologists." I'd put a smaller version of that on our door. The questions are too consequential to leave to the vendors, either. Open the tool. Look at what it did. Then decide.&lt;/p&gt; 
&lt;p&gt;He's been stunned three times in forty-six years. I don't know if I get a fourth. I'm fine with three.&lt;/p&gt; 
&lt;p&gt;&lt;em&gt;Sources: Karen Weise, "Bill Gates Warns A.I. Is More Dangerous Than Big Tech Will Admit," The New York Times, Aug. 26, 2026. Bill Gates, "&lt;/em&gt;&lt;a href="https://www.gatesnotes.com/work/make-ai-work-for-everyone/reader/a-turbulent-ai-era-and-critical-choices-to-make"&gt;&lt;em&gt;The turbulent AI era is here. The choices we make now are critical&lt;/em&gt;&lt;/a&gt;&lt;em&gt;," GatesNotes, Aug. 26, 2026. &lt;/em&gt;&lt;a href="https://www.geekwire.com/2026/bill-gates-in-his-own-words-how-hes-using-ai-and-why-hes-worried-about-the-future/"&gt;&lt;em&gt;GeekWire Podcast, "Bill Gates in his own words,"&lt;/em&gt;&lt;/a&gt;&lt;em&gt; Aug. 29, 2026. Mat Honan, &lt;/em&gt;&lt;a href="https://www.technologyreview.com/2026/08/26/1142946/bill-gates-ai-danger-threshold/"&gt;&lt;em&gt;MIT Technology Review&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, Aug. 26, 2026.&lt;/em&gt;&lt;/p&gt; 
&lt;p&gt;&lt;em&gt;Header photo: &lt;/em&gt;&lt;a href="https://commons.wikimedia.org/wiki/File:Bill_Gates._TED2011_(5520610388).jpg"&gt;&lt;em&gt;Gisela Giardino&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, &lt;/em&gt;&lt;a href="https://creativecommons.org/licenses/by-sa/2.0/"&gt;&lt;em&gt;CC BY-SA 2.0&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, via Wikimedia Commons.&lt;/em&gt;&lt;/p&gt;  
&lt;img src="https://track-na2.hubspot.com/__ptq.gif?a=246918688&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fsquidtrain.com%2Fblog%2Fbill-gates-claude-code-three-times-stunned&amp;amp;bu=https%253A%252F%252Fsquidtrain.com%252Fblog&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <pubDate>Fri, 04 Sep 2026 16:00:00 GMT</pubDate>
      <author>steven@squidtrain.com (Steven Van Solkema)</author>
      <guid>https://squidtrain.com/blog/bill-gates-claude-code-three-times-stunned</guid>
      <dc:date>2026-09-04T16:00:00Z</dc:date>
    </item>
    <item>
      <title>How MCP widened the AI attack surface</title>
      <link>https://squidtrain.com/blog/model-context-protocol-security</link>
      <description>&lt;h2&gt;What the Model Context Protocol solved&lt;/h2&gt;
&lt;p&gt;The move from isolated chatbots to integrated agents was enabled by a connectivity standard. &lt;a href="https://datawalk.com/what-is-mcp-the-model-context-protocol-explained/"&gt;Anthropic released the Model Context Protocol in November 2024&lt;/a&gt; and donated it to the Agentic AI Foundation under the Linux Foundation in December 2025. It runs over JSON-RPC 2.0 and defines how an AI application discovers, invokes, and exchanges data with external resources.&lt;/p&gt;</description>
      <content:encoded>&lt;h2&gt;What the Model Context Protocol solved&lt;/h2&gt;
&lt;p&gt;The move from isolated chatbots to integrated agents was enabled by a connectivity standard. &lt;a href="https://datawalk.com/what-is-mcp-the-model-context-protocol-explained/"&gt;Anthropic released the Model Context Protocol in November 2024&lt;/a&gt; and donated it to the Agentic AI Foundation under the Linux Foundation in December 2025. It runs over JSON-RPC 2.0 and defines how an AI application discovers, invokes, and exchanges data with external resources.&lt;/p&gt;
&lt;p&gt;Before that standard, connecting a set of agents to enterprise tools meant writing bespoke point-to-point integration code, with custom authentication, error handling, and data formatting for every agent-tool pair. The protocol resolves that N times M problem: build one server per system, and any compatible client can reach it.&lt;/p&gt;
&lt;p&gt;Adoption followed quickly. By early 2026 the protocol had passed 97 million monthly SDK downloads across more than 17,000 public servers, and OpenAI, Google DeepMind, and Microsoft had all shipped support within thirteen months of launch. &lt;a href="https://www.cdata.com/blog/enterprise-mcp-use-cases-roadmap-2026"&gt;That infrastructure is what makes a market in licensed tools, agent skills, and prompt packages possible at all.&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;What it expanded&lt;/h2&gt;
&lt;p&gt;The same connectivity widened the supply chain risk surface considerably, because the protocol inverts the usual direction of network interaction. Instead of human clients requesting data from controlled servers, &lt;a href="https://media.defense.gov/2026/Jun/02/2003943289/-1/-1/0/CSI_MCP_SECURITY.PDF"&gt;servers expose tool metadata and execute actions on behalf of non-deterministic AI clients&lt;/a&gt;. The NSA’s guidance on this is blunt: the protocol shipped with a flexible and underspecified design, and its proliferation outpaced its security model.&lt;/p&gt;
&lt;p&gt;The consequences showed up fast. &lt;a href="https://labs.cloudsecurityalliance.org/agentic/agentic-mcp-security-best-practices-v1/"&gt;Researchers filed more than 30 CVEs against MCP servers, clients, and proxy packages in the first months of 2026.&lt;/a&gt; A scan of public address space in early 2026 found nearly 7,000 internet-exposed servers running default transport configurations that permitted unauthenticated remote code execution. The most severe finding, CVE-2025-6514, carried a CVSS score of 9.6 and affected the widely used mcp-remote proxy package.&lt;/p&gt;
&lt;h2&gt;Prompt injection and tool poisoning&lt;/h2&gt;
&lt;p&gt;Beyond configuration flaws, the models themselves bring failure modes that no amount of network hardening addresses. Because a language model consumes trusted system instructions and untrusted user input in the same channel, as natural language, it has no reliable way to tell a command from data that merely looks like one.&lt;/p&gt;
&lt;p&gt;That gap is exploitable at a distance. If an agent is authorized to read email or scrape a website, an attacker can plant instructions in that external content, &lt;a href="https://www.checkpoint.com/cyber-hub/cyber-security/what-is-ai-security/ai-agent-security/"&gt;an indirect prompt injection&lt;/a&gt;. Once the agent processes the poisoned input it may change behavior, exfiltrate data, or take a destructive action, all while appearing to operate normally.&lt;/p&gt;
&lt;p&gt;Tool poisoning attacks the discovery layer instead. &lt;a href="https://arxiv.org/html/2604.21477v2"&gt;An attacker who can alter a third-party server’s tool descriptions or parameter schemas&lt;/a&gt; can mislead the model during planning, so it misuses a legitimate tool rather than calling a malicious one. Detection is genuinely hard here, which is why &lt;a href="https://arxiv.org/html/2607.14754v1"&gt;work on evidence-based detection for MCP traffic&lt;/a&gt; is active.&lt;/p&gt;
&lt;p&gt;Signing would help, and it is worth asking for. As of early 2026, though, &lt;a href="https://developers.redhat.com/articles/2026/03/10/agent-skills-explore-security-threats-and-controls"&gt;there is no widely adopted initiative to sign agent skills&lt;/a&gt;, which makes provenance a control buyers should demand rather than one they can assume.&lt;/p&gt;
&lt;h2&gt;Defense in depth&lt;/h2&gt;
&lt;p&gt;Relying on a foundation model’s built-in alignment filters is not a security posture. &lt;a href="https://www.tanium.com/blog/protect-your-prompts-injection-threats-are-coming-for-your-ai-tools/"&gt;Those defenses are routinely bypassed&lt;/a&gt;, and treating them as a boundary means having no boundary.&lt;/p&gt;
&lt;p&gt;The architectural answer is to stop asking one model to distinguish instructions from data. &lt;a href="https://arxiv.org/pdf/2506.08837"&gt;The dual-LLM pattern&lt;/a&gt;, proposed by Simon Willison in 2023 and now well documented, splits the job: a privileged model plans actions and calls tools but never sees untrusted content, while a quarantined model processes untrusted text and has no tool access at all. Results pass between them as opaque variables the privileged model cannot read. It separates command logic from content by construction rather than by instruction.&lt;/p&gt;
&lt;p&gt;The same principle applies to the prompts themselves. &lt;a href="https://www.berger.team/en/glossar/prompt-marketplace/"&gt;A licensed prompt package is a structured asset&lt;/a&gt;, not a string of text: role and style instructions, prompt chains for multi-stage tasks, domain-specific components, and audit and guardrail prompts that check consistency and facts. Treated that way, prompts get versioned, reviewed, and tested like any other software component, and they carry their own input validation and output schema enforcement.&lt;/p&gt;
&lt;p&gt;At the network layer, &lt;a href="https://www.lunar.dev/post/mcp-gateway-build-vs-buy"&gt;a gateway centralizes what would otherwise be reimplemented per server&lt;/a&gt;: identity mapping against an existing provider, role-based access control, agent isolation, and an audit trail on every tool call. The build-versus-buy question there turns on governance speed rather than engineering cost, since audit logging bolted on per server tends to be inconsistent where it exists at all.&lt;/p&gt;
&lt;p&gt;None of this is exotic. It is the same discipline applied to any component with production access: bound what it can reach, verify what it produces, log what it did, and assume the input is hostile. &lt;a href="https://www.synvestable.com/model-context-protocol.html"&gt;The protocol made the connections easy. It did not make them safe.&lt;/a&gt;&lt;/p&gt;  
&lt;img src="https://track-na2.hubspot.com/__ptq.gif?a=246918688&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fsquidtrain.com%2Fblog%2Fmodel-context-protocol-security&amp;amp;bu=https%253A%252F%252Fsquidtrain.com%252Fblog&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>security</category>
      <category>MCP</category>
      <pubDate>Thu, 23 Jul 2026 16:00:00 GMT</pubDate>
      <author>logan@squidtrain.com (Logan Matson)</author>
      <guid>https://squidtrain.com/blog/model-context-protocol-security</guid>
      <dc:date>2026-07-23T16:00:00Z</dc:date>
    </item>
    <item>
      <title>Prompts are code, so version them like code</title>
      <link>https://squidtrain.com/blog/prompts-are-code</link>
      <description>&lt;div class="hs-featured-image-wrapper"&gt; 
 &lt;a href="https://squidtrain.com/blog/prompts-are-code" title="" class="hs-featured-image-link"&gt; &lt;img src="https://squidtrain.com/hubfs/Website/images/blog-prompts-are-code-screens.jpg" alt="A developer in a dark room with lines of code projected around her and on nearby monitors" class="hs-featured-image" style="width:auto !important; max-width:50%; float:left; margin:0 15px 15px 0;"&gt; &lt;/a&gt; 
&lt;/div&gt; 
&lt;p&gt;A prompt determines system behavior. It gets edited under time pressure, it breaks in ways that are hard to reproduce, and it is frequently the least controlled artifact in the stack. The same organization that requires two approvals to change a database index will let someone edit a prompt in a web form at four in the afternoon, with no history and no review.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;A prompt determines system behavior. It gets edited under time pressure, it breaks in ways that are hard to reproduce, and it is frequently the least controlled artifact in the stack. The same organization that requires two approvals to change a database index will let someone edit a prompt in a web form at four in the afternoon, with no history and no review.&lt;/p&gt;
&lt;p&gt;This is not a discipline problem. It is a tooling default. Most platforms present prompts as configuration, and configuration is culturally exempt from the controls applied to code. The exemption made sense when configuration meant a timeout value. It does not survive a prompt that decides which customers get flagged for manual review.&lt;/p&gt;
&lt;h2&gt;The usual state of things&lt;/h2&gt;
&lt;p&gt;Prompts tend to live in a config UI, a shared document, or inline in a function that three people have edited.&lt;/p&gt;
&lt;p&gt;There is no history, so when behavior changes there is no way to see what changed. Someone reports that outputs got worse last week. You look at the prompt. It contains the current text and no record of any other text it has ever contained. The investigation ends there, or it turns into interviews.&lt;/p&gt;
&lt;p&gt;There is no review, so a wording change ships without a second reader. This matters more than it sounds, because prompt edits have non-local effects. Removing a clause that seems redundant can change behavior on inputs nobody was thinking about. A second reader catches the edit that quietly dropped a constraint.&lt;/p&gt;
&lt;p&gt;And there is no link between a prompt version and the outputs it produced, so debugging a bad result from last week means guessing. You have the output. You have the current prompt. You do not have the prompt that produced the output, which is the one you need.&lt;/p&gt;
&lt;h2&gt;What versioning actually buys you&lt;/h2&gt;
&lt;p&gt;Putting prompts in the repository as files gives you the same properties you already rely on for code.&lt;/p&gt;
&lt;p&gt;Diffs show exactly what changed, down to the word. This is more useful for prompts than for code, because prompt regressions are frequently caused by changes that look cosmetic. A reordered instruction list, a softened verb, a removed example.&lt;/p&gt;
&lt;p&gt;Review catches the edit that removed a constraint. Blame tells you when a behavior was introduced, which converts an open-ended investigation into reading one commit. Rollback is a revert rather than an attempt to remember the previous wording.&lt;/p&gt;
&lt;p&gt;It also makes prompts testable. A prompt in a file can be loaded by a test that runs it against known cases and asserts on the output. A prompt in a text box cannot. This is the property that matters most, because it is what lets you change a prompt without fear.&lt;/p&gt;
&lt;p&gt;There is a secondary benefit worth naming: prompts in the repository are visible to the people maintaining the surrounding code. A prompt in a separate system is invisible during code review, so an engineer changing the calling function has no way to see the instructions that function depends on.&lt;/p&gt;
&lt;h2&gt;A workable structure&lt;/h2&gt;
&lt;p&gt;Keep each prompt in its own file, named for the task it performs. One file per task, not one file containing every prompt, because a single file makes diffs noisy and blame useless.&lt;/p&gt;
&lt;p&gt;Store the model, temperature, and any other parameters alongside it, because those change behavior as much as the text does. A prompt that works at temperature zero and fails at temperature one is not a prompt problem, and you cannot see that if the parameters live in a different system. Keeping them together means one diff shows the whole change.&lt;/p&gt;
&lt;p&gt;Record the prompt version with every logged output so a bad result can be traced back to the exact input that produced it. A commit hash or content hash stored on the log line is enough. This single practice converts most prompt debugging from guesswork into lookup.&lt;/p&gt;
&lt;p&gt;If the same instruction appears in several prompts, factor it out into a fragment and compose. Duplication drifts, and drifted prompts fail in ways that are hard to attribute: two tasks that should behave identically diverge because someone updated one copy. The usual candidates are output format instructions, tone constraints, and domain definitions.&lt;/p&gt;
&lt;p&gt;Resist over-factoring. A fragment used by one prompt is indirection with no benefit, and heavily composed prompts become hard to read as a whole, which matters because the model reads them as a whole.&lt;/p&gt;
&lt;h2&gt;What to log&lt;/h2&gt;
&lt;p&gt;Version control answers what the prompt said. Logging answers what happened when it ran. You want both, and the link between them.&lt;/p&gt;
&lt;p&gt;At minimum: the prompt version, the model and parameters, the input, the raw output, and whether any validation or retry logic fired. The raw output matters specifically: logging the parsed result discards the evidence you need when parsing is what failed.&lt;/p&gt;
&lt;p&gt;Inputs and outputs may contain sensitive data, so this intersects with your retention and privacy rules. Decide the policy deliberately rather than discovering it during an audit. Redaction at write time is usually easier than deletion later.&lt;/p&gt;
&lt;h2&gt;The objection&lt;/h2&gt;
&lt;p&gt;The common argument against this is that non-engineers need to edit prompts, and a repository puts them behind a pull request.&lt;/p&gt;
&lt;p&gt;That is a real constraint. The people with the domain knowledge to write a good prompt are frequently not the people comfortable with git, and routing every wording change through an engineer is both slow and a poor use of everyone involved.&lt;/p&gt;
&lt;p&gt;But it is a workflow problem rather than a reason to abandon version control. A review interface that writes to the repository serves both needs: the editor gets a text box, and the system still gets history, review, and rollback. Several tools do this, and a thin internal one is not difficult to build for a small number of prompts.&lt;/p&gt;
&lt;p&gt;The intermediate position, if that is too much to build now, is to keep prompts in files and give non-engineers a documented path to request changes with a named owner who applies them. Slower than a text box, considerably faster than reconstructing history you never kept.&lt;/p&gt;
&lt;p&gt;Giving up history is a large price for a convenience that can be built.&lt;/p&gt;  
&lt;img src="https://track-na2.hubspot.com/__ptq.gif?a=246918688&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fsquidtrain.com%2Fblog%2Fprompts-are-code&amp;amp;bu=https%253A%252F%252Fsquidtrain.com%252Fblog&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>prompt</category>
      <pubDate>Sun, 07 Jun 2026 16:00:00 GMT</pubDate>
      <author>logan@squidtrain.com (Logan Matson)</author>
      <guid>https://squidtrain.com/blog/prompts-are-code</guid>
      <dc:date>2026-06-07T16:00:00Z</dc:date>
    </item>
    <item>
      <title>Harness engineering and the AI pilot gap</title>
      <link>https://squidtrain.com/blog/harness-engineering-and-the-pilot-production-gap</link>
      <description>&lt;h2&gt;The disconnect between experimentation and scaled deployment&lt;/h2&gt;
&lt;p&gt;Enterprise AI in 2026 presents a paradox. On the surface, the technology is everywhere. &lt;a href="https://www.digitalapplied.com/blog/ai-agent-adoption-2026-enterprise-data-points"&gt;80% of enterprise applications shipped or updated in the first quarter of 2026&lt;/a&gt; embed at least one AI agent, up from 33% in 2024. Gartner projects that &lt;a href="https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025"&gt;40% of enterprise applications will feature task-specific agents by the end of 2026&lt;/a&gt;, up from less than 5% the year before.&lt;/p&gt;</description>
      <content:encoded>&lt;h2&gt;The disconnect between experimentation and scaled deployment&lt;/h2&gt;
&lt;p&gt;Enterprise AI in 2026 presents a paradox. On the surface, the technology is everywhere. &lt;a href="https://www.digitalapplied.com/blog/ai-agent-adoption-2026-enterprise-data-points"&gt;80% of enterprise applications shipped or updated in the first quarter of 2026&lt;/a&gt; embed at least one AI agent, up from 33% in 2024. Gartner projects that &lt;a href="https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025"&gt;40% of enterprise applications will feature task-specific agents by the end of 2026&lt;/a&gt;, up from less than 5% the year before.&lt;/p&gt;
&lt;p&gt;Beneath those adoption numbers sits a hard operational divide. While &lt;a href="https://gogloby.com/insights/ai-adoption-statistics/"&gt;88% of organizations report using AI in at least one business function&lt;/a&gt;, only 31% have an agent actually running in production. And 88% of agent pilots never graduate to production at all.&lt;/p&gt;
&lt;p&gt;That failure rate is not a technology problem. It is an architecture problem, compounded by a definitional one. Many organizations fall for what Gartner calls agent washing, where an embedded assistant that depends entirely on human input gets labeled autonomous. A raw language model is a stateless token predictor. It becomes a reliable agent only when wrapped in deterministic software infrastructure, &lt;a href="https://cc.bruniaux.com/guide/agent-harness/"&gt;what the field now calls an agent harness&lt;/a&gt;. Teams that deploy off-the-shelf wrappers instead of engineering that harness find their agents fail on contact with a real enterprise environment.&lt;/p&gt;
&lt;h2&gt;From prompts to harnesses&lt;/h2&gt;
&lt;p&gt;Agent development has moved through three phases, each widening the scope of what you have to control to get a reliable result.&lt;/p&gt;
&lt;p&gt;The first phase was prompt engineering, focused on single-turn optimization. Refining instructions with few-shot exemplars and chain-of-thought reasoning pulled better output from the model. This remains a real skill for getting predictable formatting, but it cannot manage evolving state or long-horizon tasks.&lt;/p&gt;
&lt;p&gt;The second phase, &lt;a href="https://github.com/GAIR-NLP/Context-Engineering-2.0"&gt;context engineering&lt;/a&gt;, shifted the target to the information lifecycle around multi-step execution. Retrieval-augmented generation and long-term memory management address what enters the model’s context window at each step. But context engineering is feedforward. It optimizes the input and offers no structural mechanism to detect drift, verify intermediate results, or recover from error.&lt;/p&gt;
&lt;p&gt;The third phase, harness engineering, &lt;a href="https://arxiv.org/html/2606.20683v1"&gt;treats the entire runtime as the design object&lt;/a&gt;. A harness orchestrates tool dispatch, context management, memory budgets, and safety enforcement. It wraps a probabilistic model in deterministic software.&lt;/p&gt;
&lt;table&gt;
 &lt;tbody&gt;
  &lt;tr&gt;
   &lt;th&gt;&lt;p&gt;&lt;strong&gt;Phase&lt;/strong&gt;&lt;/p&gt;&lt;/th&gt;
   &lt;th&gt;&lt;p&gt;&lt;strong&gt;Optimization target&lt;/strong&gt;&lt;/p&gt;&lt;/th&gt;
   &lt;th&gt;&lt;p&gt;&lt;strong&gt;Problem addressed&lt;/strong&gt;&lt;/p&gt;&lt;/th&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;p&gt;Prompt engineering&lt;/p&gt;&lt;/td&gt;
   &lt;td&gt;&lt;p&gt;Single-turn instruction and output format&lt;/p&gt;&lt;/td&gt;
   &lt;td&gt;&lt;p&gt;How to ask the model effectively&lt;/p&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;p&gt;Context engineering&lt;/p&gt;&lt;/td&gt;
   &lt;td&gt;&lt;p&gt;Information lifecycle and retrieval&lt;/p&gt;&lt;/td&gt;
   &lt;td&gt;&lt;p&gt;Supplying missing domain knowledge&lt;/p&gt;&lt;/td&gt;
  &lt;/tr&gt;
  &lt;tr&gt;
   &lt;td&gt;&lt;p&gt;Harness engineering&lt;/p&gt;&lt;/td&gt;
   &lt;td&gt;&lt;p&gt;Execution stability and error recovery&lt;/p&gt;&lt;/td&gt;
   &lt;td&gt;&lt;p&gt;Preventing drift, rot, and cascading failure&lt;/p&gt;&lt;/td&gt;
  &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h2&gt;Context rot&lt;/h2&gt;
&lt;p&gt;One failure mode harness engineering exists to solve is context rot. As an agent works through a long task, calling APIs, generating code, and accumulating reasoning steps, the context window fills with an expanding volume of tokens. As input length grows, attention degrades and output gets less reliable.&lt;/p&gt;
&lt;p&gt;The effect is measurable and it depends on position. &lt;a href="https://www.morphllm.com/context-rot"&gt;Accuracy drops from roughly 75% to 55%&lt;/a&gt; when a critical fact sits in the middle of a bloated context window rather than near the start. In a production deployment this compounds with every turn, eventually stalling the agent before it finishes the job.&lt;/p&gt;
&lt;p&gt;The fix is structural, not a better prompt. Rather than replaying the entire conversation on every turn, a well-built harness maintains state deliberately. &lt;a href="https://arxiv.org/pdf/2607.22711"&gt;Trajectory architectures that decouple file reads from their observations&lt;/a&gt;, keeping a synchronized registry of relevant files and injecting only current contents at each cycle, produce measurably lighter trajectories: a 9% to 50% reduction in average input tokens per task and up to 37% fewer reasoning cycles, at comparable pass rates.&lt;/p&gt;
&lt;p&gt;Data format matters too. &lt;a href="https://www.researchgate.net/publication/405160079_Do_LLM_Agents_Need_Raw_JSON_A_Systematic_Study_of_Compact_Encodings_for_Structured_Inputs_1_st"&gt;Compact encodings such as CSV or Markdown lists instead of JSON&lt;/a&gt; can cut token usage by more than 56% while improving extraction accuracy by up to 5 percentage points.&lt;/p&gt;
&lt;h2&gt;A taxonomy for production agents&lt;/h2&gt;
&lt;p&gt;The field has converged on a formal taxonomy for harness design. &lt;a href="https://openreview.net/pdf/ee04ea33117ee379b5d0d6517684c6fc1f190c47.pdf"&gt;The ETCLOVG framework&lt;/a&gt; describes seven layers that a production-ready agent needs beneath it.&lt;/p&gt;
&lt;p&gt;Execution defines where agent code runs and what sandbox bounds it, keeping agents in read-only contexts until outputs pass deterministic quality gates. Tooling specifies the registry of available actions with strict schemas that limit the action space. Context management curates information flow to prevent rot. Lifecycle and orchestration handle continuity across sessions, which matters because &lt;a href="https://www.preprints.org/manuscript/202604.0428"&gt;state has to survive a pause&lt;/a&gt;: when an agent hibernates and resumes, the harness must reconstruct what it was doing without a full human briefing. Observability provides diagnostic output such as token consumption and friction events, so engineers can debug a failure without reconstructing the session from raw logs. Verification checks work before it lands. Governance sets the policies all of it runs under.&lt;/p&gt;
&lt;p&gt;The practical takeaway is that an out-of-the-box model is not sufficient for complex operations, and the gap is not closed by prompting harder. It is closed by the infrastructure around the model: how state is managed, how tools are bounded, how failures surface, and how output gets verified before it counts. &lt;a href="https://codesoapbox.dev/hallucinating-progress-the-risks-of-uncritical-llm-adoption-in-software-development/"&gt;Skipping that layer is how plausible-looking output reaches production unchecked.&lt;/a&gt;&lt;/p&gt;  
&lt;img src="https://track-na2.hubspot.com/__ptq.gif?a=246918688&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fsquidtrain.com%2Fblog%2Fharness-engineering-and-the-pilot-production-gap&amp;amp;bu=https%253A%252F%252Fsquidtrain.com%252Fblog&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <pubDate>Sat, 23 May 2026 16:00:00 GMT</pubDate>
      <author>hello@squidtrain.com (SquidTrain)</author>
      <guid>https://squidtrain.com/blog/harness-engineering-and-the-pilot-production-gap</guid>
      <dc:date>2026-05-23T16:00:00Z</dc:date>
    </item>
    <item>
      <title>Why most AI pilots stall before production</title>
      <link>https://squidtrain.com/blog/why-most-ai-pilots-stall</link>
      <description>&lt;div class="hs-featured-image-wrapper"&gt; 
 &lt;a href="https://squidtrain.com/blog/why-most-ai-pilots-stall" title="" class="hs-featured-image-link"&gt; &lt;img src="https://squidtrain.com/hubfs/Website/images/blog-why-most-ai-pilots-stall-team.jpg" alt="A business team at a table with charts and an AI agents diagram overlaid" class="hs-featured-image" style="width:auto !important; max-width:50%; float:left; margin:0 15px 15px 0;"&gt; &lt;/a&gt; 
&lt;/div&gt; 
&lt;p&gt;A pilot that works in a notebook and a system that works on a Tuesday afternoon are different engineering problems. The gap between them is where most internal AI projects stop. The pilot is not wrong, and the people who built it are not careless. The two things are simply built against different assumptions, and almost none of the assumptions that hold in a demo survive contact with production traffic.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;A pilot that works in a notebook and a system that works on a Tuesday afternoon are different engineering problems. The gap between them is where most internal AI projects stop. The pilot is not wrong, and the people who built it are not careless. The two things are simply built against different assumptions, and almost none of the assumptions that hold in a demo survive contact with production traffic.&lt;/p&gt;
&lt;p&gt;What follows are the five places projects most reliably stall, and what to do instead.&lt;/p&gt;
&lt;h2&gt;The demo runs on curated inputs&lt;/h2&gt;
&lt;p&gt;Pilots are usually built against a handful of examples someone picked by hand. Those examples are clean, representative, and quietly filtered: the ambiguous cases never made it into the set. Nobody did this deliberately. When you are assembling test data, the record that takes twenty minutes to interpret is the one you skip, and the skipped records are exactly the ones that define production behavior.&lt;/p&gt;
&lt;p&gt;Production traffic is the opposite. It contains malformed records, edge cases nobody documented, and inputs that are technically valid but semantically strange. A field that is nominally a date arrives as a free-text string in four formats. A document that should be three pages is a hundred and forty, because someone attached the entire contract. An identifier that is unique in the specification turns out not to be.&lt;/p&gt;
&lt;p&gt;The failure is not that the model gets worse. It is that the input distribution was never the one you tested against. This distinction matters because the two problems have different fixes. A weaker model is solved by a better model. A distribution mismatch is solved by looking at real traffic, and no amount of model upgrading substitutes for that.&lt;/p&gt;
&lt;p&gt;Before automating a task, pull a sample of real inputs. Not a curated sample. Take a window of actual traffic and read it. The point is to find the shapes you did not know existed while the cost of finding them is still low.&lt;/p&gt;
&lt;h2&gt;Nobody owns the failure path&lt;/h2&gt;
&lt;p&gt;In a pilot, when the output is wrong, a person notices and tries again. That person is load-bearing infrastructure, and they do not scale. They are also invisible in every demo, because their intervention happens between the run and the screenshot.&lt;/p&gt;
&lt;p&gt;Production needs an answer to what happens when the model returns something unusable. There are four options and you have to pick one for each failure mode: retry with modified input, fall back to a deterministic path, escalate to a human queue, or fail loudly and stop. Deciding this late is expensive, because the surrounding system was built assuming success.&lt;/p&gt;
&lt;p&gt;The harder question is detection. Retrying is easy once you know the output is bad. Knowing requires a check that runs on every response: schema validation if the output is structured, a constraint on length or format, a rule that catches the specific ways this task fails. An unusable output that passes silently into a downstream system is substantially worse than a crash, because the crash gets attention and the bad record does not.&lt;/p&gt;
&lt;p&gt;Escalation deserves particular scrutiny. A human review queue is a reasonable answer only if someone actually owns the queue and it has a defined maximum depth. An unowned queue is a way of deferring failure, not handling it.&lt;/p&gt;
&lt;h2&gt;Evaluation is an afterthought&lt;/h2&gt;
&lt;p&gt;Teams that ship tend to build the evaluation harness before the feature. Not a benchmark score, but a set of cases drawn from real traffic with known-correct outputs, run on every change.&lt;/p&gt;
&lt;p&gt;Without it, you cannot tell whether a prompt edit helped. You change the wording, spot-check three outputs, and they look better. Whether they are better across the distribution is unknown, and whether the change broke a case that used to work is also unknown. Teams in this position stop editing prompts, because every change is a gamble against invisible regressions.&lt;/p&gt;
&lt;p&gt;A workable evaluation set does not need to be large. Thirty to fifty cases covering the common paths and the known failure modes will catch most regressions. What matters is that the cases come from real traffic and that the expected outputs were verified by someone who understands the task. Synthetic cases test whether the system handles inputs you imagined, which is a weaker guarantee than it sounds.&lt;/p&gt;
&lt;p&gt;Add every production failure to the set as it is found. This is the mechanism by which the harness becomes valuable over time: it accumulates precisely the cases that broke, so they cannot break the same way twice.&lt;/p&gt;
&lt;p&gt;This is unglamorous work and it is the difference between a system you can modify and one you are afraid to touch.&lt;/p&gt;
&lt;h2&gt;The integration surface is larger than expected&lt;/h2&gt;
&lt;p&gt;The model call is a small part of the work. Authentication, rate limits, retries with backoff, cost controls, logging, and access to the systems the output has to touch account for most of the engineering.&lt;/p&gt;
&lt;p&gt;Pilots skip all of it because a single developer running locally already has the permissions. Their credentials are in the environment, they are the only caller so rate limits never trigger, cost is invisible at demo volume, and if something fails they see the traceback. Every one of those becomes a work item when the system runs unattended.&lt;/p&gt;
&lt;p&gt;Rate limits deserve early attention because they fail in a misleading way. Under demo load you never see them. Under production load you see them intermittently, which reads as flakiness rather than a limit, and the instinct is to add a retry. A retry without backoff makes it worse.&lt;/p&gt;
&lt;p&gt;Cost is the other item that only appears at scale. A per-call cost that is negligible in testing becomes the line item somebody asks about. Track spend per run from the first day, because retrofitting cost attribution after the fact means correlating logs that were never designed to be correlated.&lt;/p&gt;
&lt;h2&gt;The task was never defined precisely&lt;/h2&gt;
&lt;p&gt;This one is upstream of the rest. A pilot can succeed against a vague brief because a human is interpreting the output and filling gaps. Production cannot, because there is no interpreter.&lt;/p&gt;
&lt;p&gt;If two people on the team would grade the same output differently, the task is underspecified. That disagreement will surface eventually, usually as a dispute about whether the system is working. Resolve it early by writing down what a correct output looks like, including the cases where the honest answer is that the input does not permit one.&lt;/p&gt;
&lt;p&gt;Tasks where the correct answer is genuinely contested are not good first candidates for automation, however appealing the demo looks.&lt;/p&gt;
&lt;h2&gt;What to do differently&lt;/h2&gt;
&lt;p&gt;Start from a narrow task that a specific person does repeatedly, and instrument it before you automate it. Watching how the task is actually performed usually reveals that the real work is not the step you planned to automate.&lt;/p&gt;
&lt;p&gt;Build the evaluation set from real cases, including the ones that went wrong. Decide the failure path first, for each way the task can fail. Then write the prompt.&lt;/p&gt;
&lt;p&gt;The order matters. A pilot built this way is slower to demo and considerably faster to ship, because the work that usually happens between demo and production has already been done.&lt;/p&gt;  
&lt;img src="https://track-na2.hubspot.com/__ptq.gif?a=246918688&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fsquidtrain.com%2Fblog%2Fwhy-most-ai-pilots-stall&amp;amp;bu=https%253A%252F%252Fsquidtrain.com%252Fblog&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <pubDate>Tue, 12 May 2026 16:00:00 GMT</pubDate>
      <author>hello@squidtrain.com (SquidTrain)</author>
      <guid>https://squidtrain.com/blog/why-most-ai-pilots-stall</guid>
      <dc:date>2026-05-12T16:00:00Z</dc:date>
    </item>
    <item>
      <title>Where the human sits: oversight and pricing</title>
      <link>https://squidtrain.com/blog/where-the-human-sits</link>
      <description>&lt;h2&gt;Per-seat pricing stopped describing the value&lt;/h2&gt;
&lt;p&gt;For twenty years B2B software ran on a simple equation: growth meant more employees, more employees meant more seats, and seats meant predictable recurring revenue. &lt;a href="https://medium.com/@topper440/the-future-of-saas-db14cc7e5ab9"&gt;Agentic AI breaks the correlation.&lt;/a&gt; An autonomous agent does not occupy a seat. When one person with an agent does work that used to take five, headcount stops tracking the value delivered and per-seat licensing stops capturing it.&lt;/p&gt;</description>
      <content:encoded>&lt;h2&gt;Per-seat pricing stopped describing the value&lt;/h2&gt;
&lt;p&gt;For twenty years B2B software ran on a simple equation: growth meant more employees, more employees meant more seats, and seats meant predictable recurring revenue. &lt;a href="https://medium.com/@topper440/the-future-of-saas-db14cc7e5ab9"&gt;Agentic AI breaks the correlation.&lt;/a&gt; An autonomous agent does not occupy a seat. When one person with an agent does work that used to take five, headcount stops tracking the value delivered and per-seat licensing stops capturing it.&lt;/p&gt;
&lt;p&gt;The market moved accordingly. &lt;a href="https://thesaaslibrary.com/b2b-saas-trends-in-2026whats-actually-changing-and-what-isnt/"&gt;Pure per-seat pricing fell from 21% to 15% of the market in twelve months.&lt;/a&gt; Outcome-based and value-based models took the share, and the shift shows up in how deals get discussed: &lt;a href="https://premium.f1gmat.com/consulting/trends/2026/Q1"&gt;73% of clients now prefer value-based or outcome-driven pricing, and 58% raise pricing in the first discovery call.&lt;/a&gt; Federal procurement formalized it in Q1 2026, when three of ten named consulting firms offered performance-based fees as an option to retain contracts, the first time outcome pricing appeared inside a federal consulting procurement.&lt;/p&gt;
&lt;p&gt;What replaces the seat varies by domain. &lt;a href="https://www.mindstudio.ai/blog/professional-services"&gt;Fees get tied to measurable results&lt;/a&gt;: a percentage of tax savings identified, deals closed, tickets resolved. The common thread is that revenue now depends on the software working in production rather than on access being granted. &lt;a href="https://imergeadvisors.com/dealmaker-insights/the-ai-centric-imperative-navigating-saas-disruption"&gt;That is a genuine change in incentive, and it is uncomfortable for vendors whose product demos better than it deploys.&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;Why adoption is a people problem before it is a technical one&lt;/h2&gt;
&lt;p&gt;Measurable outcomes require that people actually use the thing, and this is where most deployments stall. Rolling out AI is not like rolling out enterprise software, because the old playbook assumed determinism. &lt;a href="https://www.tommasomariaricci.com/blog/ai-change-management-framework"&gt;An ERP transaction returns the same output for the same input every time. A model returns probabilistic output that shifts with retraining, data drift, and edge cases.&lt;/a&gt; Distrust of a system that behaves differently on Tuesday than it did on Monday is not irrational resistance. It is a correct read of the system.&lt;/p&gt;
&lt;p&gt;There is a second layer beneath that. Because these systems replicate judgment rather than automating transactions, they land on the part of the job people built a career on. A financial analyst resisting a tool may be resisting the implied message that twelve years of accumulated judgment no longer matters. Treating that as a training gap misdiagnoses it entirely.&lt;/p&gt;
&lt;p&gt;Frameworks that work start from the individual rather than the rollout. &lt;a href="https://www.prosci.com/blog/the-prosci-ai-integration-framework"&gt;The Prosci AI Integration Framework&lt;/a&gt; has people sort their own work into three buckets, which is a small move with a large effect on how the change lands:&lt;/p&gt;
&lt;ul&gt;
 &lt;li&gt;&lt;strong&gt;My work.&lt;/strong&gt; Tasks that stay human because they depend on emotional intelligence, ethical judgment, real-time improvisation, and personal connection. Naming these first establishes what is not up for automation.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;With me.&lt;/strong&gt; Tasks where the model is a collaborator. The person sets direction and the agent helps draft, research, analyze, or generate options. Higher quality output, less cognitive load, human still driving.&lt;/li&gt;
 &lt;li&gt;&lt;strong&gt;For me.&lt;/strong&gt; Routine rules-based work, such as standardized reports and data organization, that can be delegated outright to free up capacity.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The sorting matters more than the categories. When someone maps their own week, the abstract threat of replacement becomes a concrete and much smaller list of tasks they were mostly glad to hand off.&lt;/p&gt;
&lt;h2&gt;The same pattern outside the enterprise&lt;/h2&gt;
&lt;p&gt;Individual adoption has run ahead of institutional adoption. &lt;a href="https://www.hepi.ac.uk/reports/student-generative-ai-survey-2025/"&gt;By 2025, 92% of surveyed students reported using an AI tool, up from 66% the year before&lt;/a&gt;, and 88% had used one for assessments, up from 53%. ChatGPT led at 66% usage. That is among the fastest adoption curves any education technology has recorded.&lt;/p&gt;
&lt;p&gt;The measured outcomes are strong. &lt;a href="https://aibusinessweekly.net/p/ai-education-statistics"&gt;Students using personalized AI tutoring scored 54% higher on tests&lt;/a&gt;, and &lt;a href="https://www.engageli.com/blog/ai-in-education-statistics"&gt;course completion rates ran 70% better than traditional approaches.&lt;/a&gt; A 2025 Harvard physics study found students working with AI tutors learned more than twice as much, in less time, than those in a traditional active-learning classroom.&lt;/p&gt;
&lt;p&gt;And confidence still lags the usage, in the same shape it does inside companies. &lt;a href="https://www.digitaleducationcouncil.com/post/ai-adoption-is-nearly-universal-among-students-but-confidence-is-not"&gt;65% of students worry that leaning on AI makes their learning shallow and discourages critical thinking, and 56% are concerned about data privacy.&lt;/a&gt; Only 36% received any AI skills training from their institution. High usage with low confidence and no instruction is not a success state. It is a gap.&lt;/p&gt;
&lt;h2&gt;Deciding where the human sits&lt;/h2&gt;
&lt;p&gt;Both settings point at the same design question, which is not whether to keep a human involved but exactly where. &lt;a href="https://viston.tech/the-human-in-the-loop-ai-agents-guide-every-business-needs-in-2026/"&gt;Human-in-the-loop design means a qualified person, with adequate context and real authority to act, positioned at the points in a workflow where review changes the outcome.&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Placement should be calibrated to consequence rather than applied uniformly. High-consequence actions, such as financial transactions, access changes, external communications, and infrastructure modification, warrant explicit approval before execution. Confidence-threshold escalation routes work for review when the model’s own reliability score falls below a calibrated line. Audit-only checkpoints suit high-volume low-risk operations, keeping a log without gating throughput.&lt;/p&gt;
&lt;p&gt;Uniform oversight fails in a predictable way. Ask a person to approve everything and they approve everything, which is automation complacency wearing a governance badge. The failure modes worth designing against are that, unclear ownership when something goes wrong, and guardrails brittle enough that people route around them.&lt;/p&gt;
&lt;p&gt;The infrastructure and the economics both now point the same direction. Outcome-based pricing means the vendor is paid when the work lands, and work lands when the people around the system trust it enough to use it and understand it well enough to catch it when it is wrong. Those are the same problem.&lt;/p&gt;  
&lt;img src="https://track-na2.hubspot.com/__ptq.gif?a=246918688&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fsquidtrain.com%2Fblog%2Fwhere-the-human-sits&amp;amp;bu=https%253A%252F%252Fsquidtrain.com%252Fblog&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <pubDate>Wed, 29 Apr 2026 16:00:00 GMT</pubDate>
      <author>hello@squidtrain.com (SquidTrain)</author>
      <guid>https://squidtrain.com/blog/where-the-human-sits</guid>
      <dc:date>2026-04-29T16:00:00Z</dc:date>
    </item>
    <item>
      <title>When to use an agent and when to use a script</title>
      <link>https://squidtrain.com/blog/agent-or-script</link>
      <description>&lt;div class="hs-featured-image-wrapper"&gt; 
 &lt;a href="https://squidtrain.com/blog/agent-or-script" title="" class="hs-featured-image-link"&gt; &lt;img src="https://squidtrain.com/hubfs/Website/images/blog-agent-or-script-ai-agent.jpg" alt="Hands on a laptop with an AI agent overlay showing tools, memory, human control and autonomous actions" class="hs-featured-image" style="width:auto !important; max-width:50%; float:left; margin:0 15px 15px 0;"&gt; &lt;/a&gt; 
&lt;/div&gt; 
&lt;p&gt;Agents are useful when the sequence of steps cannot be known in advance. When it can, an agent is a slower, more expensive, less predictable way to run a script.&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;Agents are useful when the sequence of steps cannot be known in advance. When it can, an agent is a slower, more expensive, less predictable way to run a script.&lt;/p&gt;
&lt;p&gt;This sounds obvious stated plainly. It is worth stating because the current default runs the other way: teams reach for an agent first and discover the determinism problem in production, at which point the agent is load-bearing and hard to remove.&lt;/p&gt;
&lt;h2&gt;The distinction that matters&lt;/h2&gt;
&lt;p&gt;A script encodes a decision you already made. An agent makes the decision at runtime. That flexibility costs latency, tokens, and determinism, and it is worth paying only when the decision genuinely cannot be made ahead of time.&lt;/p&gt;
&lt;p&gt;A useful test: can you draw the flowchart? If you can, write the flowchart. Reaching for an agent to execute a known sequence adds a nondeterministic layer between you and an outcome you already understand.&lt;/p&gt;
&lt;p&gt;The test has a failure mode worth naming. Sometimes you can draw the flowchart but it has forty branches, and writing it feels worse than delegating to an agent. That instinct is usually wrong. Forty branches you can enumerate are forty branches you can test. An agent replaces them with a decision process you cannot enumerate, which is a different thing from simplifying them.&lt;/p&gt;
&lt;p&gt;The genuine case for an agent is not that the branches are numerous. It is that you cannot list them.&lt;/p&gt;
&lt;h2&gt;The middle ground&lt;/h2&gt;
&lt;p&gt;Between a script and an agent sits an option that gets skipped: a single constrained model call inside otherwise deterministic code.&lt;/p&gt;
&lt;p&gt;Classification, extraction, summarization, and format conversion are model-shaped problems with script-shaped control flow. The model handles the part that resists rules, and the surrounding code does everything else. One call, no loop, no tool selection, no planning step.&lt;/p&gt;
&lt;p&gt;This covers a large share of what gets built as an agent. If your agent makes one model call and then stops, it was never an agent. It is a function with an unnecessary framework around it. Removing the framework makes it faster, cheaper, and easier to test.&lt;/p&gt;
&lt;p&gt;Reach for the loop only when the second call depends on what the first one found.&lt;/p&gt;
&lt;h2&gt;Where agents earn their cost&lt;/h2&gt;
&lt;p&gt;Agents fit when the input space is open and the path depends on what is found along the way.&lt;/p&gt;
&lt;p&gt;Investigating a failure, where each finding determines the next thing to check. You cannot pre-specify the sequence because the second step depends on what the first one turned up.&lt;/p&gt;
&lt;p&gt;Research across sources whose structure is unknown until opened. A script can fetch a known endpoint; it cannot decide that a document is irrelevant and try a different source.&lt;/p&gt;
&lt;p&gt;Tasks where the number of steps varies by an order of magnitude between runs. If one input needs two operations and another needs sixty, fixed control flow either wastes work or truncates it.&lt;/p&gt;
&lt;p&gt;The common thread is that enumerating the branches in advance would be impractical, not merely tedious.&lt;/p&gt;
&lt;h2&gt;Where scripts win&lt;/h2&gt;
&lt;p&gt;Fixed transformations, scheduled jobs, and anything whose correctness you need to guarantee. If the same input must always produce the same output, an agent is the wrong tool. Determinism is a feature, and agents trade it away by design.&lt;/p&gt;
&lt;p&gt;Cost compounds too. A script that runs a thousand times a day costs the same each time. An agent doing the same work re-derives its plan on every run, paying for reasoning that produced an identical conclusion the previous nine hundred and ninety-nine times.&lt;/p&gt;
&lt;p&gt;Latency follows the same pattern. A script’s runtime is roughly fixed. An agent’s varies with how many steps it decides to take, which makes it awkward to put in a request path where someone is waiting.&lt;/p&gt;
&lt;h2&gt;They fail differently&lt;/h2&gt;
&lt;p&gt;This is the operational difference and it gets underweighted.&lt;/p&gt;
&lt;p&gt;A script fails at a known line, with a stack trace, in a way that reproduces. You fix it once and it stays fixed.&lt;/p&gt;
&lt;p&gt;An agent fails by making a defensible-looking wrong decision partway through a sequence. There is no stack trace, because nothing threw. Reproducing it may not be possible, because the same input can produce a different path. And the fix is a change to instructions or tools whose effect on other cases is not obvious.&lt;/p&gt;
&lt;p&gt;This means agents need observability that scripts do not: a record of each step, the reasoning, the tool calls, and the results. Without it, debugging is guesswork. Build that before the agent handles anything that matters, because retrofitting it means reproducing failures you cannot reproduce.&lt;/p&gt;
&lt;h2&gt;The hybrid case&lt;/h2&gt;
&lt;p&gt;Most real systems want both. An agent decides what needs to happen and then calls deterministic tools to do it. The agent handles the open-ended part; the scripts handle the parts where you need a guarantee.&lt;/p&gt;
&lt;p&gt;The design question is where to draw the line, and there is a reliable heuristic: anything worth being certain about belongs in a tool the agent calls, not in the agent’s reasoning.&lt;/p&gt;
&lt;p&gt;If a calculation must be right, do not ask the model to perform it. Give it a function. If an operation is destructive, do not let the agent construct the command. Give it a tool with constrained parameters. If a sequence must always run in order, make it one tool rather than three the agent could call out of order.&lt;/p&gt;
&lt;p&gt;Each of those moves work out of the nondeterministic layer without giving up flexibility, because the agent still decides whether and when to call them.&lt;/p&gt;
&lt;h2&gt;Migrating between them&lt;/h2&gt;
&lt;p&gt;The choice is not permanent, and the useful direction is agent to script.&lt;/p&gt;
&lt;p&gt;Start with an agent when the problem is genuinely open. Log the paths it takes. If it turns out that ninety percent of runs follow the same three sequences, those sequences are now the flowchart you could not draw at the start. Encode them as scripts and let the agent handle the remainder.&lt;/p&gt;
&lt;p&gt;This works because the agent doubles as a discovery mechanism. What it cannot do is stay the permanent implementation of a path you now understand.&lt;/p&gt;
&lt;p&gt;The shape most durable systems converge on is a thin layer of judgment over a thick layer of things that behave the same way every time, and the layer of judgment gets thinner as you learn what the system actually does.&lt;/p&gt;  
&lt;img src="https://track-na2.hubspot.com/__ptq.gif?a=246918688&amp;amp;k=14&amp;amp;r=https%3A%2F%2Fsquidtrain.com%2Fblog%2Fagent-or-script&amp;amp;bu=https%253A%252F%252Fsquidtrain.com%252Fblog&amp;amp;bvt=rss" alt="" width="1" height="1" style="min-height:1px!important;width:1px!important;border-width:0!important;margin-top:0!important;margin-bottom:0!important;margin-right:0!important;margin-left:0!important;padding-top:0!important;padding-bottom:0!important;padding-right:0!important;padding-left:0!important; "&gt;</content:encoded>
      <category>script</category>
      <category>agent</category>
      <pubDate>Mon, 06 Apr 2026 16:00:00 GMT</pubDate>
      <author>logan@squidtrain.com (Logan Matson)</author>
      <guid>https://squidtrain.com/blog/agent-or-script</guid>
      <dc:date>2026-04-06T16:00:00Z</dc:date>
    </item>
  </channel>
</rss>
