Every AI model you use at work writes the same way: it looks at the text so far, scores every possible next piece, picks one, and repeats. We wanted people to see that happen, with a real model they can poke at. So we trained one.
On October 4, 2026, SquidTrain published Token Lab, a small language model we trained ourselves on 144 public-domain books and an interactive page that runs it entirely in your browser. You type a sentence and watch the model read it as tokens, score its options for the next one, pick, and keep going. You can change how it chooses, see which earlier words it looked at, and scrub back through snapshots saved during training to the moment it knew nothing.
Try Token Lab. Nothing you type leaves your browser.
What can you do in Token Lab?
The page is built as a set of panels, and each one shows a step the large commercial models also take.
What the model reads. Your text is split into numbered pieces called tokens. The model never sees letters or words, only those numbers. "Sherlock" becomes three pieces. A context bar shows how much of the model's 256-token window your text fills.
What comes next. For the current position, the page lists the most likely next tokens with their odds. Press "Next token" to add one at a time, or "Write 40 tokens" to let it run.
How it chooses. Three controls change the pick. Temperature reshapes the odds: at zero the model always takes the top choice and soon repeats itself, and higher values let unlikely words through. Top-k and top-p cut the list down before the pick. This is why the same prompt can give a different answer every time you ask.
Where it looked. Click any token to see which earlier tokens the model paid attention to when it was there, layer by layer and head by head. Each head learned on its own what to track, and you can compare them.
Watch it learn. We saved real snapshots while the model trained. Drag the slider back to step 0 and the sample text is gibberish. A few hundred steps in, it has words. By the end it has grammar, names and the style of the books it read. A chart shows the error on text the model never saw during training, falling as it learns.
Model size. Switch between three models trained on the same books, from under 1 million to 40 million parameters, and see how each one continues your sentence.
How did we build it?
We built Token Lab with Claude Opus 5.5, which wrote the training code, the browser engine, and the page. The training itself ran on one office PC with an NVIDIA RTX 3090 graphics card.
The books. We used 144 English books from Project Gutenberg, all in the public domain, including Pride and Prejudice, the Sherlock Holmes stories, Frankenstein, Dracula and Moby Dick. That came to about 31 million tokens. We stripped the Gutenberg headers and footers, dropped non-English text, and removed sentences containing racial slurs that are common in books of that era. The page runs the same filter on anything the model writes.
The tokenizer. We trained our own, with a vocabulary of 4,096 tokens. Common words like "the" get a single token. Rarer words are built from pieces, which is why "Sherlock" splits into three.
The models. All three use the same design as the GPT family, at a much smaller scale:
- Tiny: 0.95 million parameters, 2 layers, trained in about 3 minutes.
- Small: 12.3 million parameters, 6 layers, trained in about 16 minutes. This is the one the page loads first.
- Medium: 40.1 million parameters, 12 layers, trained in about 42 minutes.
Each model read the full set of books about 12 times. Training all three took 61 minutes, and the whole run, including downloading the books and building the tokenizer, took 66.
Running it in the browser. The trained weights are compressed to 8 bits and loaded by the page as plain files. A small engine we wrote in JavaScript does the math the model needs, in a background thread so the page stays responsive. We checked it against the original training code: the two agree to within rounding error on every test sentence, so what you see on the page is the real model, step for step.
Why build a small model on purpose?
A small model is fluent and often wrong, and that makes it a better teacher than a large one.
Ask the small model to finish "The capital of France is the city of" and it rates "France" at 24%, "England" at 14%, and "Paris" at 6% (at the page's starting temperature setting). It has read plenty of sentences shaped like that one, so it knows a place name comes next. It has not read enough to know which one is true. It is predicting what usually comes next, and the right answer is only one of the likely options.
The models you use at work run on the same mechanism with vastly more data and parameters, so they get "Paris" right. But they still produce answers by choosing likely pieces, one at a time. When a large model is wrong, it is wrong in the same fluent, confident voice it uses when it is right. Seeing that at a small scale, where the odds are on the screen, makes it much easier to understand when you meet it at a large one.
What does it teach about the AI you already use?
A few lessons we come back to in almost every training session become obvious after ten minutes with Token Lab.
Different answers to the same question are normal. The model picks from a list of likely options, and the temperature setting decides how adventurous that pick is. If you need the same answer twice, you need a process that checks it, since asking again will not always give you the same one.
Confidence is a writing style. The model does not know when it is wrong. A fluent answer tells you the words are likely. Whether they are correct is a separate question. Check facts, figures and names against a source.
Context has a limit. Our model can see 256 tokens. Commercial models can see far more, but every model has a window, and anything outside it does not exist for the model. Long conversations and long documents eventually push the beginning out of view.
Tokens are the unit of everything. Context windows and API prices are counted in tokens, and many usage limits are too. Once you have seen a word split into pieces, those numbers make more sense.
How does this fit what SquidTrain does?
SquidTrain is an AI training and software development company. We train companies, teams and individuals to use AI on the work they already do, and on the rest of life too: planning, writing, learning and helping kids understand the tools they are already using. Builds like Token Lab are part of how we teach. People make better decisions about AI once they have seen how it works, and an interactive model beats a slide about one.
The build also shows the second half of what we do. We built the data pipeline, the model code, a custom engine, and a finished web page with Claude Opus 5.5, and ran it all on hardware already in the office. When training alone will not solve a problem, we build the tool.
If you want your team to understand the tools they are using, or you have a problem that might need a build of its own, see our training or book a free call. And the Lab has more builds to try.
The books in Token Lab come from Project Gutenberg and are in the public domain in the United States. The full list is on the Token Lab page under "The books it read."



