Breach Protocol: a Transformer You Can Read, Running on the Codes You Type
I wrote a small Transformer in C++, taught it a trick in 85 seconds, compiled the same code to run in your browser, and drew what it thinks in 3D. Here is how, in plain words.
TL;DR. A Transformer is the kind of program inside every chatbot you have used. I built a tiny one, 44,335 numbers instead of billions, that reads a row of codes from the Cyberpunk 2077 hacking minigame and writes the row backwards. It runs in the page, and while it runs, every beam you see is a real number it just computed. The full explainer, with the code, is at designsbyduhart.org/infographics/12-breach-protocol, and the deck on its own is at designsbyduhart.org/demos/breach.
Most explanations of how these models work show you a diagram: boxes and arrows, "attention" written in one of the boxes. A diagram tells you where the pieces go. It cannot tell you what is inside them, and what is inside them is the part everybody gets wrong. So I did not draw a Transformer. I built one small enough to read, and I made it draw itself.
You type codes like 1C 55 7A. The model answers 7A 55 1C. That is a silly job on purpose. It is the smallest job I could find where you can watch the three things a Transformer does, and see whether it is doing them right.
Background
Three pieces set the bar. Brendan Bycroft's LLM Visualization walks through a small GPT in 3D with its real numbers. Polo Club's Transformer Explainer runs GPT-2 in your browser. And Harvard's Annotated Transformer rewrote the original paper as a working program, which is how I learned this and what mine follows line by line.
I wanted one on my own site, with the rule the rest of my demos follow: if it is on the screen, it is running. And I wanted it to look like a netrunner's deck out of Cyberpunk 2077, yellow on black, because the game's Breach Protocol minigame is a grid of codes and a buffer, and that is roughly what a model like this is too.
Scope
What this covers
- What a Transformer actually does with a row of codes, one step at a time, with the code that does it
- How I taught it, and how I checked that the teaching was working
- Why I wrote it once in C++ and compiled it twice, once for my machine and once for your browser
- Where it breaks, with the numbers
What it does not cover
- The chatbot kind of Transformer. Those only have the second half of this one. I built both halves because the second half reading the first is the best part to watch
- How words become tokens. My model knows fifteen codes and that is the whole vocabulary
- Anything about size. This is 44 thousand numbers. The models you have used are a hundred thousand times bigger, and nothing here explains what happens at that scale
The challenge
Two things were hard. Neither was the model.
The first was picking the job. Attention is only worth watching when the right answer has a shape. The original paper's toy job was copying the input, and a copy looks like a straight line: correct, and boring. Sorting is more interesting but the patterns a tiny model learns for it are a mess. Mirroring the row turned out to be perfect. To write the first output the model has to look at the last input, so a correct run draws an X. The half that writes has to be stopped from peeking ahead, and you can see the stop. And the model has to know where each code sits in the row, which is the thing students nod through and never really believe until they see a model fail without it.
The second was resisting a second model. My first version had a trainer in JavaScript and a separate copy of the model in TypeScript for drawing. Two copies have to agree on every detail, so I needed a test to catch them disagreeing. It worked, and it was the wrong shape. Also, the first deck drew everything: three panels, a grid of numbers, ribbons, thumbnails. The person who asked for it called it busy. He was right.
Solutions and process
-
Write the model once, in C++, and compile it twice. One file,
model.cpp, is the model: turning codes into numbers, adding the position clock, attention with its mask, the think step, and picking the next code. Plain arrays of floats, no libraries. Compiled with g++ it is what the trainer drives on my machine. Compiled with Emscripten, the same file becomes a 35 KB WebAssembly module with no dependencies, which the browser downloads and calls through a dozen C functions. One source cannot disagree with itself, so the test that guarded the seam went away with the seam. -
Write the learning by hand, then check it by hand. Training means guessing, measuring how wrong the guess was, and working out for every one of the 44,335 numbers whether nudging it up or down would help. The working out is called the gradient, and
train.cppcomputes it by running the whole calculation backwards. I wrote every step of that myself. Then I checked it the dumb way: nudge each number a tiny bit, measure the change, and compare with what my code claimed. The worst disagreement across every number was 3.3 parts in a million. I did that before trusting a single training run, because a wrong gradient does not crash. It learns slowly and badly and you blame something else. -
Train it the way the paper did, small. Batches of 32 random rows, every row in a batch the same length so there is no padding to handle. Nudges that start small, grow for 400 rounds, then ease off, which is the schedule the original paper used. 4,000 rounds in 85 seconds on one CPU core. I kept the model from round 3,250, the first time it got every held-out row right, and it stayed right at every length from 2 to 8 codes.
-
Ship the numbers with a receipt. The trained numbers go out as a 121 KB file, as a list in the exact order the C++ expects, so the browser lays them end to end and hands them over without knowing what any of them is. The same file carries one fixed input, what the model answered, and the fifteen scores it produced for the first code. A test replays that input through the WebAssembly build and checks the answer and the scores match. Same source on both sides, so this checks the compiler, not my code, and it costs nothing to keep.
-
Draw beats, not frames, in three dimensions. The deck walks one run in stages: turn the codes into numbers, run the reading half, then for every output code the three steps the writing half takes. The reading half appears once because it runs once, which people get wrong and which the deck teaches without a sentence. The scene is three.js: the reading half in front with your codes on the floor and one thin slab per layer, the writing half behind it with the answer filling in, and attention drawn as beams between slots. Beams along a slab are the model looking at its own row. Beams across the gap are the writing half reading the memory.
-
Keep it sparse. The second version of the deck draws no numbers at all. A beam's thickness and brightness are its weight, beams under six percent are not drawn, and each finished output slot leaves its strongest beam behind so the X builds up. Hover a beam and its two numbers appear in one line, and nowhere else. Less on the screen turned out to be more of the model.
-
Explain it like a book. Under the deck there is a table of contents, a start page and fifteen short chapters, one per idea, each with the real lines of C++ that do that idea, cut from the source when the page is built so they cannot go stale. Most chapters have a button that jumps the deck to the moment they describe. I wrote them for someone in high school who has never seen a neural network, because if I cannot explain attention with multiplying and adding, I do not understand it well enough.
What it gets wrong
Every row the model saw in training was 2 to 8 codes long. Type 9 or 10 and it has a position pattern for those slots, because the clock is a formula and not a table, but it was never once asked what to do with them. Usually the first few beams still go to the right place and then the X falls apart in the middle.
I left that in. It is the honest edge of what the model knows, and it shows something the accuracy table cannot: a model that scores 100% on the test can still be confidently wrong one step outside it. The big models have the same edge. It is just further away and harder to find.
Takeaway
- If the point is what is inside the boxes, compute what is inside the boxes. A drawing of attention teaches you the drawing.
- Pick the toy job for the shape of its right answer. Mirroring gives you an X, a visible stop sign, and a reason for position, all on one screen.
- Write a model once and compile it twice. A trainer and a runtime that share a source cannot disagree, and the test that policed the disagreement can go.
- Check gradients by nudging. It is slow, dumb, and the only way to know.
- Show where it breaks. The length where it fails taught me more than the table where it passes.
The bill: every operation written by hand, forwards and backwards. A second compiler to install. And nothing here transfers to the big models except the shapes.
Over to you
If you have ever tried to explain attention to someone, in a classroom or a code review: what did they get wrong first? The mask, the percentages, or that the reading half only runs once? I picked a job to make those three visible, and I would like to know if I picked the right three.
References
- Vaswani et al., Attention Is All You Need (2017)
- Harvard NLP, The Annotated Transformer
- Brendan Bycroft, LLM Visualization
- Polo Club, Transformer Explainer
- The book and the deck: infographic 12 and designsbyduhart.org/demos/breach
Next: the same deck learning to sort, and why looking things up by what they are is harder to draw than looking them up by where they are.
More: linkedin.com/in/jmichaelduhart · instagram.com/designsbyduhart Portfolio and case studies: designsbyduhart.org
#MachineLearning #SoftwareEngineering #SystemDesign #WebAssembly