Infographic 06 · D3
PyTorch only, no HuggingFace, no transformers, no pretrained anything. The interesting part is not the model; it is that evaluation has the authority to stop the release.
The ring is generated from the gate list rather than drawn, adding a tenth gate changes one array. Every segment must be green for the arrow into promotion to exist at all, which is the entire argument of the diagram.
| # | Gate | Result |
|---|---|---|
| 01 | Validation perplexity | 40.0 against a 42.0 ceiling |
| 02 | LAMBADA accuracy | loaded from the original release, not a mirror |
| 03 | HellaSwag accuracy | no HuggingFace datasets anywhere |
| 04 | Contamination | eval text absent from the training shards |
| 05 | Determinism | two runs, same seed, identical logits |
| 06 | KV-cache equivalence | cached and uncached decode agree to tolerance |
| 07 | Latency budget | p95 first-token under the serving target |
| 08 | Memory ceiling | resident set fits the container limit |
| 09 | Artifact integrity | checkpoint SHA256 matches the manifest |
| Stage | What it produces |
|---|---|
| Corpus | 130M tokens · WikiText-103, licence recorded |
| Tokenizer | SentencePiece BPE · 0 UNK on held-out text |
| Training | 8,000 steps · 3.0 h · val PPL 40.0 · CPU |
| Evaluation | nine gates with the authority to block |
| Promotion | checkpoint → serving.pt, SHA256 matched |
| Serving | k3s + Helm + HPA · observed 3 replicas |
The full write-up is Solomon, Evaluation Gates & Promotion.