Skip to content
AI systemsFree

Build a working language model from scratch.

Build a Tiny LLM: From Tokens to Text

Implement the tokenizer, embeddings, attention, transformer block, decoding, training, and interpretability in NumPy, then run StoryByte in your browser.

What you will be able to do

Leave with capability, not just vocabulary.

Build and inspect a tokenizer

Implement embeddings and causal self-attention

Assemble a transformer block and sampler

Connect training loss and internal activations to generated text

Running example

StoryByte, a real roughly one-million-parameter GPT trained on TinyStories and shipped with the course.

Prerequisites

Basic Python and comfort with arrays. No PyTorch or GPU is required for the browser course.

Curriculum

Every module earns the next one.

Open any module to review its exact sections. Progress and completion follow you through the course.

8 modules · 6h 45m
01
Module 1

What an LLM Really Is, and How Text Becomes Numbers

BeginnerFree

Topics include autocomplete loop, tokens, byte-pair encoding, and more.

View 5 sections
  1. 1Autocomplete, Scaled, The Whole Idea
  2. 2Why a Model Can't Read Letters, Tokens
  3. 3Build BPE: Most-Frequent-Pair → Merge → Repeat
  4. 4Byte-Level BPE, 256 Bytes, Nothing Unknown
  5. 5Encode, Decode, and the Strawberry Problem
45 min5 sections
Open module
02
Module 2

From Tokens to Meaning: Embeddings & Position

BeginnerFree

Topics include embedding table, meaning as direction, positional encoding, and more.

View 5 sections
  1. 1The Embedding Table, A Lookup, Not a Computation
  2. 2Directions Carry Meaning
  3. 3Why Attention Is Order-Blind (The Shuffle Test)
  4. 4Injecting Position, Learned, Sinusoidal, RoPE
  5. 5The Residual Stream, The Shared Conveyor Belt
45 min5 sections
Open module
03
Module 3

Attention: How Words Read Each Other

IntermediateFree

Topics include query/key/value, scaled dot-product, the √dₖ scaling, and more.

View 5 sections
  1. 1The Intuition, Query, Key, Value
  2. 2Scaled Dot-Product Attention
  3. 3Why Divide by √dₖ, The Volume Knob
  4. 4Causal Masking, Blindfold the Future
  5. 5Multi-Head Attention, A Committee of Readers
55 min5 sections
Open module
04
Module 4

The Transformer Block: Stacking the Machine

IntermediateFree

Topics include the block, feed-forward network, GELU, and more.

View 5 sections
  1. 1Anatomy of One Block, Two Stations
  2. 2The Feed-Forward Network, Per-Word Thinking
  3. 3Residuals + LayerNorm, and Why Pre-LN
  4. 4Stacking N Blocks, Understanding in Layers
  5. 5Final LayerNorm + Unembedding → Logits
55 min5 sections
Open module
05
Module 5

Making It Speak: From Logits to Words

IntermediateFree

Topics include softmax, greedy, temperature, and more.

View 5 sections
  1. 1Logits → Softmax → Probabilities
  2. 2Greedy Decoding (and Why It Repeats)
  3. 3Temperature, The Creativity Dial
  4. 4Top-k vs Top-p, A Smarter Shortlist
  5. 5The Decode Loop, Generate StoryByte's First Story
50 min5 sections
Open module
06
Module 6

How It Learned: Training a Tiny LLM

AdvancedFree

Topics include next-token prediction, cross-entropy, AdamW, and more.

View 5 sections
  1. 1StoryByte's Training Game: Predict the Next Token
  2. 2Cross-Entropy = Surprise (and Perplexity)
  3. 3The Training Loop, Batch, Loss, Update
  4. 4Why Scale Helps, and How to Spend It
  5. 5The Twist: Data Beats Size, How We Trained StoryByte
55 min5 sections
Open module
07
Module 7

Opening the Box: What's Actually Inside

AdvancedFree

Topics include logit lens, attention-head patterns, induction-head candidates, and more.

View 5 sections
  1. 1The Logit Lens, Reading a Hunch Become an Answer
  2. 2Attention-Head Roles → Induction Heads
  3. 3Features vs Neurons, Superposition
  4. 4The FFN as Key-Value Memory
  5. 5Honest Narration, Evidence, Never Mind-Reading
55 min5 sections
Open module
08
Module 8

The Whole Machine & The Frontier

AdvancedFree

Topics include end-to-end pipeline, parameter budget, documented Llama 2 components, and more.

View 5 sections
  1. 1The Full Pipeline, End to End
  2. 2Where the Parameters Live
  3. 3How Modern LLMs Differ from StoryByte
  4. 4The Honest Scorecard, What It Can and Cannot Do
  5. 5Where to Go Next
45 min5 sections
Open module
Who this course is for

Built for people who need to use the skill.

01

Engineers who want transformer internals to click

02

Python learners ready for a serious build

03

AI practitioners tired of architecture diagrams alone

Start the course

Begin with What an LLM Really Is, and How Text Becomes Numbers.

The first module establishes the language and example used throughout the rest of the course.

Open Module 1
Build a Tiny LLM: From Tokens to Text | Let's Data Science