Skip to content
AI EngineeringFree

Build a working language model from scratch.

Build a Tiny LLM: From Tokens to Text

Implement the tokenizer, embeddings, attention, transformer block, decoding, training, and interpretability in NumPy, then run StoryByte in your browser.

What you will leave with

You finish with a working mental and code model of GPT-style generation, from raw text to one predicted token at a time.

Modules
8
Duration
6h 45m
Level
Beginner to Advanced
Access
Free course
  • StoryByte: a real 1M-param GPT
  • Build it from scratch in NumPy
  • Runs live in your browser
  • Real attention maps & loss curve

What you will be able to do

What this course prepares you to do.

The curriculum is organized around these 4 practical outcomes.

Build and inspect a tokenizer

Implement embeddings and causal self-attention

Assemble a transformer block and sampler

Connect training loss and internal activations to generated text

Curriculum

Every module earns the next one.

Open a module to inspect every section before you start. Your progress follows you through the course.

01
Module 1

What an LLM Really Is, and How Text Becomes Numbers

BeginnerFree

autocomplete loop, tokens, byte-pair encoding, and more.

View 5 sections
  1. 1Autocomplete, Scaled, The Whole Idea
  2. 2Why a Model Can't Read Letters, Tokens
  3. 3Build BPE: Most-Frequent-Pair → Merge → Repeat
  4. 4Byte-Level BPE, 256 Bytes, Nothing Unknown
  5. 5Encode, Decode, and the Strawberry Problem
45 min5 sections
Open module
02
Module 2

From Tokens to Meaning: Embeddings & Position

BeginnerFree

embedding table, meaning as direction, positional encoding, and more.

View 5 sections
  1. 1The Embedding Table, A Lookup, Not a Computation
  2. 2Directions Carry Meaning
  3. 3Why Attention Is Order-Blind (The Shuffle Test)
  4. 4Injecting Position, Learned, Sinusoidal, RoPE
  5. 5The Residual Stream, The Shared Conveyor Belt
45 min5 sections
Open module
03
Module 3

Attention: How Words Read Each Other

IntermediateFree

query/key/value, scaled dot-product, the √dₖ scaling, and more.

View 5 sections
  1. 1The Intuition, Query, Key, Value
  2. 2Scaled Dot-Product Attention
  3. 3Why Divide by √dₖ, The Volume Knob
  4. 4Causal Masking, Blindfold the Future
  5. 5Multi-Head Attention, A Committee of Readers
55 min5 sections
Open module
04
Module 4

The Transformer Block: Stacking the Machine

IntermediateFree

the block, feed-forward network, GELU, and more.

View 5 sections
  1. 1Anatomy of One Block, Two Stations
  2. 2The Feed-Forward Network, Per-Word Thinking
  3. 3Residuals + LayerNorm, and Why Pre-LN
  4. 4Stacking N Blocks, Understanding in Layers
  5. 5Final LayerNorm + Unembedding → Logits
55 min5 sections
Open module
05
Module 5

Making It Speak: From Logits to Words

IntermediateFree

softmax, greedy, temperature, and more.

View 5 sections
  1. 1Logits → Softmax → Probabilities
  2. 2Greedy Decoding (and Why It Repeats)
  3. 3Temperature, The Creativity Dial
  4. 4Top-k vs Top-p, A Smarter Shortlist
  5. 5The Decode Loop, Generate StoryByte's First Story
50 min5 sections
Open module
06
Module 6

How It Learned: Training a Tiny LLM

AdvancedFree

next-token prediction, cross-entropy, AdamW, and more.

View 5 sections
  1. 1StoryByte's Training Game: Predict the Next Token
  2. 2Cross-Entropy = Surprise (and Perplexity)
  3. 3The Training Loop, Batch, Loss, Update
  4. 4Why Scale Helps, and How to Spend It
  5. 5The Twist: Data Beats Size, How We Trained StoryByte
55 min5 sections
Open module
07
Module 7

Opening the Box: What's Actually Inside

AdvancedFree

logit lens, attention-head patterns, induction-head candidates, and more.

View 5 sections
  1. 1The Logit Lens, Reading a Hunch Become an Answer
  2. 2Attention-Head Roles → Induction Heads
  3. 3Features vs Neurons, Superposition
  4. 4The FFN as Key-Value Memory
  5. 5Honest Narration, Evidence, Never Mind-Reading
55 min5 sections
Open module
08
Module 8

The Whole Machine & The Frontier

AdvancedFree

end-to-end pipeline, parameter budget, documented Llama 2 components, and more.

View 5 sections
  1. 1The Full Pipeline, End to End
  2. 2Where the Parameters Live
  3. 3How Modern LLMs Differ from StoryByte
  4. 4The Honest Scorecard, What It Can and Cannot Do
  5. 5Where to Go Next
45 min5 sections
Open module

Who this course is for

Built for people who need to use the skill.

Start with the background you have. The prerequisite notes above tell you exactly what is assumed.

01

Engineers who want transformer internals to click

02

Python learners ready for a serious build

03

AI practitioners tired of architecture diagrams alone

Start the course

Begin with What an LLM Really Is, and How Text Becomes Numbers.

Module 1 introduces the language and example used throughout the rest of the course.

Open Module 1
Build a Tiny LLM: From Tokens to Text | Let's Data Science