Skip to content
AI systemsFree

Understand what happens between tokens and text.

LLM Foundations

Build a connected mental model of tokenization, embeddings, attention, transformer blocks, training, decoding, post-training, and interpretability.

What you will be able to do

Leave with capability, not just vocabulary.

Explain how text becomes tokens and vectors

Trace information through attention and transformer blocks

Connect training objectives to model behavior

Reason about decoding, post-training, and interpretability

Running example

Small token sequences and model traces that make each internal operation visible before scaling the idea up.

Prerequisites

No prior LLM course is required. Basic algebra and Python familiarity help with optional code.

Curriculum

Every module earns the next one.

Open any module to review its exact sections. Progress and completion follow you through the course.

8 modules · ~8 hours
01
Module 1

Tokens: The Words a Model Sees

BeginnerFree

Topics include BPE, Vocabulary, Glitch tokens, and more.

View 5 sections
  1. 1Why Tokens Exist
  2. 2Byte-Pair Encoding (BPE): The Algorithm
  3. 3Vocabulary, Special Tokens, and the Pre-Tokenizer
  4. 4Glitch Tokens and Tokenizer Fragility
  5. 5Tokenizer Choice in 2026: Tiktoken vs SentencePiece vs Llama
60 min5 sections
Open module
02
Module 2

Embeddings & Positional Encoding

BeginnerFree

Topics include Word embeddings, Sinusoidal positions, RoPE, and more.

View 5 sections
  1. 1Tokens Become Vectors
  2. 2The Embedding Table: A Lookup, Not a Computation
  3. 3Why Order Matters: The Need for Positional Encoding
  4. 4Sinusoidal Positions: The Original Trick
  5. 5Rotary Position Embedding (RoPE): The Modern Default
60 min5 sections
Open module
03
Module 3

Attention is Information Routing

IntermediateFree

Topics include Q/K/V, Scaled dot-product, Multi-head, and more.

View 5 sections
  1. 1The Intuition: Tokens Reading From Each Other
  2. 2Q, K, V: Three Lenses on the Same Vector
  3. 3The Scaled Dot-Product Attention Formula
  4. 4Multi-Head Attention: Subspace Specialization
  5. 5Causal Masking and the Residual Stream
60 min5 sections
Open module
04
Module 4

The Transformer Block

IntermediateFree

Topics include LayerNorm vs RMSNorm, SwiGLU FFN, Pre-norm, and more.

View 5 sections
  1. 1Anatomy of One Block
  2. 2Normalization: LayerNorm, RMSNorm, and Why Order Matters
  3. 3The Feed-Forward Network: GELU, SwiGLU, and Width
  4. 4Inference-Efficient Attention: MQA, GQA, and MLA
  5. 5Mixture of Experts: Routing for Scale
60 min5 sections
Open module
05
Module 5

Training at Scale

AdvancedFree

Topics include Next-token loss, AdamW, Scaling laws, and more.

View 5 sections
  1. 1Next-Token Prediction and Cross-Entropy Loss
  2. 2Optimization: AdamW, LR Schedules, Gradient Clipping
  3. 3Scaling Laws: Kaplan to Chinchilla and Beyond
  4. 42026 Reality: Deliberate Over-Training for Inference Economics
  5. 5Distributed Training: FSDP / ZeRO and Mixed Precision
60 min5 sections
Open module
06
Module 6

Inference & Decoding

IntermediateFree

Topics include Greedy vs sampling, Temperature, Top-k / top-p / min-p, and more.

View 5 sections
  1. 1From Logits to Tokens: The Inference Loop
  2. 2Sampling Strategies: Temperature, Top-k, Top-p, Min-p
  3. 3The KV Cache: Why Inference is Memory-Bound
  4. 4Speculative Decoding and Other Production Tricks
  5. 5Test-Time Compute: Sequential CoT and Best-of-N
60 min5 sections
Open module
07
Module 7

Post-Training: SFT, DPO, GRPO

AdvancedFree

Topics include SFT, DPO, GRPO / RLVR, and more.

View 5 sections
  1. 1Why Pre-Training Is Not Enough
  2. 2Supervised Fine-Tuning (SFT): Instruction Following
  3. 3Direct Preference Optimization (DPO): RLHF Without the RL
  4. 4Group Relative Policy Optimization (GRPO) and RLVR
  5. 5Reasoning Models: o1, R1, and Emergent Chain-of-Thought
60 min5 sections
Open module
08
Module 8

Looking Inside & The Frontier

AdvancedFree

Topics include Mechanistic interpretability, Induction heads, Sparse autoencoders, and more.

View 5 sections
  1. 1The Residual Stream as Shared Memory
  2. 2Induction Heads: The First Real Circuit
  3. 3Sparse Autoencoders and Monosemantic Features
  4. 4Evaluation: MMLU, HumanEval, lm-eval-harness, Holistic Eval
  5. 5The Frontier: Mamba, World Models, and What Comes Next
60 min5 sections
Open module
Who this course is for

Built for people who need to use the skill.

01

Engineers entering LLM application work

02

Data professionals who need model-level intuition

03

Technical leaders evaluating LLM claims

Start the course

Begin with Tokens: The Words a Model Sees.

The first module establishes the language and example used throughout the rest of the course.

Open Module 1
LLM Foundations, How Large Language Models Actually Work | Free Interactive Course | Let's Data Science