Skip to content
AI EngineeringFree

Understand what happens between tokens and text.

LLM Foundations

Build a connected mental model of tokenization, embeddings, attention, transformer blocks, training, decoding, post-training, and interpretability.

What you will leave with

You can trace an LLM request through the major mechanisms, explain the tradeoffs, and reason about failures without relying on slogans.

Modules
8
Duration
~8 hours
Level
Beginner to Advanced
Access
Free course
  • Interactive Tokenizer & Attention
  • RoPE / DPO / GRPO from scratch
  • Induction Heads + SAE demo

What you will be able to do

What this course prepares you to do.

The curriculum is organized around these 4 practical outcomes.

Explain how text becomes tokens and vectors

Trace information through attention and transformer blocks

Connect training objectives to model behavior

Reason about decoding, post-training, and interpretability

Curriculum

Every module earns the next one.

Open a module to inspect every section before you start. Your progress follows you through the course.

01
Module 1

Tokens: The Words a Model Sees

BeginnerFree

BPE, Vocabulary, Glitch tokens, and more.

View 5 sections
  1. 1Why Tokens Exist
  2. 2Byte-Pair Encoding (BPE): The Algorithm
  3. 3Vocabulary, Special Tokens, and the Pre-Tokenizer
  4. 4Glitch Tokens and Tokenizer Fragility
  5. 5Tokenizer Choice in 2026: Tiktoken vs SentencePiece vs Llama
60 min5 sections
Open module
02
Module 2

Embeddings & Positional Encoding

BeginnerFree

Word embeddings, Sinusoidal positions, RoPE, and more.

View 5 sections
  1. 1Tokens Become Vectors
  2. 2The Embedding Table: A Lookup, Not a Computation
  3. 3Why Order Matters: The Need for Positional Encoding
  4. 4Sinusoidal Positions: The Original Trick
  5. 5Rotary Position Embedding (RoPE): The Modern Default
60 min5 sections
Open module
03
Module 3

Attention is Information Routing

IntermediateFree

Q/K/V, Scaled dot-product, Multi-head, and more.

View 5 sections
  1. 1The Intuition: Tokens Reading From Each Other
  2. 2Q, K, V: Three Lenses on the Same Vector
  3. 3The Scaled Dot-Product Attention Formula
  4. 4Multi-Head Attention: Subspace Specialization
  5. 5Causal Masking and the Residual Stream
60 min5 sections
Open module
04
Module 4

The Transformer Block

IntermediateFree

LayerNorm vs RMSNorm, SwiGLU FFN, Pre-norm, and more.

View 5 sections
  1. 1Anatomy of One Block
  2. 2Normalization: LayerNorm, RMSNorm, and Why Order Matters
  3. 3The Feed-Forward Network: GELU, SwiGLU, and Width
  4. 4Inference-Efficient Attention: MQA, GQA, and MLA
  5. 5Mixture of Experts: Routing for Scale
60 min5 sections
Open module
05
Module 5

Training at Scale

AdvancedFree

Next-token loss, AdamW, Scaling laws, and more.

View 5 sections
  1. 1Next-Token Prediction and Cross-Entropy Loss
  2. 2Optimization: AdamW, LR Schedules, Gradient Clipping
  3. 3Scaling Laws: Kaplan to Chinchilla and Beyond
  4. 42026 Reality: Deliberate Over-Training for Inference Economics
  5. 5Distributed Training: FSDP / ZeRO and Mixed Precision
60 min5 sections
Open module
06
Module 6

Inference & Decoding

IntermediateFree

Greedy vs sampling, Temperature, Top-k / top-p / min-p, and more.

View 5 sections
  1. 1From Logits to Tokens: The Inference Loop
  2. 2Sampling Strategies: Temperature, Top-k, Top-p, Min-p
  3. 3The KV Cache: Why Inference is Memory-Bound
  4. 4Speculative Decoding and Other Production Tricks
  5. 5Test-Time Compute: Sequential CoT and Best-of-N
60 min5 sections
Open module
07
Module 7

Post-Training: SFT, DPO, GRPO

AdvancedFree

SFT, DPO, GRPO / RLVR, and more.

View 5 sections
  1. 1Why Pre-Training Is Not Enough
  2. 2Supervised Fine-Tuning (SFT): Instruction Following
  3. 3Direct Preference Optimization (DPO): RLHF Without the RL
  4. 4Group Relative Policy Optimization (GRPO) and RLVR
  5. 5Reasoning Models: o1, R1, and Emergent Chain-of-Thought
60 min5 sections
Open module
08
Module 8

Looking Inside & The Frontier

AdvancedFree

Mechanistic interpretability, Induction heads, Sparse autoencoders, and more.

View 5 sections
  1. 1The Residual Stream as Shared Memory
  2. 2Induction Heads: The First Real Circuit
  3. 3Sparse Autoencoders and Monosemantic Features
  4. 4Evaluation: MMLU, HumanEval, lm-eval-harness, Holistic Eval
  5. 5The Frontier: Mamba, World Models, and What Comes Next
60 min5 sections
Open module

Who this course is for

Built for people who need to use the skill.

Start with the background you have. The prerequisite notes above tell you exactly what is assumed.

01

Engineers entering LLM application work

02

Data professionals who need model-level intuition

03

Technical leaders evaluating LLM claims

Start the course

Begin with Tokens: The Words a Model Sees.

Module 1 introduces the language and example used throughout the rest of the course.

Open Module 1
LLM Foundations, How Large Language Models Actually Work | Free Interactive Course | Let's Data Science