Skip to content
Statistics & MLPro

Train models the way working data scientists do.

Machine Learning Fundamentals (scikit-learn)

Learn the supervised-learning loop with regression, classification, regularization, metrics, cross-validation, and leak-free scikit-learn pipelines.

What you will be able to do

Leave with capability, not just vocabulary.

Frame a supervised-learning problem correctly

Train and interpret regression and classification models

Select metrics that match the decision

Build cross-validated pipelines without data leakage

Running example

Lendly, a peer-to-peer lending dataset used to predict and evaluate borrower outcomes.

Prerequisites

Comfort with Python and basic statistics. Pandas familiarity is recommended.

Curriculum

Every module earns the next one.

Open any module to review its exact sections. Progress and completion follow you through the course.

8 modules · ~9 hours
01
Module 1

The ML Frame & the Lendly Dataset

BeginnerFree preview

Topics include Supervised vs Unsupervised vs Reinforcement, Features, Targets, and the Hypothesis Class, Regression vs Classification, and more.

View 5 sections
  1. 1What Machine Learning Actually Is
  2. 2Features, Targets, and the Hypothesis Class
  3. 3Regression vs Classification: Same Frame, Different Output
  4. 4Train, Validate, Test: Why You Need Three Sets
  5. 5Meet Lendly: Your Running Dataset for All 8 Modules
65 min5 sections
Open module
02
Module 2

Linear Regression: From Line-Fitting to Loss

BeginnerPro

Topics include The Linear Model y = Xβ + ε, Squared Error and Why, The Normal Equation (Closed Form), and more.

View 6 sections
  1. 1The Linear Model: A Weighted Sum of Features
  2. 2Mean Squared Error: Why Squared, Why Mean
  3. 3The Normal Equation: Closed-Form Solution
  4. 4Gradient Descent: When Closed Form Won't Scale
  5. 5Reading the Output: Coefficients, R², Residuals
  6. 6Lendly Case: Predicting Interest Rate from Borrower Features
75 min6 sections
Open module
03
Module 3

Regularization & the Bias–Variance Tradeoff

IntermediatePro

Topics include Overfitting, The Symptom, The Bias-Variance Decomposition, Ridge (L2), Shrink Coefficients, and more.

View 6 sections
  1. 1Overfitting in Lendly: The Symptom
  2. 2Bias and Variance: The Two Failure Modes
  3. 3Ridge Regression: Squeeze the Coefficients
  4. 4Lasso: Squeeze Some Coefficients to Zero
  5. 5Picking Alpha and Why Standardization Is Non-Negotiable
  6. 6Lendly Case: Regularized Rate Prediction
65 min6 sections
Open module
04
Module 4

Classification: Logistic, kNN, Decision Tree

IntermediatePro

Topics include From Regression to Classification, Logistic Regression & The Sigmoid, k-Nearest Neighbors, The Lazy Learner, and more.

View 6 sections
  1. 1The Classification Setup: Predicting Default on Lendly
  2. 2Logistic Regression: Sigmoid, Log-Odds, and Why
  3. 3k-Nearest Neighbors: Geometric, Lazy, Powerful
  4. 4Decision Trees: Recursive Yes/No Splits
  5. 5Three Models, One Dataset: How They Disagree and Why
  6. 6Lendly Case: Default Prediction Bake-Off
70 min6 sections
Open module
05
Module 5

Classification Metrics & Threshold Choice

IntermediatePro

Topics include Confusion Matrix, Four Numbers That Matter, Precision, Recall, F1, Accuracy, ROC and PR Curves, and more.

View 6 sections
  1. 1The Confusion Matrix: Where Every Other Metric Comes From
  2. 2Precision, Recall, F1: Tradeoffs, Not Synonyms
  3. 3ROC and PR Curves: Whole-Model Reports
  4. 4Threshold Choice: The Most Underrated Lever in ML
  5. 5Class Imbalance: When Accuracy Lies
  6. 6Lendly Case: Picking a Default-Risk Threshold
65 min6 sections
Open module
06
Module 6

Cross-Validation, Leakage & Stratification

IntermediatePro

Topics include Why One Split Is Not Enough, K-Fold Cross-Validation, Stratified K-Fold, and more.

View 6 sections
  1. 1The Problem with One Train/Test Split
  2. 2K-Fold Cross-Validation: One Idea, Many Variants
  3. 3Stratified K-Fold: Keep Class Balance Across Folds
  4. 4Data Leakage: The Silent Model Killer
  5. 5Time-Series CV: Why Random Folds Are Wrong for Temporal Data
  6. 6Lendly Case: 5-Fold CV the Right Way
65 min6 sections
Open module
07
Module 7

Pipelines, Preprocessing & GridSearchCV

IntermediatePro

Topics include Why Pipelines, The Anti-Leakage Pattern, ColumnTransformer for Mixed Feature Types, OneHotEncoder, StandardScaler, SimpleImputer, and more.

View 6 sections
  1. 1The Anti-Leakage Pattern: Why Pipelines Exist
  2. 2ColumnTransformer: One Transformer per Column Type
  3. 3Imputation, Scaling, Encoding: The Trio
  4. 4GridSearchCV: Searching the Hyperparameter Space
  5. 5RandomizedSearchCV: Smarter When Grids Explode
  6. 6Lendly Case: A Production-Style Pipeline
65 min6 sections
Open module
08
Module 8

Capstone: End-to-End on Lendly

IntermediatePro

Topics include The Full ML Loop, EDA in 15 Minutes, Splitting, Pipelining, Tuning, and more.

View 6 sections
  1. 1The Full Loop: From CSV to Decision
  2. 2Fast EDA: What to Look for in 15 Minutes
  3. 3Build, Tune, Evaluate: The Capstone Pipeline
  4. 4Threshold and Calibration: The Last Mile
  5. 5The Model Card: How to Hand Off a Model
  6. 6Where to Go Next: A/B Testing, LLMs, and Beyond
65 min6 sections
Open module
Who this course is for

Built for people who need to use the skill.

01

Aspiring data scientists

02

Analysts moving from description to prediction

03

ML practitioners repairing gaps in their workflow

Start the course

Begin with The ML Frame & the Lendly Dataset.

The first module establishes the language and example used throughout the rest of the course.

Open Module 1
Machine Learning Fundamentals (scikit-learn) | Let's Data Science | Let's Data Science