🌱 AI Engineering · AI Foundations (No Math Fear) · cozy lesson
Data, Training and Testing
9 min · 1 min read · no scary math, promise
🤖
You’ve got this. Read a little, play a little — I’ll wait. No rush.
Split or it didn't happen
70/15/15. Shuffle, stratify labels. Time data? Split by time, not random.
Under vs over
- Underfit: too simple, both bad → bigger model, better features
- Overfit: train great, test bad → more data, dropout, early stop
- Just right: gap small, both good
Leakage horror
Including future info (e.g. “returned?” column when predicting returns) gives 99% fake accuracy. Always ask: “Would I know this at prediction time?”
💛 Enjoying? Try 5 playful quizzes or watch it move.
Check your understanding
Correct answers earn XP (once each).
1. Why keep test set locked?
2. Signal of overfitting?
My notes (saved in this browser)
Select text above → Save selection, or write your own. AlgoMaster-style notebook, local-first for MVP.
No notes yet. Your highlights will live here.
Finished reading? Seal it with a tick ✅
The checkbox in the explorer turns green too — same progress.