🌱SproutStack
👤 Guest

🌱 AI Engineering · AI Foundations (No Math Fear) · cozy lesson

Data, Training and Testing

9 min · 1 min read · no scary math, promise

🤖
You’ve got this. Read a little, play a little — I’ll wait. No rush.

Split or it didn't happen

70/15/15. Shuffle, stratify labels. Time data? Split by time, not random.

Under vs over

  • Underfit: too simple, both bad → bigger model, better features
  • Overfit: train great, test bad → more data, dropout, early stop
  • Just right: gap small, both good

Leakage horror

Including future info (e.g. “returned?” column when predicting returns) gives 99% fake accuracy. Always ask: “Would I know this at prediction time?”

Check your understanding

Correct answers earn XP (once each).

1. Why keep test set locked?

2. Signal of overfitting?

My notes (saved in this browser)

Select text above → Save selection, or write your own. AlgoMaster-style notebook, local-first for MVP.

No notes yet. Your highlights will live here.

Finished reading? Seal it with a tick ✅

The checkbox in the explorer turns green too — same progress.