One question at a time
A decision tree asks about the data the way you would play twenty questions. Is x less than 0.8? Yes, go left. No, go right. On each side it asks another question, and it keeps going until it reaches a leaf. The leaf says which class most of the training points that landed there belonged to.
Two features, x and y, means every question is a vertical or horizontal line on the map. The regions the tree carves out are rectangles. Watch the heat map in the demo and you will see them.
Which question first
Every candidate split gets a score: how much purer are the two sides than the whole was? Purity is measured by Gini impurity:
gini = 1 - p0² - p1²
A group that is all one class scores 0. A fifty-fifty mix scores 0.5. The gain of a split is the impurity before it minus the impurity after, weighted by how many points went each way. The engine tries every threshold on both features and keeps the biggest gain. Then it does the same on each side.
Deeper is not better
Given enough depth, a tree can fit any training set perfectly. It just keeps asking until every leaf is pure, even if that means carving a region around a single point whose label was flipped by noise. On the training set that looks like a win: 100 percent. It has not learned the shape of the data. It has memorised the list.
That is why the demo keeps a test set the tree never sees. Filled points are training data. Hollow points are test data. Press sweep depths and read the two lines:
- Train accuracy only ever goes up with depth. It has to.
- Test accuracy climbs to a peak around depth 2 or 3, then slides.
The gap between the lines is the overfit. The depth where test peaks is where you stop.
Where it goes wrong
- Depth 0. One leaf, no questions, majority class for everyone. A coin flip on balanced data.
- Depth 12 on noisy data. Perfect on train, worse than depth 2 on test. The tree is fitting the flipped labels.
- Noise going up. Raise the label noise slider and the best depth moves down. Noisy data wants simpler models. That is not a tree rule, it is a modelling rule.
Grow the tree at depth 1, 3 and 12, and look at the map each time. Smooth borders are learning. Jagged islands around single points are memorising.