What it is
Every model you have ever used learned by doing one thing over and over: measure how wrong it is, then nudge its settings in the direction that makes it less wrong. That nudge is gradient descent. It is the engine behind the feed ranking on your phone, the autocomplete in your editor, and the model that decided your last playlist.
In this module the hill is a surface you can see, and the ball is a ball. Two settings, x and y, instead of the millions a real model has. The loop is identical.
The slope is the hint
The gradient is the slope under the ball. It points uphill, toward "more wrong". So the update goes the other way:
x = x - learningRate * slopeX
y = y - learningRate * slopeY
That minus sign is the most important character in machine learning. Subtract the slope and you go downhill. Add it and you climb. The Debug mission has exactly that bug in it, so keep the sign in mind.
Learning rate
The learning rate is the step size. It is one number, and it decides whether the run lands or dies.
- 0.1 on the bowl: each step removes twenty percent of the distance to the bottom. Smooth landing in about thirty steps.
- 1.1 on the bowl: each step overshoots to the other side, further out than before. The loss goes up every step. That is divergence, and the demo stops when the numbers blow up.
- 0.001 on the bowl: nothing breaks and nothing happens. After a hundred steps the ball has barely moved.
There is no universal right value. There is a right value for this surface, and finding it is a skill. The Tune mission is that skill.
Momentum
Real loss surfaces are not bowls. They have long flat valleys where the slope is tiny, so a plain step is tiny too, and the ball crawls. Momentum keeps a fraction of the last step's velocity and adds it to this one:
velocity = momentum * velocity - learningRate * slope
x = x + velocity
With momentum 0.9 the ball keeps ninety percent of its speed. It carries across the flat parts and rolls through small bumps. It also overshoots, so it can take a while to settle. Try Rosenbrock, the curved valley, with and without momentum. Same learning rate, very different trip.
When it goes wrong
Three things you will see in the demo, and later in real training logs:
- Loss goes up every step. Learning rate too high. Divide it by ten.
- Loss flat, ball not moving. Learning rate too low, or the ball is stuck on a saddle where the slope is zero in the useful direction. Momentum helps with the second one.
- Loss bounces around a value and never settles. Slightly too high. Lower it a little, or let momentum decay.
That is the whole module. Run the demo until the three shapes of the loss curve look familiar, then go clear the missions.