Skip to content

Foundations

Gradient Descent

Every model you have ever heard of learned by rolling downhill. This is the hill, the ball, and the one number that decides whether it lands or flies off the map.

difficulty
easy
time
30 min
xp available
720

02briefing

What it is

Every model you have ever used learned by doing one thing over and over: measure how wrong it is, then nudge its settings in the direction that makes it less wrong. That nudge is gradient descent. It is the engine behind the feed ranking on your phone, the autocomplete in your editor, and the model that decided your last playlist.

In this module the hill is a surface you can see, and the ball is a ball. Two settings, x and y, instead of the millions a real model has. The loop is identical.

The slope is the hint

The gradient is the slope under the ball. It points uphill, toward "more wrong". So the update goes the other way:

x = x - learningRate * slopeX
y = y - learningRate * slopeY

That minus sign is the most important character in machine learning. Subtract the slope and you go downhill. Add it and you climb. The Debug mission has exactly that bug in it, so keep the sign in mind.

Learning rate

The learning rate is the step size. It is one number, and it decides whether the run lands or dies.

  • 0.1 on the bowl: each step removes twenty percent of the distance to the bottom. Smooth landing in about thirty steps.
  • 1.1 on the bowl: each step overshoots to the other side, further out than before. The loss goes up every step. That is divergence, and the demo stops when the numbers blow up.
  • 0.001 on the bowl: nothing breaks and nothing happens. After a hundred steps the ball has barely moved.

There is no universal right value. There is a right value for this surface, and finding it is a skill. The Tune mission is that skill.

Momentum

Real loss surfaces are not bowls. They have long flat valleys where the slope is tiny, so a plain step is tiny too, and the ball crawls. Momentum keeps a fraction of the last step's velocity and adds it to this one:

velocity = momentum * velocity - learningRate * slope
x = x + velocity

With momentum 0.9 the ball keeps ninety percent of its speed. It carries across the flat parts and rolls through small bumps. It also overshoots, so it can take a while to settle. Try Rosenbrock, the curved valley, with and without momentum. Same learning rate, very different trip.

When it goes wrong

Three things you will see in the demo, and later in real training logs:

  1. Loss goes up every step. Learning rate too high. Divide it by ten.
  2. Loss flat, ball not moving. Learning rate too low, or the ball is stuck on a saddle where the slope is zero in the useful direction. Momentum helps with the second one.
  3. Loss bounces around a value and never settles. Slightly too high. Lower it a little, or let momentum decay.

That is the whole module. Run the demo until the three shapes of the loss curve look familiar, then go clear the missions.

03live demo

04missions

Predict

Which way does the ball roll?

Read the state, call the outcome before the engine does.

unranked
up to 120 xpclear Vectors and Tensors first
Tune

Find a learning rate that settles

Turn the knobs until the loss behaves.

unranked
up to 160 xpclear Vectors and Tensors first
Debug

The loss goes up. Why?

Something is wrong on purpose. Find it.

unranked
up to 200 xpclear Vectors and Tensors first
Code

Write descendBowl()

Write the function. The tests are the judge.

unranked
up to 240 xpclear Vectors and Tensors first

05in production

Every training run, everywhere

When a recommendation model on a music app gets retrained overnight, gradient descent is what runs. The loss is "how wrong were the recommendations", the parameters are millions of weights instead of two, and the update is the same line of code you write in this module: parameter minus learning rate times gradient. The fancy optimisers with names like Adam are this loop with a smarter learning rate per weight. If you can read this one, you can read those.

The learning rate is a production setting

Teams keep a learning rate schedule in version control the same way they keep a config file. Too high and a training job that costs thousands of dollars of GPU time diverges an hour in. Too low and it never gets there. The Tune mission is the exact judgement call an ML engineer makes before pressing run.

06code samples

engine/cpp/gradient_descent.hc++
// Gradient descent with heavy-ball momentum on a 2-D surface, one point per step.
#pragma once

namespace ne {

/** Beyond this on either axis the run is called diverged and stopped. */
constexpr double GD_DIVERGE_LIMIT = 1e6;

/**
 * Descend `surface` from (x0, y0) for `steps` updates.
 *
 * Writes (x, y, loss) triples to out_xyl: the starting point first, then one triple per
 * accepted step, so a full run writes steps + 1 triples and out_xyl must hold
 * 3 * (steps + 1) floats. Returns the number of triples written.
 *
 * Update rule, with v starting at zero:
 *   v <- momentum * v - lr * grad(x, y)
 *   (x, y) <- (x, y) + v
 * momentum 0 is plain gradient descent.
 *
 * A step that lands outside GD_DIVERGE_LIMIT on either axis, or on a non-finite value, is
 * not written; the run stops there and *out_diverged (when non-null) is set to 1. An
 * unknown surface id writes nothing and is reported as diverged, since the gradient is NaN.
 */
int gd_run(int surface, double x0, double y0, double lr, double momentum, int steps,
           float* out_xyl, int* out_diverged);

}  // namespace ne

07debrief

module tier

unranked

0 xp earned here

PredictWhich way does the ball roll?unranked
TuneFind a learning rate that settlesunranked
DebugThe loss goes up. Why?unranked
CodeWrite descendBowl()unranked

Run the demo for Bronze. Pass the quiz for Silver.

08share kit

Every AI model you use learned the same way: a ball, a hill, one number. Here is the number.

  1. Slide 1. Gradient descent is a ball rolling down a hill. The hill is how wrong the model is. The bottom is the answer.
  2. Slide 2. The slope tells the ball which way is down. That slope has a name: the gradient. Subtract it and you move downhill.
  3. Slide 3. The learning rate is the step size. 0.1 on this bowl: smooth landing. 1.1: the ball flies out of the bowl on the second step.
  4. Slide 4. Momentum keeps some of the last step. It carries the ball across flat stretches where plain descent gets stuck.
  5. Slide 5. This is the same loop that trains the model behind your feed, with millions of knobs instead of two. Run it yourself, link in bio.

story

  1. Frame 1. A neon bowl, a ball at the rim. Text: which way does it roll?
  2. Frame 2. The ball bounces out of the bowl. Text: learning rate 1.1. this is what a broken training run looks like.
  3. Frame 3. The skill tree with Gradient Descent lit up. Text: clear it, get the XP. NEURAL//RUN.

#machinelearning #gradientdescent #deeplearning #learntocode #ai #datascience #python #cplusplus #webassembly #threejs #codinglife #developer #mlengineer #neuralnetworks #learnml #techeducation #cyberpunk #interactivelearning

completion card

run the demo to earn a card
+50 credits, once
back to the tree