Designs by Duhart All writing

·2 min read·machinelearning · ai · deeplearning · datascience · python · numpy · mlengineering · softwareengineering

Optimizers Aren't Models: They Move the Weights

Linear regression is the model. SGD, momentum and Adam are optimizers: the rule that updates its weights. Run all three on the same model and they land on the same answer.

Slide 1 of 6 of "Optimizers Aren't Models: They Move the Weights", dark notebook background with copper and teal accents. Cover: Optimizers aren't models. They move the weights.

The conversation

I say: we need to optimize the loss function. The junior developer says: oh, like linear regression? And then: so Adam is the better model?

It is an easy mix-up, and it makes conversations about training go in circles, because the two words name different layers of the same system.

The 43 second version

Meme cold open, the explainer, and the payoff. Captions on.
Opening frame of the video: Optimizers Aren't Models: They Move the Weights
Meme cold open, the explainer, and the payoff. Captions on. Watch the video: https://designsbyduhart.org/blog/optimizers-are-not-models/

Optimizers aren't models

Slide 1 of 6 of "Optimizers Aren't Models: They Move the Weights", dark notebook background with copper and teal accents. Cover: Optimizers aren't models. They move the weights.
They move the weights.

The model is the function

Linear regression is a model: a function with weights, y = w x + b. Given weights it makes predictions. It says nothing about how you find good weights.

The optimizer is the update rule

An optimizer is the rule that moves the weights after every mistake: w <- w - learning rate x gradient of the loss. Plain gradient descent (SGD) is exactly that line. Momentum adds a running average of past gradients so the update keeps moving in a consistent direction. Adam keeps running averages of both the gradient and its square, and scales each weight's step by them.

None of them can predict anything. They only map (weights, gradient) to new weights, and the repo's test test_optimizer_has_no_prediction_method makes that literal.

c10_optimizers/optimizers.py

python
class SGD:
    def __init__(self, lr=0.1):
        self.lr = lr

    def step(self, p, g):
        return p - self.lr * g

class Momentum:
    def __init__(self, lr=0.1, beta=0.9):
        self.lr, self.beta, self.v = lr, beta, None

    def step(self, p, g):
        self.v = g if self.v is None else self.beta * self.v + g
        return p - self.lr * self.v

class Adam:
    def __init__(self, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
        self.lr, self.b1, self.b2, self.eps = lr, b1, b2, eps
        self.m = self.v = None
        self.t = 0

    def step(self, p, g):
        if self.m is None:
            self.m, self.v = np.zeros_like(g), np.zeros_like(g)
        self.t += 1
        self.m = self.b1 * self.m + (1 - self.b1) * g
        self.v = self.b2 * self.v + (1 - self.b2) * g * g
        m_hat = self.m / (1 - self.b1 ** self.t)
        v_hat = self.v / (1 - self.b2 ** self.t)
        return p - self.lr * m_hat / (np.sqrt(v_hat) + self.eps)

Same model, three optimizers

The repo fits one linear regression, on the same data as part 4, with each optimizer for 200 steps:

  • SGD: w = [1.997, -2.994], b = 0.5
  • Momentum: w = [1.998, -2.994], b = 0.5
  • Adam: w = [1.997, -2.994], b = 0.5

Same answer, because it is the same model and the same loss. What differs is the path: SGD reaches a loss under 0.01 in 19 steps, momentum in 66 and Adam in 69. On a small, well-scaled problem with two weights, plain SGD is the fastest. Adam earns its reputation on large, noisy, badly scaled problems, which is why transformers are trained with AdamW, not because it is a better model.

Same model, three optimizers

Slide 3 of 6 of "Optimizers Aren't Models: They Move the Weights", dark notebook background with copper and teal accents. Same model, three optimizers: SGD; Momentum; Adam, each with its value printed by the companion repo.
Final weights printed by the repo.

Only the path differs

Slide 4 of 6 of "Optimizers Aren't Models: They Move the Weights", dark notebook background with copper and teal accents. Only the path differs: SGD, steps to a loss under 0.01; Momentum; Adam, each with its value printed by the companion repo.
Steps to a loss under 0.01.

c10_optimizers/optimizers.py

python
def fit(optimizer, X, y, steps=200):
    """Same model every time: p = [w1, w2, b]. Only `optimizer.step` differs."""
    p = np.zeros(X.shape[1] + 1)
    history = []
    for _ in range(steps):
        gw, gb = grads(p[:-1], p[-1], X, y)
        p = optimizer.step(p, np.append(gw, gb))
        history.append(mse(X @ p[:-1] + p[-1], y))
    return p, history

Run it yourself

The folder is c10_optimizers/ in the companion repo, with three tests: every optimizer finds the same weights, an optimizer has no predict method, and Adam's first bias corrected step has the size of the learning rate whatever the gradient's scale.

Run this concept

bash
git clone https://gitlab.com/j.michaelduhart/our-saviors.git
cd our-saviors && pip install -r requirements.txt
python c10_optimizers/optimizers.py
python -m pytest c10_optimizers

Next: learning rates, the one number that decides whether any optimizer works at all.

More: LinkedIn · Instagram. Portfolio and case studies: designsbyduhart.org.

If any of this saved you an afternoon, Buy me a coffee.