Optimizers Aren't Models: They Move the Weights
Linear regression is the model. SGD, momentum and Adam are optimizers: the rule that updates its weights. Run all three on the same model and they land on the same answer.

The conversation
I say: we need to optimize the loss function. The junior developer says: oh, like linear regression? And then: so Adam is the better model?
It is an easy mix-up, and it makes conversations about training go in circles, because the two words name different layers of the same system.
The 43 second version

Optimizers aren't models

The model is the function
Linear regression is a model: a function with weights, y = w x + b. Given weights it makes predictions. It says nothing about how you find good weights.
The optimizer is the update rule
An optimizer is the rule that moves the weights after every mistake: w <- w - learning rate x gradient of the loss. Plain gradient descent (SGD) is exactly that line. Momentum adds a running average of past gradients so the update keeps moving in a consistent direction. Adam keeps running averages of both the gradient and its square, and scales each weight's step by them.
None of them can predict anything. They only map (weights, gradient) to new weights, and the repo's test test_optimizer_has_no_prediction_method makes that literal.
c10_optimizers/optimizers.py
class SGD:
def __init__(self, lr=0.1):
self.lr = lr
def step(self, p, g):
return p - self.lr * g
class Momentum:
def __init__(self, lr=0.1, beta=0.9):
self.lr, self.beta, self.v = lr, beta, None
def step(self, p, g):
self.v = g if self.v is None else self.beta * self.v + g
return p - self.lr * self.v
class Adam:
def __init__(self, lr=0.1, b1=0.9, b2=0.999, eps=1e-8):
self.lr, self.b1, self.b2, self.eps = lr, b1, b2, eps
self.m = self.v = None
self.t = 0
def step(self, p, g):
if self.m is None:
self.m, self.v = np.zeros_like(g), np.zeros_like(g)
self.t += 1
self.m = self.b1 * self.m + (1 - self.b1) * g
self.v = self.b2 * self.v + (1 - self.b2) * g * g
m_hat = self.m / (1 - self.b1 ** self.t)
v_hat = self.v / (1 - self.b2 ** self.t)
return p - self.lr * m_hat / (np.sqrt(v_hat) + self.eps)Same model, three optimizers
The repo fits one linear regression, on the same data as part 4, with each optimizer for 200 steps:
- SGD: w = [1.997, -2.994], b = 0.5
- Momentum: w = [1.998, -2.994], b = 0.5
- Adam: w = [1.997, -2.994], b = 0.5
Same answer, because it is the same model and the same loss. What differs is the path: SGD reaches a loss under 0.01 in 19 steps, momentum in 66 and Adam in 69. On a small, well-scaled problem with two weights, plain SGD is the fastest. Adam earns its reputation on large, noisy, badly scaled problems, which is why transformers are trained with AdamW, not because it is a better model.
Same model, three optimizers

Only the path differs

c10_optimizers/optimizers.py
def fit(optimizer, X, y, steps=200):
"""Same model every time: p = [w1, w2, b]. Only `optimizer.step` differs."""
p = np.zeros(X.shape[1] + 1)
history = []
for _ in range(steps):
gw, gb = grads(p[:-1], p[-1], X, y)
p = optimizer.step(p, np.append(gw, gb))
history.append(mse(X @ p[:-1] + p[-1], y))
return p, historyRun it yourself
The folder is c10_optimizers/ in the companion repo, with three tests: every optimizer finds the same weights, an optimizer has no predict method, and Adam's first bias corrected step has the size of the learning rate whatever the gradient's scale.
Run this concept
git clone https://gitlab.com/j.michaelduhart/our-saviors.git
cd our-saviors && pip install -r requirements.txt
python c10_optimizers/optimizers.py
python -m pytest c10_optimizersNext: learning rates, the one number that decides whether any optimizer works at all.
More: LinkedIn · Instagram. Portfolio and case studies: designsbyduhart.org.
If any of this saved you an afternoon, Buy me a coffee.