Skip to main content
Optimization Algorithms

Optimization Algorithms

The Optimization Landscape

Training a neural network means searching for a good set of weights in a landscape with millions of dimensions. Imagine you are blindfolded on a mountain range, and you can only feel the slope directly beneath your feet. Your goal is to find the lowest valley — but you cannot see the terrain ahead, and there are countless valleys, ridges, and plateaus in every direction. Training neural networks = finding good minima in a non-convex loss landscape. Challenges:
  • Saddle points (more common than local minima in high dimensions — in a 100-million parameter model, a true local minimum requires the loss to curve upward in all 100 million directions simultaneously, which is astronomically unlikely)
  • Flat regions with vanishing gradients (you are on a plateau and cannot feel which direction is downhill)
  • Sharp vs flat minima (sharp minima generalize poorly because tiny weight perturbations send you uphill; flat minima are robust)
  • Ill-conditioned Hessians (the loss surface curves steeply in some directions and gently in others, making a single learning rate suboptimal)
A useful mental model: In high-dimensional spaces, most critical points are saddle points, not local minima. SGD with momentum naturally escapes saddle points because the momentum carries you past the “saddle” even when the gradient is momentarily zero. This is one reason why simple optimizers work surprisingly well in practice.

Gradient Descent Variants

Vanilla SGD

wt+1=wtηwL(wt)w_{t+1} = w_t - \eta \nabla_w \mathcal{L}(w_t)
Problems: Oscillates in steep-walled ravines (zigzags back and forth instead of heading down the valley), painfully slow in flat regions (tiny gradients = tiny steps), and uses the same learning rate for all parameters (but some parameters need big updates while others need small ones).

SGD with Momentum

vt=μvt1+wL(wt)v_t = \mu v_{t-1} + \nabla_w \mathcal{L}(w_t) wt+1=wtηvtw_{t+1} = w_t - \eta v_t

Adaptive Learning Rates

RMSprop

Adapt learning rate per-parameter based on gradient history: vt=βvt1+(1β)gt2v_t = \beta v_{t-1} + (1 - \beta) g_t^2 wt+1=wtηvt+ϵgtw_{t+1} = w_t - \frac{\eta}{\sqrt{v_t + \epsilon}} g_t

Adam (Adaptive Moment Estimation)

Combines momentum and adaptive learning rates: mt=β1mt1+(1β1)gt(first moment)m_t = \beta_1 m_{t-1} + (1 - \beta_1) g_t \quad \text{(first moment)} vt=β2vt1+(1β2)gt2(second moment)v_t = \beta_2 v_{t-1} + (1 - \beta_2) g_t^2 \quad \text{(second moment)} m^t=mt1β1t,v^t=vt1β2t(bias correction)\hat{m}_t = \frac{m_t}{1 - \beta_1^t}, \quad \hat{v}_t = \frac{v_t}{1 - \beta_2^t} \quad \text{(bias correction)} wt+1=wtηv^t+ϵm^tw_{t+1} = w_t - \frac{\eta}{\sqrt{\hat{v}_t} + \epsilon} \hat{m}_t

AdamW (Decoupled Weight Decay)

Standard Adam applies L2 regularization incorrectly with adaptive learning rates. The problem: Adam divides gradients by the square root of the second moment (essentially adapting the learning rate per parameter). When you add L2 regularization to the loss, the regularization gradient also gets divided by this adaptive term, meaning frequently-updated parameters get less regularization than rare ones. This is the opposite of what you want. AdamW decouples weight decay from the adaptive update, applying it directly to the weights:
Always use AdamW over Adam for training modern deep networks. The decoupling was shown by Loshchilov and Hutter (2019) to improve generalization across almost every setting. If you see code using Adam with a weight_decay argument, that is the wrong formulation — switch to AdamW.
Pitfall — learning rate and weight decay interaction: When using AdamW, the effective regularization strength is lr * weight_decay. If you double the learning rate, you should halve the weight decay to maintain the same effective regularization. This coupling catches many people off guard when tuning hyperparameters. Some frameworks (like timm) use an absolute weight decay convention to avoid this confusion.

Learning Rate Schedules

Step Decay

Cosine Annealing

Warmup + Cosine (Transformer Standard)

This is the de facto standard schedule for Transformers and modern deep learning. The intuition: at the beginning of training, the model’s weights are random, so gradients are noisy and unreliable. Starting with a high learning rate would send the model careening in random directions. Warmup starts with a tiny learning rate and linearly increases it, giving the optimizer time to accumulate reliable gradient statistics (the first and second moments in Adam) before taking large steps. After warmup, cosine decay gradually reduces the learning rate, which acts like a fine-tuning phase — large steps explore broadly, small steps refine the solution.
How long should warmup be? A common rule of thumb: 1-5% of total training steps for large models (LLMs typically use 2000 warmup steps), or 5-10 epochs for vision tasks. Too little warmup and you may get early training instability (loss spikes); too much warmup wastes compute on a suboptimally low learning rate. If you see loss spikes in the first few hundred steps, try increasing warmup.

Optimizer Comparison


Modern Optimizers

Lion (Evolved Sign Momentum)

Sharpness-Aware Minimization (SAM)

Seeks flat minima by perturbing weights:

Practical Recipes

Vision Transformers

Large Language Models


Exercises

Train CIFAR-10 with SGD, Adam, and AdamW. Compare convergence speed and final accuracy.
Compare constant LR vs step decay vs cosine annealing on the same model.
Train a transformer with and without warmup. Observe gradient norms and loss curves.

What’s Next

Module 19: Training Strategies at Scale

Mixed precision, gradient accumulation, distributed training, and more.