1. Pick a Start
Click anywhere on the 2D loss surface. That's your initial guess — usually far from the minimum.
Click anywhere on the loss surface to drop a starting point, then watch the optimizer find the minimum — one step at a time.
Click anywhere on the 2D loss surface. That's your initial guess — usually far from the minimum.
The gradient tells us which direction is "downhill." It's the vector of partial derivatives of the loss with respect to each parameter.
We move opposite the gradient by a factor of the learning rate. Too big and you overshoot; too small and you crawl.
SGD bounces around, Momentum smooths the path, and Adam adapts per-parameter. Each finds the valley differently.