Gradient descent
Learning rate
Each step moves the parameters against the gradient, scaled by a step size η called the learning rate: x ← x − η∇f(x). 1The whole method in one line: the gradient picks the direction, η picks how far to go.Choosing η is a trade-off between speed and stability.
2This is why the η = 1.1 example bounces: on f(x) = x², any η above 1 makes each step overshoot further.If the rate is too large, each update overshoots the minimum and the iterates can oscillate or diverge. If it is too small, progress is slow and training may stall on flat regions.
In practice the rate is often decayed over time, or adapted per parameter, as in AdaGrad and Adam.