Gradient descent

From the free encyclopedia · Optimization

Learning rate

Each step moves the parameters against the gradient, scaled by a step size η called the learning rate: x ← x − η∇f(x). 1The whole method in one line: the gradient picks the direction, η picks how far to go.Choosing η is a trade-off between speed and stability.

2This is why the η = 1.1 example bounces: on f(x) = x², any η above 1 makes each step overshoot further.If the rate is too large, each update overshoots the minimum and the iterates can oscillate or diverge. If it is too small, progress is slow and training may stall on flat regions.

Figure 2. With a large step size, iterates jump from one side of the valley to the other before settling.

In practice the rate is often decayed over time, or adapted per parameter, as in AdaGrad and Adam.