One ReLU Neuron, Four Learning Rates

One neuron is usually used to make backpropagation look simple. It is more interesting for the opposite reason: it is small enough to solve. With one ReLU, one example, and four learning rates, we can watch training converge, ring, enter a loop, or erase its own gradient.

A model small enough to solve

Consider a scalar input x=2x=2, target y=4y=4, and one ReLU neuron:

y^t=ReLU⁡(at),at=wtx+bt\hat y_t=\operatorname{ReLU}(a_t),\qquad a_t=w_tx+b_t

Initialize w0=1w_0=1 and b0=0b_0=0. The first pre-activation is therefore a0=2a_0=2. While at>0a_t>0, ReLU is the identity, so the squared loss Lt=(y^t−y)2L_t=(\hat y_t-y)^2 has gradients

∂Lt∂wt=2(at−y)x∂Lt∂bt=2(at−y)\begin{aligned} \frac{\partial L_t}{\partial w_t} &= 2(a_t-y)x \\ \frac{\partial L_t}{\partial b_t} &= 2(a_t-y) \end{aligned}

We could update the weight and bias separately. A cleaner view is to ask what one update does to ata_t, the only quantity that affects this example. Substituting the two gradient steps gives

at+1=x[wt−2η(at−y)x]+[bt−2η(at−y)]=at−2η(x2+1)(at−y),at+1−4=(1−10η)(at−4).\begin{aligned} a_{t+1} &= x\left[w_t-2\eta(a_t-y)x\right] \\ &\quad+ \left[b_t-2\eta(a_t-y)\right] \\ &= a_t-2\eta(x^2+1)(a_t-y), \\ a_{t+1}-4 &= (1-10\eta)(a_t-4). \end{aligned}

The last line is the whole training run, as long as the neuron remains active. If et=at−4e_t=a_t-4 is the prediction error, every step multiplies it by 1−10η1-10\eta. One example cannot determine the weight and bias independently; it only asks them to reach the line 2w+b=42w+b=4. The recurrence tells us how the neuron approaches, crosses, or misses that line.

Write that multiplier as q=1−10ηq=1-10\eta. Its magnitude determines whether the error shrinks, and its sign determines which side of the target comes next. When 0<q<10<q<1, the approach is monotone. When −1<q<0-1<q<0, the error changes sign but contracts. At q=−1q=-1, it never contracts; below −1-1, it grows. These are exact boundaries for the active part of this problem, not qualitative rules of thumb.

The learning rate chooses the behavior

Four learning-rate trajectories for one ReLU neuronFour panels plot pre-activation over six training steps. Learning rates 0.01 and 0.15 converge, 0.20 alternates between two values, and 0.25 enters the negative region where the ReLU gradient is zero.η = 0.01smooth convergence40η = 0.15alternating convergence40η = 0.20two-cycle40η = 0.25dead ReLU40gradient = 00246training step t
Pre-activation over six updates from the same initial parameters. Gray marks the inactive half-space; red is reserved for the trajectory that enters it.

With η=0.01\eta=0.01, the error multiplier is 0.90.9. The pre-activation moves monotonically: 2→2.2→2.38→⋯2\to2.2\to2.38\to\cdots. Learning is slow but uneventful.

At η=0.15\eta=0.15, the multiplier is −0.5-0.5. The sign change makes the neuron cross the target on every step, while the magnitude below one damps the crossings: 2→5→3.5→4.25→⋯2\to5\to3.5\to4.25\to\cdots. At the special value η=0.10\eta=0.10, the multiplier is zero and this single example is fit in one update.

At η=0.20\eta=0.20, the multiplier reaches −1-1. Nothing decays: the neuron repeats 2↔62\leftrightarrow6 forever. A gradient exists at every step, and each step faithfully follows it, but the step length prevents any progress.

The step that removes its own correction

The last trajectory is worse than ordinary oscillation. With η=0.25\eta=0.25, the active-region multiplier is −1.5-1.5. The sequence begins 2→7→−0.52\to7\to-0.5. That second update crosses the ReLU boundary. The output becomes zero, the loss becomes (0−4)2=16(0-4)^2=16, and the ReLU derivative is zero. Both parameter gradients vanish.

A linear neuron at a=−0.5a=-0.5 would still receive a gradient pointing back toward four. The ReLU neuron does not. It has entered a flat half-space in the loss surface: wrong, but without a local signal that can make it less wrong. The large step did not merely increase the error; it removed its own correction.

The geometry is simple enough to state exactly. In the active half-plane, the loss is (2w+b−4)2(2w+b-4)^2: curved across the line 2w+b=42w+b=4 and flat along it. In the inactive half-plane 2w+b<02w+b<0, the output is always zero, so the loss is the constant 16. The ReLU boundary joins a sloped surface to a plateau. Gradient descent can cross onto that plateau even though every preceding update pointed downhill locally.

Feature scale changes the stability bound

The factor of ten above is not arbitrary. It is 2(x2+1)2(x^2+1) for x=2x=2; the extra one comes from updating the bias. For a vector input, the same active-region calculation gives

et+1=[1−2η(∥x∥2+1)]ete_{t+1} =\left[1-2\eta\left(\lVert \mathbf{x}\rVert^2+1\right)\right]e_t

The error contracts only when the multiplier has magnitude below one. Therefore, for this one-example problem, the stable range is

0<η<1∥x∥2+10<\eta<\frac{1}{\lVert \mathbf{x}\rVert^2+1}

The rate that lands on the target in one active-region step is exactly half that upper bound: η=1/[2(∥x∥2+1)]\eta=1/[2(\lVert\mathbf{x}\rVert^2+1)]. This is not a generally optimal learning rate; it works because a single example produces a one-dimensional quadratic. It does make the dependence on feature magnitude explicit.

Rescaling the input changes that interval even if the optimizer is untouched. A learning rate that is conservative for small features can oscillate or diverge after those features are enlarged. The feature norm enters the error multiplier quadratically and lowers the upper bound of the stable interval.

What survives in larger models

A dataset with many examples no longer collapses to one scalar recurrence. The stability idea remains. For a quadratic objective, each eigendirection of the curvature has its own multiplier of the form 1−ηλ1-\eta\lambda. The direction with the largest curvature sets the strictest useful step size. Our neuron has only one relevant direction, with λ=2(∥x∥2+1)\lambda=2(\lVert\mathbf{x}\rVert^2+1).

A ReLU network is not globally quadratic. Its activation pattern divides parameter space into regions, and an update can cross from one region into another. That is precisely what the red trajectory does.

The recurrence describes one example only while its ReLU remains active. “Follow the negative gradient” is only a local instruction. The learning rate decides whether the next point remains in the region where that instruction was useful.