What an Evaluation Harness Decides
The model is only one part of a closed-loop benchmark. Parsing, recovery, timing, and rollout accounting determine what its score means.
Notes on evaluation, learning systems, and the assumptions that determine their behavior.
The model is only one part of a closed-loop benchmark. Parsing, recovery, timing, and rollout accounting determine what its score means.
Convolution is a wager: nearby values interact, and one rule should work everywhere. That explains both its efficiency and its blind spots.
A one-line recurrence separates smooth convergence, oscillation, a two-cycle, and a dead ReLU.