The Residual Stream's Effective Depth
In the authors' words
We introduce effective depth (), a scalar diagnostic that treats the layer-wise residual stream of a transformer as a discrete-time process, measures how representation similarity decays with layer distance, and aggregates that profile into one number. Across sixteen decoder-only language models, separates a structural consequence of residual accumulation from an empirical one: even maximally diverse orthogonal updates have the closed-form reference , yet fifteen of sixteen default measurements lie below (Qwen3.5: 32--44%, OLMo-2: 40--41%, Pythia: 23--28%). Matched references show that the gap is not caused by the persistent initial state or update-size imbalance, but is largely a calibrated signature of correlated residual updates rather than evidence that depth is unused. Symmetric position-0, token-normalisation, and top-PC controls show the regime is not reducible to BOS or top-PC artefacts: the lone above-reference default outlier joins the same regime, and all sixteen models are sub-reference after token-normalisation or top-1-PC removal. Intermediate checkpoints show that the regime is established early in OLMo-2 and stable through 5T tokens, while Pythia-1.4B follows a distinct decreasing trajectory. A controlled residual-carry intervention supports the mechanism, and is best read as a global accumulated-state diagnostic, not as a capability score or pruning method.
Appeared: Monday, September 28. arXiv. Preprint, not yet peer-reviewed.
Authors' comment: Accepted at the 17th Asian Conference on Machine Learning (ACML 2026)