Why average scale can hide uneven units
The latest initialization manuscript studies independent Gaussian weight rows and anisotropic inputs. Matching average activation scale does not ensure similar energy for units whose rows are sampled once and kept fixed. Input effective rank describes that variation, while an operator analysis examines how spectral range is allocated between forward computation and a backward surrogate. The study separates that surrogate from shared-weight backpropagation and develops an explicit two-layer ReLU analysis with Gaussian readout and squared loss. Its reported experiments do not establish a positive accuracy gain from covariance shaping alone.
Representation collapse and backward disagreement
The companion manuscript studies two linked effects in deep ReLU networks: forward representations of different inputs become alike, while their backward gradients lose agreement. It relates them to the value and derivative of the same correlation map and studies a usable-depth budget. The exponents depend on infinite-width and explicit independence assumptions; finite-width observations do not resolve the convergence rate, and an attention-sink extension remains a prediction. This collapse concerns representation propagation through depth, distinct from model collapse under recursive synthetic-data training.
Current progress and validation scope
The covariance-shaping experiments establish no accuracy benefit. Infinite-width laws, two-layer closed forms and finite-width observations are kept separate.
Discuss this research
I welcome conversations about the questions, methods, and ways to test them.
lancer20060105@gmail.com