Asking what problem the operator solves

Batch normalization subtracts a batch mean and divides by a standard deviation. The manuscript starts from scale redundancy: several parameter settings can represent the same network function, and normalization chooses a representative. It expresses that choice as a strictly convex problem for a fixed batch and channel, then checks the relation between its solution and the familiar operator.

Testing explanations with distinguishable interventions

The study compares normalization rules with identical forward outputs but different backward derivatives, using symbolic checks, float64 operator checks and limited training comparisons. The purpose is to identify which effect an explanation actually supports. Evidence covers limited architectures, batch sizes and budgets; the backward intervention also changes weight-norm drift, so its differences cannot all be attributed to a single gradient mechanism.

Placing normalization inside a tracking family

The manuscript views batch normalization as an exact solution of a per-batch objective, then constructs a family that tracks the objective over time. Standard BN is the one-step corner; other members retain mean and scale states and use feedback gains to determine their response to a new batch. A Kalman view connects assumed observation noise with tracking speed, making the amount of smoothing testable.

How the noise model changes actual behaviour

Controlled ResNet-20 and CIFAR experiments separate training statistics, stored inference state and test-time re-estimation. Treating spatial activations as independent samples can understate noise and keep a tracker too close to single-batch statistics. Image-level noise estimates change its small-batch behaviour, but this does not establish general superiority over BN or group normalization. The study distinguishes failed controller claims, benefits of test statistics and exploratory findings.

Current progress and validation scope

The theory and experiments depend on fixed batches, explicit noise assumptions and limited comparisons; they do not establish general superiority over BN or other normalization methods.

Discuss this research

I welcome conversations about the questions, methods, and ways to test them.

lancer20060105@gmail.com