What recursive training can lose
This research examines what generative models lose when repeatedly trained on their own data, and why apparent sample quality can hide distributional degradation. Scale, macroscopic structure and effective support are observed separately, with attention to interactions among recursive pipelines, finite sampling and long tails. A dataset can pass individual-sample checks while losing diversity or rare modes. The project distinguishes sample validity from population health and seeks diagnostics that locate causes rather than report one aggregate score.
Keeping the research objects separate
The folder holds two kinds of work. WDPR is a proposal for distribution-level repair under recursive training, organizing reweighting and supplementary generation around a potential function. The SAE manuscript concerns statistical detection of weak features in superposition and how noise and an adjustable null affect difficulty. One studies long-term data-distribution degradation, the other feature identification in activations; their unverified performance claims should not be pooled.
Distributional diagnosis and branch interventions
The research combines tractable distributions, synthetic-data experiments and branch comparisons in generative pipelines. Diagnostics track multiscale structure and tail deficits. Candidate interventions include moment calibration, sampler changes, tail protection, conditioning anchors and windowed resampling, each tied to a stated degradation channel. Earlier architecture also explored cluster monitoring, conditional generation and data supplementation. Routes retain their own assumptions and versions; ablations distinguish restoration of scale, coverage and task performance.
From controlled phenomena to real-model evidence
Existing materials include revised theory, runnable toy experiments and a broader research synthesis for generative pipelines. Some decompositions rely on surrogate distributions or lower-bound assumptions, and geometric diagnostics still need a testable bridge to semantic quality. Controlled effects do not automatically generalize to long-term language-model self-training. Further work emphasizes frozen real pipelines, strong controls, multiple seeds and explicit costs, preserving failed predictions and adverse results. The project remains mechanism research and method exploration.
Discuss this research
I welcome conversations about the questions, methods, and ways to test them.
lancer20060105@gmail.com