Intermediate credit beyond terminal reward
CIDS-GA addresses attribution in reasoning training. A successful trajectory does not identify which intermediate operation actually improved its prospects. Starting from errors made by the current policy, the method actively compares alternatives at the same state and prefix, using functional tests to certify differences. Training computation is used to discover localized repairs and transfer them into ordinary forward capability, so deployment relies on learned parameters rather than repeating training-time search and verification.
Same-state experiments, probability transport and rollback
Repairs are represented as complete typed transactions. Cards preserve policy versions, states, prefixes, alternatives, paired effects and functional evidence. Certified targets move probability mass from harmful operations toward better operations or functional classes; older certificates are reconsidered after policy changes. Candidate parameters face paired evaluation against both the current policy and a historical high-water mark, with atomic rollback on failure. A deterministic micro-executor separates legal actions, transaction prediction, functional correctness and retention.
The tension between learning and protection
Sandbox, small-model and fixed-pool real-code studies provide different layers of mechanism evidence, not a single universal score. Recent diagnosis found that constraints protecting historical behaviour can also suppress useful stopping repairs; the complete high-water-mark gate still rejects the current candidate. Further work explores protecting correct functionality while allowing certified action changes. Causal attribution, learnability, autonomous repair and capability retention are therefore assessed separately to test whether each component performs its intended role.
Discuss this research
I welcome conversations about the questions, methods, and ways to test them.
lancer20060105@gmail.com