01 / VERIFIER-BIT LEDGER
Where does the improvement come from?
A higher hidden-test pass rate can reflect better repair or information supplied by the evaluation channel. We ask two questions: how much improvement can a given amount of feedback information buy, and can an auditor use controlled interactions and samples to identify information that the logged channel cannot explain?
02 / VERIFIER-BIT LEDGER
Account for every round
Each round is charged for the extra information it reveals about the hidden specification, after conditioning on the task, earlier feedback, and policy randomness. The charges sum to the information in the entire interaction. This is an information budget rather than a word count: a verified structured renderer can produce a long message whose only specification-dependent input is one verdict bit.
RESEARCH FIGURE / LEDGER
The verifier-bit ledger
03 / VERIFIER-BIT LEDGER
Keep the interaction, replace the answer
The ghost baseline runs the full repair loop with the verifier keyed to an independently drawn decoy specification, then scores the output against the true specification. It preserves the generic scaffolding supplied by interaction. Improvement over this baseline is the quantity the ledger must explain, with imperfect decoy distributions accounted for in the audit's error allowance.
04 / VERIFIER-BIT LEDGER
From a ceiling to a certificate
The theory bounds uplift for any policy. Its information-based bound has a first-order optimal constant at small budgets; a separate cardinality bound is exactly attainable for declared feedback alphabets. For example, one verdict bit permits at most 0.5 uplift. A certificate requires improvement beyond the applicable ceiling plus sampling uncertainty. An audit that does not trigger remains inconclusive.
Scores in [0,1]; B is measured in nats. A separate declared-alphabet bound gives 0.5 for a one-bit transcript.
05 / VERIFIER-BIT LEDGER
Amortized gains are prepaid
Two conservation results account for re-solving training instances and transferring information across instances. Information acquired in training and fresh information acquired at test time jointly pay for improvement. In a shared-secret synthetic family, amortization replaces repeated information payments with a single payment, giving a separation that grows with the number of instances. This characterizes information reuse; it does not by itself establish a performance advantage for a particular model size.
06 / VERIFIER-BIT LEDGER
Three layers of validation
An exactly solvable family checks the accounting and audit calibration. Real code-repair experiments compare feedback prices, while controlled real-task protocols test declared-budget certificates. A complementary residual audit conditions on information already disclosed by legal feedback. With randomized hidden suites and a model in the repair loop, it detects a scripted policy exploiting one additional leaked test.
RESEARCH FIGURE / HERO
Synthetic metrology and conservation
07 / VERIFIER-BIT LEDGER
Observed code-repair outcomes
A controlled feedback comparison uses 378 paired MBPP+ tasks and Qwen2.5-Coder-7B, with deterministic decoding and four stateful repair rounds from shared initial candidates. The figure below compares ghost feedback, ordinary visible full input/output feedback, and a planted hidden failing example.
Paired bootstrap 95% interval: 2.38–7.94 points; exact McNemar p=0.00055. Compared with ordinary full-I/O feedback, the 2.65-point difference has p=0.0755. These free-text experiments provide pricing and behavioural evidence.
RESEARCH FIGURE / REALTASK
Repair outcomes and feedback prices
08 / VERIFIER-BIT LEDGER
Controlled audit certificates
A separate test uses declared or structurally capped feedback. At a one-bit cap, the current alphabet ceiling is 0.5. Each true/ghost audit contains 2,000 episodes, with a 5% false-alarm level and a 0.060736 sampling threshold.
| Protocol | Honest G | Planted G |
|---|---|---|
| Leaderboard | -0.3135 | +0.3050 |
| Verdict only | -0.4580 | +0.3510 |
| Structured feedback | -0.4985 | +0.3795 |
The planted arms exceed the threshold; the tested honest aggregate audits stay below it. The leaderboard protocol uses controlled programs without an LLM. The two Qwen protocols score hidden-convention compliance, rather than ordinary benchmark code-repair success. These are planted positive controls.
A checked renderer may emit a long message while depending on the hidden specification through a single verdict bit. The contextual encoding price is 782.0 nats; it measures a different object from the structural cap.
09 / VERIFIER-BIT LEDGER
Beyond legally disclosed feedback
The residual audit conditions on what legal first-failure feedback already disclosed. The source records 321 randomized-suite episodes from 107 MBPP+ tasks, with Qwen2.5-Coder-7B in a two-round repair loop. An offline script constructs positive controls from cached honest final-program pass vectors.
Median time to certification for the full scripted one-test exploit among 400 reorderings of the same 321 episodes. All 400 reorderings certify; the honest trajectory crosses the sequential threshold in 1/400.
The 400 reorderings reuse the same data. The successful controls are offline scripted transformations; the LLM did not effectively exploit the supplied leaks. The certificate depends on randomized suites and a known verifier mechanism.
10 / VERIFIER-BIT LEDGER
Current scope and open questions
The verifier-bit ledger is the information-accounting and audit part of the broader Ω / OMEGA-AO research program. The current manuscript develops ceilings, conservation results, and protocol experiments. General free-text feedback is primarily priced; residual certificates require randomized suites and a known verifier mechanism. Certified black-box estimation, realistic threats, and tightness at intermediate information budgets remain open research questions.
