Research / Information theory & reasoning

VERIFICATION / INFORMATION / REASONING

The Verifier-Bit Ledger

Budget, audit, and price the information carried by verifier feedback.

Research manuscript · Ongoing research

An information ledger for iterative reasoning: account for verifier feedback, establish an uplift ceiling, and inspect gains that the logged channel cannot explain.

The manuscript is not publicly available yet.

Concept artwork: measured feedback paths entering an ivory ledger, with a copper path outside the gate
A visual study of accounted feedback and information beyond the ledger.

01 / VERIFIER-BIT LEDGER

Where does the improvement come from?

A higher hidden-test pass rate can reflect better repair or information supplied by the evaluation channel. We ask two questions: how much improvement can a given amount of feedback information buy, and can an auditor use controlled interactions and samples to identify information that the logged channel cannot explain?

02 / VERIFIER-BIT LEDGER

Account for every round

Each round is charged for the extra information it reveals about the hidden specification, after conditioning on the task, earlier feedback, and policy randomness. The charges sum to the information in the entire interaction. This is an information budget rather than a word count: a verified structured renderer can produce a long message whose only specification-dependent input is one verdict bit.

RESEARCH FIGURE / LEDGER

The verifier-bit ledger

Original research schematic: feedback, a decoy-driven ghost baseline, the information ceiling and an executable audit. The declared-alphabet and residual extensions are described below.

03 / VERIFIER-BIT LEDGER

Keep the interaction, replace the answer

The ghost baseline runs the full repair loop with the verifier keyed to an independently drawn decoy specification, then scores the output against the true specification. It preserves the generic scaffolding supplied by interaction. Improvement over this baseline is the quantity the ledger must explain, with imperfect decoy distributions accounted for in the audit's error allowance.

04 / VERIFIER-BIT LEDGER

From a ceiling to a certificate

The theory bounds uplift for any policy. Its information-based bound has a first-order optimal constant at small budgets; a separate cardinality bound is exactly attainable for declared feedback alphabets. For example, one verdict bit permits at most 0.5 uplift. A certificate requires improvement beyond the applicable ceiling plus sampling uncertainty. An audit that does not trigger remains inconclusive.

Information-based ceiling|S − qghost| ≤ min{√(B / 2), √(1 − e−B)}

Scores in [0,1]; B is measured in nats. A separate declared-alphabet bound gives 0.5 for a one-bit transcript.

05 / VERIFIER-BIT LEDGER

Amortized gains are prepaid

Two conservation results account for re-solving training instances and transferring information across instances. Information acquired in training and fresh information acquired at test time jointly pay for improvement. In a shared-secret synthetic family, amortization replaces repeated information payments with a single payment, giving a separation that grows with the number of instances. This characterizes information reuse; it does not by itself establish a performance advantage for a particular model size.

06 / VERIFIER-BIT LEDGER

Three layers of validation

An exactly solvable family checks the accounting and audit calibration. Real code-repair experiments compare feedback prices, while controlled real-task protocols test declared-budget certificates. A complementary residual audit conditions on information already disclosed by legal feedback. With randomized hidden suites and a model in the repair loop, it detects a scripted policy exploiting one additional leaked test.

RESEARCH FIGURE / HERO

Synthetic metrology and conservation

Exact accounting, held-out calibration, audit power and training conservation. The calibration panel shows one held-out seed; the maximum anchor error across three held-out seeds is 0.65%.

07 / VERIFIER-BIT LEDGER

Observed code-repair outcomes

A controlled feedback comparison uses 378 paired MBPP+ tasks and Qwen2.5-Coder-7B, with deterministic decoding and four stateful repair rounds from shared initial candidates. The figure below compares ghost feedback, ordinary visible full input/output feedback, and a planted hidden failing example.

Ghost baseline70.63%267 / 378
Ordinary full-I/O73.02%276 / 378
Planted hidden failcase75.66%286 / 378
+5.03 percentage points vs ghost

Paired bootstrap 95% interval: 2.38–7.94 points; exact McNemar p=0.00055. Compared with ordinary full-I/O feedback, the 2.65-point difference has p=0.0755. These free-text experiments provide pricing and behavioural evidence.

RESEARCH FIGURE / REALTASK

Repair outcomes and feedback prices

Original figure, 378 paired MBPP+ tasks. Counts: ghost 267, ordinary full-I/O 276, planted hidden failcase 286. Uplift over ghost: 5.03 percentage points (95% paired bootstrap CI 2.38–7.94; p=0.00055). The difference from ordinary full-I/O is not significant (p=0.0755). The original rate axis spans 65–85%; full rates are listed above.

08 / VERIFIER-BIT LEDGER

Controlled audit certificates

A separate test uses declared or structurally capped feedback. At a one-bit cap, the current alphabet ceiling is 0.5. Each true/ghost audit contains 2,000 episodes, with a 5% false-alarm level and a 0.060736 sampling threshold.

G = S − qghost − 0.5; a certificate requires G > 0.060736
ProtocolHonest GPlanted G
Leaderboard-0.3135+0.3050
Verdict only-0.4580+0.3510
Structured feedback-0.4985+0.3795

The planted arms exceed the threshold; the tested honest aggregate audits stay below it. The leaderboard protocol uses controlled programs without an LLM. The two Qwen protocols score hidden-convention compliance, rather than ordinary benchmark code-repair success. These are planted positive controls.

AVERAGE FEEDBACK LENGTH193.8 tokens
HIDDEN-DEPENDENT CAP1 bit

A checked renderer may emit a long message while depending on the hidden specification through a single verdict bit. The contextual encoding price is 782.0 nats; it measures a different object from the structural cap.

09 / VERIFIER-BIT LEDGER

Beyond legally disclosed feedback

The residual audit conditions on what legal first-failure feedback already disclosed. The source records 321 randomized-suite episodes from 107 MBPP+ tasks, with Qwen2.5-Coder-7B in a two-round repair loop. An offline script constructs positive controls from cached honest final-program pass vectors.

35 episodes

Median time to certification for the full scripted one-test exploit among 400 reorderings of the same 321 episodes. All 400 reorderings certify; the honest trajectory crosses the sequential threshold in 1/400.

The 400 reorderings reuse the same data. The successful controls are offline scripted transformations; the LLM did not effectively exploit the supplied leaks. The certificate depends on randomized suites and a known verifier mechanism.

10 / VERIFIER-BIT LEDGER

Current scope and open questions

The verifier-bit ledger is the information-accounting and audit part of the broader Ω / OMEGA-AO research program. The current manuscript develops ceilings, conservation results, and protocol experiments. General free-text feedback is primarily priced; residual certificates require randomized suites and a known verifier mechanism. Certified black-box estimation, realistic threats, and tightness at intermediate information budgets remain open research questions.

This page presents a research overview and selected figures. The full manuscript and submission materials remain unpublished here.
Discuss the research

Swipe horizontally to read the full figure.