A research proposal generated by Ariadne. The experiments below are planned, and the review scores describe this internal selection.
核心洞见: SP 的有效性来自"受访者私有信息不同",而不是"让模型预测人群"。同一模型的链可以通过是否获得独立执行信号来制造这种差异:被工具知情链比朴素总体更多背书的答案,是多数投票看不到的真值证据。预测项取 N 类实测分布,判据是似然比检验;正确性的充分条件是执行信号对真答案的"提升"大于对任何具体错误吸引子的提升(r(a*)>r(a)),该条件可在标注开发集上预先检验。
本轮 top-k 排名稳定性 1.00 · BT 强度 1.86 · AC 加权分 3.15/5(borderline)· 查新 NEAR(风险 medium,证据等级 corpus)· 扛过红队 1 轮 · direction: Heterogeneous voter ensemble methods
用"是否接触独立执行信号"把同一模型的采样链人为分成信息不对称的两类,以纯文本链的实测答案分布为基线,选出被工具知情链过度支持的答案(Dirichlet 后验 LCB 对数似然比),并用可在标注开发集上检验的单调信息性条件界定适用范围。
Core insight: SP is effective because respondents have different private information, rather than because a model is asked to predict the crowd. This difference can be created among chains from the same model by varying whether they receive an independent execution signal: an answer endorsed more strongly by tool-informed chains than by the naive population constitutes evidence of truth that majority voting cannot see. The prediction term uses the measured class-N distribution, and the decision criterion is a likelihood-ratio test. A sufficient condition for correctness is that the execution signal provides a greater "lift" to the true answer than to any particular incorrect attractor (r(a*)>r(a)); this condition can be tested in advance on a labeled development set.
Top-k ranking stability 1.00 · BT strength 1.86 · AC weighted score 3.15/5 (borderline) · Novelty search NEAR (risk medium, evidence level corpus) · Survived 1 red-team round · direction: Heterogeneous voter ensemble methods
Artificially divide sampling chains from the same model into two classes with asymmetric information according to whether they encounter an independent execution signal. Use the measured answer distribution of text-only chains as the baseline, select the answer disproportionately supported by tool-informed chains (the Dirichlet-posterior LCB of the log likelihood ratio), and define the scope of applicability through a monotonic-informativeness condition that can be tested on a labeled development set.
Why the contribution could be memorable
研究视角:跨领域结构迁移(机制设计/众包的 SP 迁移到 LLM 采样)。评审可能记住的一句话是:"SP 在 LLM 上失败,不是因为 SP 错,而是因为样本没有信息差;信息差可以被制造并被检验。"如果再附上可复现的诊断(各领域 r 的实测值、单调性成立与失效的位置、随模型规模的变化),它是一个有观点的分析性结论,而不只是又一个投票规则。反过来,如果结果仅是"工具链更准",这个记忆点就会消失,因此必须有对 tool-only 和加权投票的明确胜出或明确的否定结果。
Pattern: cross-domain structural transfer (transferring SP from mechanism design/crowdsourcing to LLM sampling). A sentence reviewers might remember is: "SP fails on LLMs because the samples lack an information gap, rather than because SP is wrong; that gap can be created and tested." If accompanied by reproducible diagnostics—measured r values across domains, where monotonicity holds and fails, and how it changes with model scale—this would be an analytical conclusion with a clear perspective, rather than merely another voting rule. Conversely, if the result is only that "tool chains are more accurate," that memorable contribution disappears. It therefore requires a clear win over tool-only and weighted voting, or a clear negative result.
Abstract
Surprisingly-Popular(SP)类聚合要求受访者拥有不同的私有信息,而同一模型的 i.i.d. 推理链不满足这一点,因此基于一致性的无标签选择在多数答案"自信但错误"时会失效。本文提出人为制造信息不对称:对每道题同时采样 n0 条纯文本长链(class N)和 n1 条可调用 Python 沙箱的链(class I,做枚举、代入检验、小规模验证、测试执行)。我们不使用口头化元预测,而把 N 类的实测答案分布 q0 当作"预期",对每个候选答案计算 Dirichlet 后验下 log q1 − log q0 的 20% 下置信界(LCB)。仅当候选 LCB>0 且 c1≥2 时才推翻合并后的多数票。超参数(alpha、分位数、阈值)在 MATH-train 上一次性固定。同时给出单调信息性条件 r(a*)>r(a_wrong),可在标注开发集上先行检验。我们在 MATH-500 hard、AIME、OlympiadBench 数值子集、LiveCodeBench 上做等 token 比较,基线包括合并 tool-vote、tool-only、加权投票、CoT/PoT 选择、口头化 SP/ISP、DeepConf 与跨模型 jury。预期收益有限(+2–4 点),主要价值在于刻画该条件何时成立;Phase 1(10 天)的 go/no-go 不达标即终止。
Surprisingly-Popular (SP) aggregation requires respondents to have different private information, a condition not met by i.i.d. reasoning chains from the same model. Consequently, agreement-based selection without labels can fail when the majority answer is "confident but wrong." This work proposes artificially creating information asymmetry: for each problem, sample n0 long text-only chains (class N) and n1 chains that can call a Python sandbox (class I, for enumeration, substitution checks, small-scale verification, and test execution). Instead of verbalized meta-predictions, we use the measured class-N answer distribution q0 as the "expectation" and calculate, for each candidate answer, the 20% lower confidence bound (LCB) of log q1 − log q0 under a Dirichlet posterior. The combined majority vote is overridden only when the candidate has LCB>0 and c1≥2. Hyperparameters (alpha, quantile, and thresholds) are fixed once on MATH-train. We also provide a monotonic-informativeness condition, r(a*)>r(a_wrong), that can be checked beforehand on a labeled development set. Planned equal-token comparisons on MATH-500 hard, AIME, the numerical subset of OlympiadBench, and LiveCodeBench include combined tool-vote, tool-only, weighted voting, CoT/PoT selection, verbalized SP/ISP, DeepConf, and cross-model jury baselines. Expected gains are limited (+2–4 percentage points); the principal value lies in characterizing when the condition holds. The project will terminate if Phase 1 (10 days) fails its go/no-go criteria.
Motivation
多数投票在难题上会固化错误答案:样本共享同一误解时,错误答案就是共同吸引子,而同一模型的 i.i.d. 样本无法提供"错误不相同"的参照总体。跨模型 jury 可以去相关,但需要第二个模型族、成本高;"LLMs and the Madness of Crowds"也指出不同 LLM 的错误本身相关。SP/BTS 直接迁移到 LLM(口头化元预测)已被报告劣于多数投票("Bayesian Truth Serum for LLMs?":共享预训练破坏了条件独立性)。本工作检验:与其让模型"预测人群",不如真实地制造一个信息不同的总体,SP 的识别逻辑能否重新成立。这对应简报里"自信但错误的共识"这一瓶颈,并为无标签、无 PRM 的选择与自适应预算分配提供信号。
Majority voting can entrench incorrect answers on hard problems: when samples share the same misconception, the incorrect answer becomes a common attractor, and i.i.d. samples from the same model cannot provide a reference population with different errors. Cross-model juries can decorrelate errors, but require a second model family and are costly; "LLMs and the Madness of Crowds" also notes that errors across different LLMs are themselves correlated. Directly transferring SP/BTS to LLMs through verbalized meta-predictions has already been reported to perform worse than majority voting ("Bayesian Truth Serum for LLMs?": shared pretraining violates conditional independence). This work tests whether SP's identification logic can be restored by actually creating a population with different information, rather than asking a model to "predict the crowd." This addresses the brief's bottleneck of "confident but wrong consensus" and could provide a signal for selection without labels or a PRM, and for adaptive budget allocation.
The proposed gap in prior work
专家的默认做法是"让模型口头预测人群",这对可交换样本不自洽,因为元预测只是模型对自身分布的看法;另一个默认预期是"工具链更准,所以多信任工具链"。非显然之处在于既放弃口头元预测,也不把工具链当作更强的投票者,而是把两类总体的"对比"当作证据。需要坦白的是:这一点尚未被证明。红队指出,系统性的工具 bug(枚举上界过小、off-by-one、浮点误差、错误建模)会产生 c0≈0 的答案,正好抬高错误答案的 r,这恰是单调性假设最容易失败之处。对比相对"只信任工具"是否有增量,必须由实验回答。
The default expert approach is to ask the model to verbally predict the crowd. This is not self-consistent for exchangeable samples, because the meta-prediction is simply the model's view of its own distribution. Another default expectation is that "tool chains are more accurate, so they should be trusted more." The non-obvious move is both to abandon verbalized meta-predictions and to avoid treating tool chains as stronger voters, instead using the contrast between the two populations as evidence. We must be candid: this has not yet been demonstrated. The red team noted that systematic tool bugs—enumeration bounds that are too small, off-by-one errors, floating-point errors, and incorrect modeling—can produce answers with c0≈0, thereby increasing r for incorrect answers. This is precisely where the monotonicity assumption is most likely to fail. Experiments must determine whether the contrast adds value beyond simply trusting tools.
Proposed method
- 两类受访者采样:每题采样 n0 条纯文本长链(class N)与 n1 条可调用沙箱 Python 的链(class I,同一模型),缓存所有链、token 数(含解释器输出)和规范化答案;使用 vLLM 与 prefix caching。
- 答案规范化与聚类:整数/表达式做符号等价归一;代码题用测试输出签名作为答案;统计每个候选 a 的 c0(a)、c1(a)。
- 后验与得分:q_j(a) ~ Dirichlet(c_j + alpha),alpha=0.5;s(a)=E[log q1(a) − log q0(a)],取 20% 下置信界(LCB)处理稀有答案;c0+c1<3 的候选不具备翻转资格。
- 选择规则:默认取合并投票的多数票;若存在候选 a 满足 LCB s(a)>0 且 c1(a)≥2,则选 LCB 最大者。所有阈值在 MATH-train 开发集一次性固定,之后不再调整,全程不使用第二个模型族。
- 单调信息性诊断:在标注开发数据上测量 r(a*)/r(a_wrong),按类别(算术、组合、坐标几何、代码 vs 证明类、工具不透明题)报告条件成立与失效的位置,并统计 top-r 错误答案中源自工具特有 bug 的比例。
- 红队驱动的加固(新增,需消融):要求 c0(a)≥1 才可翻转,或对候选做代入交叉检验;两者均与不加约束版本对比。
- 相关错误标记与路由:合并多数票 s≤0 且 I 类与 N 类多数票不一致时标记为相关错误题,仅在这些题上追加 I 链;与简单的分歧式 adaptive-SC 分配对比,并把追加链计入等 token 账目。
- Phase 2(仅在 Phase 1 通过后):前缀级版本,在分支节点上对该前缀的续写做 I-vs-N 对比,得到节点值用于 beam 剪枝;必须在等 token 下优于答案级选择,否则放弃。
- Sampling two classes of respondents: for each problem, sample n0 long text-only chains (class N) and n1 chains that can call sandboxed Python (class I, from the same model). Cache all chains, token counts (including interpreter output), and normalized answers; use vLLM and prefix caching.
- Answer normalization and clustering: normalize integers/expressions by symbolic equivalence; use test-output signatures as answers for coding problems; count c0(a) and c1(a) for each candidate a.
- Posterior and score: q_j(a) ~ Dirichlet(c_j + alpha), alpha=0.5; s(a)=E[log q1(a) − log q0(a)]. Use the 20% lower confidence bound (LCB) to handle rare answers; candidates with c0+c1<3 are ineligible to overturn the vote.
- Selection rule: default to the majority answer in the combined vote. If a candidate a satisfies LCB s(a)>0 and c1(a)≥2, select the candidate with the largest LCB. Fix all thresholds once on the MATH-train development set and make no subsequent adjustments; no second model family is used at any point.
- Monotonic-informativeness diagnostics: measure r(a*)/r(a_wrong) on labeled development data; report where the condition holds and fails by category (arithmetic, combinatorics, coordinate geometry, code versus proof-based problems, and problems opaque to tools), and measure the proportion of top-r incorrect answers that originate from tool-specific bugs.
- Red-team-driven safeguards (new, requiring ablation): require c0(a)≥1 before an answer can overturn the vote, or cross-check candidates by substitution; compare both against the unconstrained version.
- Correlated-error flagging and routing: flag a problem as having correlated errors when the combined majority answer has s≤0 and the class-I and class-N majority answers disagree. Add further I chains only for these problems; compare against a simple disagreement-based adaptive-SC allocation, and include the additional chains in the equal-token accounting.
- Phase 2 (only after Phase 1 passes): a prefix-level version compares I-versus-N continuations of the prefix at branching nodes to obtain node values for beam pruning. It must outperform answer-level selection at equal token cost; otherwise, abandon it.
Distinction from nearest work
| paper | difference |
|---|---|
| Bayesian Truth Serum for LLMs? (2026) | 该文报告 BTS 各变体都不如多数投票,原因是共享预训练破坏条件独立性。本文把它当作动机:不再依赖口头化预测,而是制造信息不对称。但我们只拿到摘录式证据,需要通读原文以确认其变体是否已包含类似的异质受访者设计。 |
| Beyond Majority Voting: Efficient Best-Of-N with Radial Consensus Score (2026) | 该文改进多数投票对主导答案的偏置,机制上是另一种共识打分,并非构造异质总体。本文提供的是'两类总体的似然比对比'。注意:输入里把 Ai et al. 2025 的 'Beyond Majority Voting' 作为 SP/ISP 基线,而此处提供的同名论文是 Radial Consensus Score,两者是否同一工作需先核对,不要混淆引用。 |
| LLMs and the Madness of Crowds (2024) | 该文发现 LLM 的错误答案跨模型相关,聚合得不到稳健真值信号。这对本文是风险与动机:它质疑'去相关'的可能性;本文的区别是用执行信号而非更换模型来制造差异,是否真的去相关仍待实验。 |
| LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning (2026) | 该文用跨模型共识打破相关错误导致的虚假多数。本文用单模型内的工具/文本异质性,不需要第二个模型族,并把 jury 作为带标签的对比基线,不用于调参。 |
| When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2026) | 该文展示多数投票在难题上固化错误模式,是本问题设定的证据。它没有提出基于对比总体的补救;GPQA 在本文中作为工具弱信息下的预期失败检查。 |
| 整体证据说明 | 提供的五篇近邻没有覆盖 CoT-vs-PAL/PoT 混合选择(Zhao et al. 2023、XoT)与 CodeT/MBR-exec,而区域主席指出这些与本文重叠较大。这些不在给定近邻列表中,我无法核实其细节,写作前必须做文献检索;新颖性证据目前偏薄。 |
| paper | difference |
|---|---|
| Bayesian Truth Serum for LLMs? (2026) | This paper reports that all BTS variants underperform majority voting because shared pretraining violates conditional independence. The present work treats this as motivation: instead of relying on verbalized predictions, it creates information asymmetry. However, we have only excerpt-level evidence and need to read the full paper to establish whether its variants already include a similar heterogeneous-respondent design. |
| Beyond Majority Voting: Efficient Best-Of-N with Radial Consensus Score (2026) | This paper addresses majority voting's bias toward dominant answers through another consensus-scoring mechanism, rather than by constructing heterogeneous populations. The present work offers a likelihood-ratio contrast between two populations. Note: the input lists Ai et al. 2025's 'Beyond Majority Voting' as an SP/ISP baseline, whereas the paper with the same main title supplied here concerns Radial Consensus Score. Whether these are the same work must be checked first; the citations must not be conflated. |
| LLMs and the Madness of Crowds (2024) | This paper finds that incorrect LLM answers are correlated across models, preventing aggregation from yielding a robust truth signal. This is both a risk and a motivation for the present work: it questions whether decorrelation is possible. The distinction here is that differences are created using execution signals rather than a change of model; whether this actually decorrelates errors remains to be tested. |
| LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning (2026) | This paper uses cross-model consensus to break false majorities caused by correlated errors. The present work uses tool/text heterogeneity within a single model, requires no second model family, and includes the jury as a labeled comparison baseline rather than using it for hyperparameter tuning. |
| When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2026) | This paper shows that majority voting entrenches incorrect patterns on hard problems, providing evidence for the problem setting. It does not propose a remedy based on contrasting populations; GPQA is used here as an expected-failure check when tools provide weak information. |
| Overall evidence note | The five supplied neighboring papers do not cover CoT-versus-PAL/PoT mixed selection (Zhao et al. 2023, XoT) or CodeT/MBR-exec, which the area chair identified as substantially overlapping with this work. These are absent from the supplied neighboring-work list, so I cannot verify their details. A literature search is required before writing; the current novelty evidence is thin. |
Planned experiments
主模型 Qwen3-32B(单模型,无第二模型族),另用 Qwen3-8B 做规模对比(至少两个尺寸)。每题 32 条 N 链 + 32 条 I 链,沙箱 Python 执行。超参数在 MATH-train 开发集固定,测试集不调参。主指标为等总 token(含解释器输出和路由追加的 I 链)下的准确率,跨数据集与 seed 汇总并做 bootstrap 95% CI。
- datasets: MATH-500 hard split(levels 4-5); AIME 2024/2025(并入 AIME 2022-23 以增加统计功效); OlympiadBench(数值子集); GPQA-Diamond(次要;工具信号弱,作为预期失败检查); LiveCodeBench(答案=测试输出签名;I 类运行自生成测试); ProcessBench(仅 Phase 2)
- baselines: 等 token 的纯文本多数投票; 等 token 的合并 tool-integrated 多数投票(隔离对比效应的关键基线); tool-only 多数投票; 带一个调参权重的工具链加权合并投票; CoT-vs-PAL/PoT 选择与混合自一致性(Zhao et al. 2023、XoT;需复现,并先核实文献); Ai et al. 的 SP/ISP 口头化元预测聚合(重实现); 原版口头化元预测 SP(消融基线); DeepConf / self-certainty 加权投票; 跨模型 jury(第二个 32B 级模型族,等总 token); CodeT / MBR-exec / 执行聚类(仅 LiveCodeBench); Qwen2.5-Math-PRM-7B 与 Math-Shepherd 的 best-of-N / beam(Phase 2)
- metrics: 等总 token 总体准确率(主要指标); plurality-wrong 题上正确答案按 LCB s 在少数答案中的排名百分位 vs 随机(go/no-go 统计量); plurality-wrong 上翻转为正确的比例 vs plurality-right 上翻转为错误的比例; 相关错误标记的 AUROC,并与分歧式 adaptive-SC 对比; 各题型类别的 r(a*)/r(a_wrong)(识别条件有效性); top-r 错误答案中源自工具特有 bug 的比例; ProcessBench step AUROC/F1(Phase 2)
- ablations: 口头化元预测 p 替代实测 q0; 合并投票 vs 对比(隔离非可加交互); 无工具访问的 class I(同提示、不执行),证明驱动效应的是异质性而非提示; Dirichlet/LCB 开关及最小支持阈值; 不同 persona/提示的 class I vs 执行信号; 要求 c0≥1 / 代入交叉检验的有无; 第二模型 class I(jury 式),仅作带标签对比,不用于调参; 模型规模(两个尺寸); 答案级 vs 前缀级节点值
- expected: 预期适中:在 AIME/OlympiadBench/hard MATH 上等 token 比合并 tool-vote 与文本投票中较好者高 +2–4 点;在满足 r 条件的题上救回 10–20% 的 plurality-wrong 题,flip-to-wrong 低于 2%;标记 AUROC 高于 0.7。GPQA 预期无增益,作为边界条件报告。需要诚实预期:较强模型下工具链占优,对比可能退化为 tool-only 投票,增益随规模可能缩小;代码上的增益可能主要来自已知的执行聚类。
- 否证条件: 第 10 天 Phase 1 终止条件:在单模型、测试集不调参的情况下,于合并后的 plurality-wrong 题上,正确答案按 LCB s 在少数答案中的中位百分位排名不高于 65%(随机=50%),或 bootstrap CI 未排除 55%;或选择规则相对合并 tool-vote 的等 token 增益不到 1 点或 CI 不高于 0。另加区域主席提出的证伪:若 LCB 比选择在不增加 flip-to-wrong 的前提下翻转的正确题数不多于 tool-only 多数票或单参数加权合并投票,则 SP 框架无增量,终止。Phase 2 若前缀级节点值不优于答案级选择则放弃。
The primary model is Qwen3-32B (a single model, with no second model family); Qwen3-8B is also used for a scale comparison (at least two sizes). Each problem receives 32 N chains + 32 I chains, with sandboxed Python execution. Hyperparameters are fixed on the MATH-train development set, with no tuning on the test sets. The primary metric is accuracy at equal total token cost (including interpreter output and additional I chains introduced by routing), aggregated across datasets and seeds, with bootstrap 95% CIs.
- datasets: MATH-500 hard split (levels 4-5); AIME 2024/2025 (combined with AIME 2022-23 to increase statistical power); OlympiadBench (numerical subset); GPQA-Diamond (secondary; tool signals are weak, so it serves as an expected-failure check); LiveCodeBench (answer = test-output signature; class I runs self-generated tests); ProcessBench (Phase 2 only)
- baselines: equal-token text-only majority voting; equal-token combined tool-integrated majority voting (the key baseline for isolating the contrast effect); tool-only majority voting; combined weighted voting with one tuned weight for tool chains; CoT-versus-PAL/PoT selection and mixed self-consistency (Zhao et al. 2023, XoT; require reproduction and prior bibliographic verification); Ai et al.'s SP/ISP aggregation using verbalized meta-predictions (reimplemented); original SP with verbalized meta-predictions (ablation baseline); DeepConf / self-certainty weighted voting; a cross-model jury (a second model family at approximately the 32B scale, at equal total token cost); CodeT / MBR-exec / execution clustering (LiveCodeBench only); Qwen2.5-Math-PRM-7B and Math-Shepherd best-of-N / beam (Phase 2)
- metrics: overall accuracy at equal total token cost (primary metric); the correct answer's percentile rank among minority answers according to LCB s on plurality-wrong problems, versus random ranking (go/no-go statistic); the proportion of plurality-wrong problems flipped to correct versus the proportion of plurality-right problems flipped to incorrect; AUROC of correlated-error flagging, with comparison against disagreement-based adaptive-SC; r(a*)/r(a_wrong) for each problem category (validity of the identification condition); the proportion of top-r incorrect answers originating from tool-specific bugs; ProcessBench step AUROC/F1 (Phase 2)
- ablations: verbalized meta-prediction p in place of measured q0; combined voting versus contrast (to isolate the non-additive interaction); class I without tool access (the same prompt, without execution), to demonstrate that heterogeneity rather than prompting drives the effect; Dirichlet/LCB on versus off and minimum-support thresholds; class I with different personas/prompts versus execution signals; with versus without the c0≥1 requirement / substitution cross-check; class I from a second model (jury-style), used only for labeled comparison and not for tuning; model scale (two sizes); answer-level selection versus prefix-level node values
- expected: moderate expectations: at equal token cost, +2–4 percentage points on AIME/OlympiadBench/hard MATH above the better of combined tool-vote and text voting; recovery of 10–20% of plurality-wrong problems among those satisfying the r condition; flip-to-wrong below 2%; flagging AUROC above 0.7. No gain is expected on GPQA, which will be reported as a boundary condition. Expectations must remain candid: with stronger models, tool chains may dominate and the contrast may collapse to tool-only voting, with gains potentially shrinking as scale increases; code gains may primarily arise from established execution clustering.
- falsification criteria: Phase 1 termination criteria on day 10: with a single model and no test-set tuning, on problems where the combined plurality is wrong, the correct answer's median percentile rank among minority answers according to LCB s is no higher than 65% (random = 50%), or the bootstrap CI does not exclude 55%; or the selection rule's equal-token gain over combined tool-vote is less than 1 percentage point or its CI is not above 0. An additional falsification test proposed by the area chair is that if LCB-ratio selection flips no more problems to correct than tool-only majority voting or single-parameter weighted combined voting without increasing flip-to-wrong, the SP framework adds no value and the project terminates. Abandon Phase 2 if prefix-level node values do not outperform answer-level selection.
Two-week pilot plan
- 第 1–3 天:用 Qwen3-32B(及 Qwen3-8B 做廉价检查)在 MATH-500 levels 4-5、AIME 2022-25、OlympiadBench 数值子集上各采样 32 条文本链与 32 条工具链;缓存所有链、token 数(含解释器输出)与规范化答案;同时在 MATH-train 上固定 alpha、分位数、阈值。
- 第 4–5 天:构建带标签的 plurality-wrong 分层(仅用于评估);计算 r(a*) 与 r(a_wrong) 检验识别条件;统计 top-r 错误答案中工具 bug 的占比;计算正确答案按 LCB s 的排名 vs 随机。这一步就是最便宜的 go/no-go 先验。
- 第 6–8 天:用固定超参运行选择规则,与文本投票、合并 tool-vote、tool-only、单参数加权投票、重实现的 SP/ISP 在等 token 下比较;加入 c0≥1 与代入检验的消融。
- 第 9–10 天:Bootstrap CI、flip-to-wrong 分析,按 kill criterion 做 go/no-go,并写出一页结论(包括是否做 Phase 2、是否把代码移出主结论)。同时并行完成 CoT/PoT 与 CodeT 相关文献检索。
- Days 1–3: using Qwen3-32B (and Qwen3-8B for inexpensive checks), sample 32 text chains and 32 tool chains per problem on MATH-500 levels 4-5, AIME 2022-25, and the numerical subset of OlympiadBench. Cache all chains, token counts (including interpreter output), and normalized answers; simultaneously fix alpha, the quantile, and thresholds on MATH-train.
- Days 4–5: construct a labeled plurality-wrong stratum (for evaluation only); calculate r(a*) and r(a_wrong) to test the identification condition; measure the proportion of top-r incorrect answers attributable to tool bugs; calculate the correct answer's ranking by LCB s versus random ranking. This step provides the least expensive initial go/no-go check.
- Days 6–8: run the selection rule with fixed hyperparameters and compare it against text voting, combined tool-vote, tool-only, single-parameter weighted voting, and reimplemented SP/ISP at equal token cost; add ablations for c0≥1 and substitution checks.
- Days 9–10: bootstrap CIs, analyze flip-to-wrong, make the go/no-go decision according to the kill criterion, and write a one-page conclusion (including whether to undertake Phase 2 and whether to remove code from the principal conclusions). In parallel, complete the CoT/PoT and CodeT literature searches.
Risks and responses
- 系统性工具 bug(枚举上界过小、off-by-one、浮点误差、错误建模)产生 c0≈0 的错误答案,似然比反而偏向这些错误(红队指出的致命问题) → 测量 top-r 错误中的工具 bug 比例;要求 c0≥1 或代入交叉检验并做消融;若仍无法区分则如实报告为负面结果。
- 收益退化为'工具链更准',无法归因于 SP 对比 → 必须包含 tool-only、合并 tool-vote、单参数加权投票与 CoT/PoT 选择基线;以 kill criterion 中的证伪实验为准。
- 工具链与文本链共享同一误解,r≈1,没有对比信号 → 先在开发集测量 r,按题型分层;仅在条件成立的类别上声明方法适用。
- AIME 上稀有答案噪声大,go/no-go 统计量可能不达标 → 合并 AIME 2022-25 与其他数据集,使用 LCB 与 bootstrap;不达标则终止而非调参。
- 适用面窄(仅工具可检验题),随模型增强而缩小;还有数据污染(MATH、AIME 2022-24) → 至少两个模型尺寸;报告污染风险并优先使用较新的 AIME 2025;把贡献定位为条件诊断分析。
- 相关工作(CoT/PoT 混合、CodeT、SP-for-LLM)可能已覆盖该对比,且他人可能在 3 个月内发表 → 在写作前完成文献检索;尽快完成 Phase 1;代码结果不计入主结论或加入执行聚类基线。
- 算力口径:解释器输出和路由追加链被漏计会让等 token 比较不公平 → 统一记账所有 token,并与分歧式 adaptive-SC 的分配做同预算比较。
- Systematic tool bugs (enumeration bounds that are too small, off-by-one errors, floating-point errors, and incorrect modeling) produce incorrect answers with c0≈0, causing the likelihood ratio to favor those errors instead (a potentially fatal issue identified by the red team). → Measure the proportion of tool bugs among top-r errors; require c0≥1 or substitution cross-checks and conduct ablations; if they still cannot be distinguished, report the negative result faithfully.
- The gain reduces to 'tool chains are more accurate,' and cannot be attributed to the SP contrast. → Tool-only, combined tool-vote, single-parameter weighted voting, and CoT/PoT selection baselines are mandatory; use the falsification experiment in the kill criterion as the decision standard.
- Tool and text chains share the same misconception, r≈1, and there is no contrast signal. → Measure r on the development set first and stratify by problem type; claim applicability only for categories where the condition holds.
- Rare-answer noise on AIME is large, and the go/no-go statistic may fail the threshold. → Combine AIME 2022-25 with other datasets and use LCBs and bootstrap; terminate if the threshold is not met, rather than tuning parameters.
- The scope is narrow (only problems verifiable by tools) and may shrink as models improve; data contamination is also a concern (MATH, AIME 2022-24). → Use at least two model sizes; report contamination risks and prioritize the newer AIME 2025; frame the contribution as an analysis of condition diagnostics.
- Related work (CoT/PoT mixtures, CodeT, SP-for-LLM) may already cover this contrast, and others may publish within 3 months. → Complete the literature search before writing; finish Phase 1 promptly; exclude code results from the principal conclusions or include execution-clustering baselines.
- Compute accounting: omitting interpreter output and chains added by routing would make equal-token comparisons unfair. → Account uniformly for every token and compare against disagreement-based adaptive-SC allocation under the same budget.
Reviewer questions and responses
- Q: 似然比 q1/q0 对仅由工具产生的错误最大:工具代码有自己的系统性失败模式(枚举上界过小、off-by-one、浮点、误建模),文本链几乎不产生这些答案,c0≈0,少数几条 I 链就能让错误答案获得很高的 r。 A: 这是真实且可能致命的风险,我们不能声称已解决。计划是直接测量:在开发集上统计 top-r 错误答案来自工具 bug 的比例;加入 c0≥1 和代入交叉检验的约束并消融;LCB 与 c1≥2 的门槛只是部分缓解。如果测得比例高且约束后增益消失,就如实报告为负面结论。
- Q: LiveCodeBench 上 I 类机制就是基于自生成测试的执行聚类,属于 CodeT / MBR-exec / 执行投票这一成熟家族,且都不是基线,代码增益会被误记为 I-vs-N 对比的功劳。 A: 同意。要么加入 CodeT/MBR-exec/执行聚类基线,要么把代码移出主结论。我们的计划是先加基线;若对比没有超过它们,则代码只作为说明性结果,不进入核心主张。
- Q: 如果这是 A+B:单独 tool-only 投票已经拿到大部分收益,口头化 SP 又不自洽,对比的增量没有被证明。 A: 确实未证明。因此把 tool-only 与带一个调参权重的合并投票设为必须击败的基线,并以'LCB 比选择是否比它们翻转更多正确题而不增加翻转错误'作为证伪实验。如果输了,论文只剩下'SP 需要制造信息差'这一概念性观点和 r 诊断,可能不足以作为方法论文。
- Q: 与 CoT-vs-PAL/PoT 选择与混合(Zhao et al. 2023、XoT、hybrid self-consistency)重叠较大,这些没有被引用。 A: 这些文献不在我拿到的近邻列表里,我无法在此核实其细节。写作前必须检索并复现作为基线;新颖性只能收窄为'用实测朴素分布做似然比并检验可识别性条件',且必须由实验证明它超过这些方法。
- Q: +2–4 点的增益在 AIME 规模下处于 seed 噪声内,并且基准存在污染。 A: 通过跨数据集汇总、多 seed 和 bootstrap 来控制,主指标是等 token 总体准确率而不是只看 plurality-wrong 子集;污染无法完全排除,会明确披露并优先使用较新的数据。若 CI 仍包含 0,则按 kill criterion 终止。
- Q: 优势会随模型变强而消失,工具链占优,对比退化为 tool-only 投票;适用面窄。 A: 同意这可能发生。至少两种尺寸的结果会直接展示趋势;适用域限制(算术、组合、代码)和 GPQA 的预期失败会如实写明。把贡献定位为'条件何时成立'的分析,而不是通用方法。
- Q: The likelihood ratio q1/q0 is largest for errors produced only by tools: tool code has its own systematic failure modes (enumeration bounds that are too small, off-by-one errors, floating-point errors, and incorrect modeling). Text chains almost never produce these answers, so c0≈0, and just a few I chains can give an incorrect answer a very high r. A: This is a real and potentially fatal risk, and we cannot claim to have resolved it. The plan is to measure it directly: on the development set, quantify the proportion of top-r incorrect answers that originate from tool bugs; add c0≥1 and substitution cross-check constraints and ablate them. The LCB and c1≥2 thresholds offer only partial mitigation. If the measured proportion is high and the gains disappear after imposing the constraints, we will report the negative conclusion faithfully.
- Q: On LiveCodeBench, the class-I mechanism is execution clustering based on self-generated tests, belonging to the established CodeT / MBR-exec / execution-voting family. None of these is included as a baseline, so code gains could be wrongly credited to the I-versus-N contrast. A: Agreed. Either add CodeT/MBR-exec/execution-clustering baselines or remove code from the principal conclusions. Our plan is to add the baselines first; if the contrast does not outperform them, code will serve only as an illustrative result and will not enter the central claim.
- Q: If this is A+B, tool-only voting already delivers most of the gains, verbalized SP is not self-consistent, and the added value of the contrast has not been demonstrated. A: It has indeed not been demonstrated. We therefore make tool-only voting and combined voting with one tuned weight mandatory baselines to beat, and use the question of whether 'LCB-ratio selection flips more problems to correct than they do without increasing flips to incorrect' as the falsification experiment. If it loses, the paper retains only the conceptual point that 'SP requires creating an information gap' and the r diagnostics, which may be insufficient for a methods paper.
- Q: There is substantial overlap with CoT-versus-PAL/PoT selection and mixtures (Zhao et al. 2023, XoT, hybrid self-consistency), which are not cited. A: These papers are absent from the neighboring-work list I received, so I cannot verify their details here. They must be retrieved and reproduced as baselines before writing. The novelty claim can only be narrowed to 'using the measured naive distribution for a likelihood ratio and testing identifiability conditions,' and experiments must show that it outperforms those methods.
- Q: Gains of +2–4 percentage points are within seed noise at AIME's scale, and the benchmarks are contaminated. A: We will control this through aggregation across datasets, multiple seeds, and bootstrap. The primary metric is overall accuracy at equal token cost, rather than just the plurality-wrong subset. Contamination cannot be completely excluded; we will disclose it explicitly and prioritize newer data. If the CI still includes 0, we will terminate according to the kill criterion.
- Q: The advantage will disappear as models become stronger: tool chains will dominate, the contrast will collapse to tool-only voting, and the scope is narrow. A: Agreed that this may happen. Results at a minimum of two sizes will directly show the trend. We will state the domain restrictions (arithmetic, combinatorics, and code) and expected failure on GPQA faithfully. The contribution should be positioned as an analysis of 'when the condition holds,' rather than as a general method.
Conference fit
NeurIPS/ICLR/ACL 的推理方向可接受,但前提是写成"理解型"论文:核心是 SP 需要制造信息不对称、以及可检验的单调信息性条件和其在各领域/规模下的实测。若只写成选择规则,会被视为 A+B 且效应可预期。区域主席给出的综合判断为 borderline(significance 3、insight 3、generality 2),我认为这个判断合理:值得先投入 10 天的低成本 Phase 1,只有通过 kill criterion(尤其是击败 tool-only 与加权投票)后再投入数月。
The reasoning tracks of NeurIPS/ICLR/ACL could be suitable, provided the paper is written to develop understanding: its core is that SP requires creating information asymmetry, together with a testable monotonic-informativeness condition and measurements across domains and model scales. If presented only as a selection rule, it will be viewed as A+B with a predictable effect. The area chair's overall assessment was borderline (significance 3, insight 3, generality 2), which I consider reasonable. A low-cost, 10-day Phase 1 is worth undertaking first; invest several months only after passing the kill criterion, especially by beating tool-only and weighted voting.