A research proposal generated by Ariadne. The experiments below are planned, and the review scores describe this internal selection.
核心洞见: 正确的少数答案可能是 surprisingly popular:它出现的频率高于模型自己的 meta-prediction 所预期的,纯 agreement 投票看不到这一点。对单模型来说,这成立的条件是模型在“预测 naive 求解者会错在哪”上掌握的知识,多于它自己求解时避开该错误的能力(knowledge asymmetry)。这是一个可检验的经验假设,不是已被保证的定理。
本轮 top-k 排名稳定性 — · BT 强度 — · AC 加权分 2.75/5(reject)· 查新 DISTINCT(风险 low,证据等级 corpus)· 扛过红队 1 轮 · direction: Forced divergence for consensus validation
AC 指出的致命问题: The method depends on the meta-prediction carrying information that differs from the model's own sampling distribution. Surprisingly Popular theory gets that difference from respondents with heterogeneous private signals, and i.i.d. samples from one model's posterior do not have them. So s(a) mostly measures meta-prediction miscalibration and noise on rare answers. In addition, λ is tuned against a different model family, which confounds any gain with a cross-model consensus signal. The novelty claim is also undercut by existing LLM work using higher-order or inverse-SP aggregation.
AC 的改进建议:
- Run a cheap go/no-go study first. On pooled hard problems, measure whether s(correct) > s(wrong plurality) on plurality-wrong items, using a single model and no cross-model tuning. Report the correlation between meta-predicted and empirical frequencies.
- Remove the cross-model λ selection, or add an explicit cross-model jury baseline with matched tokens. Otherwise gains cannot be attributed to SP.
- Supply the missing heterogeneity in a principled way, for example with different models, different prefixes or partial information, or distinct personas, and connect it to a formal condition from SP theory under which the score is truth-tracking.
- Position the work against 'Beyond Majority Voting' (higher-order or inverse-SP aggregation) and other SP-for-LLM papers. Narrow the contribution to the prefix-node search value and the correlated-error budget router.
- Pool plurality-wrong items across datasets and seeds and report bootstrap confidence intervals. Make equal-token overall accuracy the primary metric.
把 Surprisingly Popular / Bayesian Truth Serum 的“实际频率高于预测频率”用于单个 LLM 的 rollouts 加自生成 meta-prediction,作为无标签的答案选择信号、prefix 节点价值和 correlated-error 预算路由器。核心前提(单模型 meta-prediction 携带不同于采样分布的信息)尚未验证,应先用廉价实验证伪。
Core insight: A correct minority answer may be surprisingly popular: it occurs more often than the model's own meta-prediction expects, something that pure agreement voting cannot detect. For a single model, this requires that its knowledge of "predicting where naive solvers will go wrong" exceed its ability to avoid that error when solving the problem itself (knowledge asymmetry). This is a testable empirical hypothesis, not a theorem with an established guarantee.
top-k ranking stability — · BT strength — · AC weighted score 2.75/5 (reject) · Novelty search DISTINCT (risk low, evidence level corpus) · Survived 1 red-team round · direction: Forced divergence for consensus validation
Fatal issue identified by the AC: The method depends on the meta-prediction carrying information that differs from the model's own sampling distribution. Surprisingly Popular theory gets that difference from respondents with heterogeneous private signals, and i.i.d. samples from one model's posterior do not have them. So s(a) mostly measures meta-prediction miscalibration and noise on rare answers. In addition, λ is tuned against a different model family, which confounds any gain with a cross-model consensus signal. The novelty claim is also undercut by existing LLM work using higher-order or inverse-SP aggregation.
AC recommendations for improvement:
- Run a cheap go/no-go study first. On pooled hard problems, measure whether s(correct) > s(wrong plurality) on plurality-wrong items, using a single model and no cross-model tuning. Report the correlation between meta-predicted and empirical frequencies.
- Remove the cross-model λ selection, or add an explicit cross-model jury baseline with matched tokens. Otherwise gains cannot be attributed to SP.
- Supply the missing heterogeneity in a principled way, for example with different models, different prefixes or partial information, or distinct personas, and connect it to a formal condition from SP theory under which the score is truth-tracking.
- Position the work against 'Beyond Majority Voting' (higher-order or inverse-SP aggregation) and other SP-for-LLM papers. Narrow the contribution to the prefix-node search value and the correlated-error budget router.
- Pool plurality-wrong items across datasets and seeds and report bootstrap confidence intervals. Make equal-token overall accuracy the primary metric.
Apply the Surprisingly Popular / Bayesian Truth Serum principle of "actual frequency exceeding predicted frequency" to a single LLM's rollouts plus self-generated meta-prediction, as a label-free answer-selection signal, prefix-node value, and budget router for correlated errors. The central premise (that single-model meta-prediction carries information different from the sampling distribution) has not been verified and should first be subjected to a cheap falsification experiment.
Why the contribution could be memorable
研究视角:跨领域结构迁移,从机制设计/众包的 peer prediction 迁移结构。评审会记住的是一句反共识论点:判断共识是否错误,看“实际减预测”的差距,而不是一致率。即使方法失败,一个干净的诊断结果(单模型 meta-prediction 是否携带超出样本频率的信息)也有引用价值。但这是退路,不是主要卖点,而且需要诚实面对:若结果为负,论文只能是诊断或负结果型。
Research perspective: cross-domain structural transfer, transferring a structure from peer prediction in mechanism design/crowdsourcing. What reviewers would remember is a claim that challenges consensus: to judge whether consensus is wrong, examine the "actual minus predicted" gap rather than the agreement rate. Even if the method fails, a clean diagnostic result (whether single-model meta-prediction carries information beyond sample frequencies) would have citation value. But this is a fallback, not the main selling point, and the boundary must be faced honestly: if the result is negative, the paper can only be a diagnostic or negative-results paper.
Abstract
Majority voting 和其他基于一致性的无标签价值信号,在多条推理链共享同一误解时,会自信地给出错误共识。Peer-prediction 理论(Prelec 的 Surprisingly Popular、Bayesian Truth Serum)指出,有信息的量不是原始一致率,而是实际一致率与被预测一致率之间的差距。我们提出 SP-Search。对每个问题采样 N 条链并聚类最终答案,得到经验频率 f(a)。再让同一模型在 naive/expert persona 下预测“其他独立求解者的答案分布”,平均得到 p(a)。用 Dirichlet 收缩后的 s(a)=log f(a)−log p(a) 做选择:argmax f(a)·exp(λ s(a))。同一分数可作为 prefix 分支节点的价值,用于剪枝和扩展。“f 高而 s 为负”的问题被标记为 correlated-error,额外预算路由到多样化 prefix 采样。关键风险是:单模型的 i.i.d. 样本缺少 SP 理论所需的异质私有信号,s(a) 可能只是 meta-prediction 的校准误差加噪声。因此第一阶段是不含跨模型调参的 go/no-go 诊断:在 plurality 错误的问题上,s(正确答案)>s(错误 plurality) 的比例是否显著高于随机。只有通过后才推进完整方法。
Majority voting and other agreement-based label-free value signals can confidently produce a wrong consensus when multiple reasoning chains share the same misconception. Peer-prediction theory (Prelec's Surprisingly Popular and Bayesian Truth Serum) indicates that the informative quantity is not the raw agreement rate, but the gap between actual and predicted agreement. We propose SP-Search. Sample N chains per problem and cluster their final answers to obtain empirical frequencies f(a). Then ask the same model, under naive/expert personas, to predict "other independent solvers' answer distributions," averaging the predictions to obtain p(a). Select using Dirichlet-shrunk s(a)=log f(a)−log p(a): argmax f(a)·exp(λ s(a)). The same score can serve as the value of prefix branch nodes for pruning and expansion. Problems with "high f but negative s" are flagged as correlated-error cases, and extra budget is routed to diversified prefix sampling. The key risk is that a single model's i.i.d. samples lack the heterogeneous private signals required by SP theory, so s(a) may simply be meta-prediction calibration error plus noise. The first stage is therefore a go/no-go diagnostic without cross-model tuning: on problems with a wrong plurality, is the proportion for which s(correct answer)>s(wrong plurality) significantly above chance? Proceed with the full method only if this test passes.
Motivation
PRM 需要昂贵的 step 级监督,在 OOD 下脆弱,还会被重度搜索 reward-hack。agreement 类无标签信号隐含“多数通常正确”的假设,而 closest papers 反复表明:在难题上,多数票会固化错误众数,LLM 的错误还跨模型相关。目前缺少一个无标签检验来区分“正确共识”和“相关性错误”。若这样的检验存在,可以(a)在选择阶段纠正一部分 plurality-wrong 的题,(b)作为风险信号,把预算从易题转向被误导的难题。area chair 也指出,可回收的 plurality-wrong 子集本身不大,所以实际收益上限有限。
PRMs require expensive step-level supervision, are brittle under OOD conditions, and can be reward-hacked by intensive search. Agreement-based label-free signals implicitly assume that "the majority is usually correct," whereas the closest papers repeatedly show that majority voting locks in the wrong modal answer on difficult problems and that LLM errors are correlated across models. A label-free test that distinguishes "correct consensus" from "correlated errors" is currently missing. If such a test exists, it could (a) correct some plurality-wrong problems at the selection stage and (b) serve as a risk signal that shifts budget from easy problems to difficult problems being misled by wrong consensus. The area chair also notes that the recoverable plurality-wrong subset is itself small, limiting the upper bound on practical gains.
The proposed gap in prior work
领域默认 agreement 对正确性单调递增。SP 理论来自受访者异质的人类群体,没有人把它直接用在同一模型的 rollouts 上。但专家的先验是相反的:单模型 meta-prediction 多半只是复述自己的采样分布,s(a)≈0 加噪声。所以“这能行”不显然,“这不行”也没被证明。这正是应先做廉价诊断的原因。
The field assumes by default that agreement increases monotonically with correctness. SP theory originates in human populations with heterogeneous respondents, and no one has applied it directly to rollouts from the same model. However, experts' prior belief points the other way: single-model meta-prediction is likely merely to restate its own sampling distribution, giving s(a)≈0 plus noise. Thus "this can work" is not obvious, but "this cannot work" has not been proved either. This is precisely why a cheap diagnostic should come first.
Proposed method
- 答案聚类与经验频率:每题采样 N(如 32)条链,数学用答案等价、代码用 test-output signature、GPQA 用选项,得到 f(a)。
- Meta-prediction:同一模型 K 次(如 8 次)被提示“估计其他独立求解者的答案分布”,比较无 persona、naive、expert 三种条件,平均得到 p(a)。用 prefix caching 控制成本。
- SP 打分与收缩:s(a)=log f(a)−log p(a),对稀有答案用 Dirichlet 先验收缩(这一步在 plurality-wrong 子集上可能把分数拉回多数票,需单独检验)。
- 最终选择:argmax f(a)·exp(λ s(a))。λ 必须只用单模型信号选择(例如用少量带标签的开发集,或固定默认值并报告敏感度),不使用跨模型一致性,否则会与 cross-model jury 信号混淆。
- Prefix 节点价值:在分支节点,以该 prefix 为条件采样补全得到答案分布,与以该 prefix 为条件的 meta-prediction 比较,得到 surprise 作为剪枝/扩展的价值。
- Correlated-error 检测器:标记 f 高而 s 为负的问题,并把额外预算路由到多样化 prefix 采样。
- 异质性来源(为回应 fatal issue 而新增):比较 persona、不同 prefix、部分信息(隐去部分题干)、不同模型家族作为 meta-predictor 的效果,并尽量对应到 SP 理论的形式化条件。
- Answer clustering and empirical frequencies: sample N chains per problem (for example, 32); use answer equivalence for mathematics, test-output signatures for code, and answer options for GPQA to obtain f(a).
- Meta-prediction: prompt the same model K times (for example, 8 times) to "estimate other independent solvers' answer distributions," comparing three conditions: no persona, naive, and expert. Average the predictions to obtain p(a). Use prefix caching to control cost.
- SP scoring and shrinkage: s(a)=log f(a)−log p(a), with Dirichlet-prior shrinkage for rare answers (on the plurality-wrong subset, this step may pull the scores back toward majority voting and must be tested separately).
- Final selection: argmax f(a)·exp(λ s(a)). Choose λ using single-model signals only (for example, a small labeled development set, or a fixed default with reported sensitivity), without cross-model agreement; otherwise, it becomes confounded with a cross-model jury signal.
- Prefix-node value: at a branch node, sample completions conditioned on that prefix to obtain an answer distribution, and compare it with meta-prediction conditioned on the same prefix; use the resulting surprise as a value for pruning/expansion.
- Correlated-error detector: flag problems with high f and negative s, and route extra budget to diversified prefix sampling.
- Sources of heterogeneity (added in response to the fatal issue): compare the effects of personas, different prefixes, partial information (withholding parts of the problem statement), and different model families as meta-predictors, aligning these as far as possible with the formal conditions of SP theory.
Distinction from nearest work
| paper | difference |
|---|---|
| The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus (2026) | PoLR 聚类短 prefix、取主导簇并只扩展其中的路径,本质是 prefix 层面的多数共识。本想法用实际与预测频率的差距评分,理论上可以逆转主导簇的优先级。机制相似、目的不同。如果 SP 信号无效,本想法会退化成 PoLR 类方法。 |
| LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning (2026) | 该文用跨模型共识缓解相关性错误,是本想法最强的竞争基线。本想法使用单模型自生成 meta-prediction。如果 λ 用跨模型一致性调参,就会混入 jury 信号,因此必须去掉或加入匹配 token 的 jury 基线。 |
| LLMs and the Madness of Crowds (2024) | 该文发现 LLM 的错误答案跨模型相关,聚合不能给出稳健的真值信号。这既是本想法的动机,也是风险:若 meta-prediction 与答案分布共享同一误解,SP 同样会失效。该文属于诊断,未提出基于 meta-prediction 的修正。 |
| When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals (2026) | 该文审计 agreement 与正确性的弱相关,提供动机和评估思路(高一致性自动接受规则仍会放过错误答案)。它不提供替代信号。本想法尝试提供替代信号,但证据只来自摘录,需读全文核对。 |
| When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2026) | 该文展示多数票在难题上固化错误众数,支持 plurality-wrong 子集的存在和规模。它是诊断性的。本想法针对该子集,但这也意味着可回收的题目数有限。 |
| 补充说明(证据薄弱) | 提供的 closest papers 中没有任何一篇使用 SP/BTS 打分。但 area chair 提到 'Beyond Majority Voting'(2025)等使用二阶/高阶信息或 inverse-SP 聚合的 LLM 工作,以及 SP-for-LLM-crowd 研究,这些不在提供的论文列表内,我没有核实。因此“首次使用 SP/BTS 与 LLM meta-prediction”的说法很可能不成立,应删去,并把贡献收窄为单模型 prefix 节点价值和 correlated-error 预算路由,投稿前必须检索核实。 |
| paper | difference |
|---|---|
| The Path of Least Resistance: Guiding LLM Reasoning Trajectories with Prefix Consensus (2026) | PoLR clusters short prefixes, takes the dominant cluster, and expands only paths within it; this is essentially majority consensus at the prefix level. This idea scores the gap between actual and predicted frequencies and could, in principle, reverse the dominant cluster's priority. The mechanism is similar, but the objective differs. If the SP signal is ineffective, this idea will collapse into a PoLR-like method. |
| LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning (2026) | This paper uses cross-model consensus to mitigate correlated errors and is the strongest competing baseline for this idea. This idea uses self-generated meta-prediction from a single model. Tuning λ using cross-model agreement would mix in the jury signal, so this tuning must be removed or a token-matched jury baseline added. |
| LLMs and the Madness of Crowds (2024) | This paper finds that wrong LLM answers are correlated across models and that aggregation cannot provide a robust truth signal. This is both the motivation and a risk for this idea: if meta-prediction and the answer distribution share the same misconception, SP will also fail. The paper is diagnostic and does not propose a meta-prediction-based correction. |
| When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals (2026) | This paper audits the weak correlation between agreement and correctness, providing motivation and evaluation ideas (automatic acceptance rules based on high agreement still let wrong answers through). It does not provide an alternative signal. This idea attempts to provide one, but the evidence comes only from excerpts and must be checked against the full text. |
| When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs (2026) | This paper shows that majority voting locks in the wrong modal answer on hard problems, supporting the existence and size of the plurality-wrong subset. It is diagnostic. This idea targets that subset, which also means that the number of recoverable problems is limited. |
| Additional note (weak evidence) | None of the supplied closest papers uses SP/BTS scoring. However, the area chair mentions LLM work such as 'Beyond Majority Voting' (2025), which uses second-/higher-order information or inverse-SP aggregation, as well as SP-for-LLM-crowd studies. These are absent from the supplied paper list, and I have not verified them. The claim of "the first use of SP/BTS with LLM meta-prediction" is therefore likely invalid and should be removed. Narrow the contribution to single-model prefix-node value and correlated-error budget routing, and verify the literature through a search before submission. |
Planned experiments
模型:Qwen3-8B、QwQ-32B 等开源长 CoT 模型,8×A100,vLLM。主要指标为等 token 下的总体准确率,plurality-wrong 分层只作分析。额外 meta-prediction 的 token 必须计入总预算,基线获得同等数量的额外样本。跨 dataset 和 seed 汇总 plurality-wrong 题目,报告 bootstrap 置信区间。第 0 阶段为 go/no-go:仅用单模型,不做跨模型调参。
- datasets: MATH-500(难题子集); AIME 2024/2025(plurality-wrong 题数很少,仅作汇总用); OlympiadBench; GPQA-Diamond(答案空间小,meta-prediction 可能平凡匹配 f); LiveCodeBench(答案=test-output signature); ProcessBench(step 级 AUROC/F1)
- baselines: Majority vote; DeepConf / self-certainty 加权投票; Qwen2.5-Math-PRM-7B best-of-N 与 beam search; Math-Shepherd best-of-N; 同等额外 token 预算全部用于增加样本的 self-consistency; 匹配 token 的 cross-model jury 基线(新增,回应红队与 area chair); PoLR 式 prefix consensus(新增); 仅 meta-prediction 的答案分布(B 单独)
- metrics: 等总 token 下的准确率; plurality-wrong 子集上的准确率(按 oracle 分层,汇总并带 CI); P[s(correct)>s(wrong plurality)] 及 meta-prediction 与 f 的相关性; correlated-error 标记的 AUROC; ProcessBench step AUROC/F1; reward-hacking 曲线随 N 变化; plurality-right 被翻错的比例
- ablations: 无 persona 的 meta-prediction; 来自不同模型家族的 meta-prediction; 去掉 Dirichlet 收缩; 仅答案级 vs prefix 级节点价值; 随机 persona vs expert/naive persona; λ 固定 vs 带标签开发集 vs 跨模型调参(单独标注); 不同模型规模(检验强模型下信号是否消失)
- expected: 原假设:在难题集上回收 10–25% 的 plurality-wrong 题,AIME/OlympiadBench 等 token 下总体 +2–4 点,flag AUROC>0.7。我的判断:这些数字偏乐观。area chair 指出单模型 meta-prediction 很可能只复述 f,强模型下信号更弱,AIME 子集太小。更现实的预期是信号微弱或集中在少数 knowledge-asymmetry 题上,总体增益在 0–1 点,且难以超过等 token 的额外采样。
- 否证条件: 满足任一条即终止方法路线:(1)在 plurality-wrong 子集上,SP 打分在等 token 下相对额外样本多数票的绝对提升不足 5%;(2)meta-prediction 与经验频率的相关性高于 0.95(无信息不对称);(3)汇总后 P[s(correct)>s(wrong plurality)] 的 CI 包含 0.5。若只有(3)不通过,可转为发表诊断/负结果。
Models: open-source long-CoT models such as Qwen3-8B and QwQ-32B, with 8×A100 and vLLM. The primary metric is overall accuracy at equal token budgets; plurality-wrong stratification is for analysis only. Additional meta-prediction tokens must count toward the total budget, and baselines receive the same quantity of additional samples. Pool plurality-wrong problems across datasets and seeds and report bootstrap confidence intervals. Stage 0 is a go/no-go study: use only a single model, with no cross-model tuning.
- datasets: MATH-500 (hard-problem subset); AIME 2024/2025 (very few plurality-wrong problems, used only in pooled analysis); OlympiadBench; GPQA-Diamond (small answer space, so meta-prediction may trivially match f); LiveCodeBench (answer=test-output signature); ProcessBench (step-level AUROC/F1)
- baselines: Majority vote; DeepConf / self-certainty-weighted voting; Qwen2.5-Math-PRM-7B best-of-N and beam search; Math-Shepherd best-of-N; self-consistency using the entire equivalent additional-token budget to increase the sample count; a token-matched cross-model jury baseline (added in response to the red team and area chair); PoLR-style prefix consensus (added); the answer distribution from meta-prediction alone (B alone)
- metrics: Accuracy at equal total token budgets; accuracy on the plurality-wrong subset (oracle-stratified, pooled, and with CI); P[s(correct)>s(wrong plurality)] and the correlation between meta-prediction and f; AUROC of correlated-error flags; ProcessBench step AUROC/F1; reward-hacking curves as N varies; the proportion of plurality-right problems flipped to a wrong answer
- ablations: Meta-prediction without a persona; meta-prediction from a different model family; removal of Dirichlet shrinkage; answer-level-only vs prefix-level node value; random personas vs expert/naive personas; fixed λ vs a labeled development set vs cross-model tuning (labeled separately); different model sizes (testing whether the signal disappears in stronger models)
- expected: Original hypothesis: recover 10–25% of plurality-wrong problems on hard-problem sets, achieve an overall +2–4 percentage points on AIME/OlympiadBench at equal token budgets, and obtain flag AUROC>0.7. My judgment: these numbers are optimistic. The area chair notes that single-model meta-prediction is likely merely to restate f, that the signal is weaker in stronger models, and that the AIME subset is too small. A more realistic expectation is a weak signal or one concentrated on a small number of knowledge-asymmetry problems, with an overall gain of 0–1 percentage points and difficulty outperforming additional sampling at equal token budgets.
- falsification criteria: Terminate the method direction if any of the following holds: (1) on the plurality-wrong subset, SP scoring provides less than a 5% absolute improvement over majority voting with additional samples at equal token budgets; (2) the correlation between meta-prediction and empirical frequencies exceeds 0.95 (no information asymmetry); (3) the CI for pooled P[s(correct)>s(wrong plurality)] contains 0.5. If only (3) fails, the work could shift to publication as a diagnostic/negative result.
Two-week pilot plan
- 第1–2天:为 Qwen3-8B 和 QwQ-32B 在 MATH-500 难题子集和 AIME 上各采样 32 条链并缓存,建立答案聚类/等价流程,标出 oracle plurality 对错分层。
- 第3–4天:对每题收集 8 次 meta-prediction,3 种 persona(无/naive/expert),并先检查输出格式是否可解析为分布。
- 第5天:离线计算 meta-prediction 与 f 的相关性。若>0.95,立即停止并转写为诊断结论。
- 第6–7天:在 plurality-wrong 子集上计算 P[s(correct)>s(wrong plurality)](单模型、不调 λ),并加入 OlympiadBench/GPQA 汇总,做 bootstrap CI。
- 第8–9天:加入等 token 额外样本多数票基线和匹配 token 的 cross-model jury 基线(可用已缓存的第二个模型),比较翻转率:目标是翻对≥10% plurality-wrong,翻错<2% plurality-right。
- 第10天:go/no-go 评审,若通过再设计 prefix 节点价值与路由实验。预算约 150 GPU 小时。
- Days 1–2: for Qwen3-8B and QwQ-32B, sample and cache 32 chains per problem on the MATH-500 hard-problem subset and AIME; establish the answer-clustering/equivalence pipeline and label the oracle strata for correct vs wrong plurality.
- Days 3–4: collect 8 meta-predictions per problem under 3 persona conditions (none/naive/expert), first checking whether the output format can be parsed as a distribution.
- Day 5: compute the correlation between meta-prediction and f offline. If it is >0.95, stop immediately and write up the result as a diagnostic conclusion.
- Days 6–7: compute P[s(correct)>s(wrong plurality)] on the plurality-wrong subset (single model, without tuning λ), add OlympiadBench/GPQA to the pooled analysis, and compute bootstrap CI.
- Days 8–9: add majority voting with additional samples at equal token budgets and a token-matched cross-model jury baseline (the cached second model can be used), and compare flip rates: the targets are to correct ≥10% of plurality-wrong problems and incorrectly flip <2% of plurality-right problems.
- Day 10: conduct a go/no-go review; if it passes, design prefix-node-value and routing experiments. The budget is approximately 150 GPU-hours.
Risks and responses
- 模型在 meta-prediction 中只是复述自己的答案分布,s(a)≈噪声(area chair 判定为最高风险)。 → 第一周先测相关性和 P[s(correct)>s(wrong)];用 persona、部分信息、不同 prefix 引入异质性;若无效则转为诊断论文。
- 稀有答案的 log-ratio 噪声主导,而 plurality-wrong 恰好多为稀有答案;Dirichlet 收缩又会把分数拉回多数票。 → 增大 N、对 λ 与收缩强度做敏感度分析,并报告收缩强度扫描;只对 top-k 候选簇打分。
- 收益来自跨模型信号而非 SP(λ 调参混淆)。 → 主结果不使用跨模型调参;加入 jury 基线;跨模型 meta-prediction 单独作为消融。
- 可回收子集小,AIME 上的结果在 seed 方差内。 → 跨数据集汇总,多 seed,bootstrap CI,以总体等 token 准确率为主指标。
- 强模型更校准,信号随规模消失;额外 meta-call 成本不低于额外采样。 → 多规模实验;所有 meta token 计入预算,与等量额外样本对比。
- 被抢发或与现有 LLM SP/高阶聚合工作重叠。 → 立即检索 'Beyond Majority Voting' 等工作并核实;收窄贡献到 prefix 节点价值与预算路由。
- The model merely restates its own answer distribution in meta-prediction, making s(a)≈noise (the area chair's highest-ranked risk). → Measure the correlation and P[s(correct)>s(wrong)] in the first week; introduce heterogeneity through personas, partial information, and different prefixes; if ineffective, shift to a diagnostic paper.
- Log-ratio noise for rare answers dominates, and plurality-wrong cases often involve rare answers; Dirichlet shrinkage pulls the scores back toward majority voting. → Increase N, analyze sensitivity to λ and shrinkage strength, and report a sweep over shrinkage strengths; score only the top-k candidate clusters.
- Gains come from the cross-model signal rather than SP (confounding through λ tuning). → Use no cross-model tuning in the primary results; add a jury baseline; treat cross-model meta-prediction as a separate ablation.
- The recoverable subset is small, and AIME results fall within seed variance. → Pool across datasets, use multiple seeds and bootstrap CI, and make overall equal-token accuracy the primary metric.
- Stronger models are better calibrated, so the signal disappears with scale; extra meta-calls cost at least as much as additional sampling. → Experiment across multiple model sizes; count all meta-prediction tokens in the budget and compare against an equivalent quantity of additional samples.
- Others publish first, or the work overlaps with existing LLM SP/higher-order aggregation methods. → Search and verify 'Beyond Majority Voting' and related work immediately; narrow the contribution to prefix-node value and budget routing.
Reviewer questions and responses
- Q: SP 需要具有不同私有信号的受访者,单模型样本没有这种异质性,s(a) 衡量的是 meta-prediction 的校准误差而不是真值。 A: 这一反对基本成立,是本想法的核心风险。我们不声称有理论保证。只主张:若 meta-prediction 在特定提示下携带采样器没有的“常见陷阱”知识,则存在经验信号。我们用 go/no-go 实验直接检验,并用 persona、部分信息、不同模型等方式引入异质性。无法形式化对应 SP 条件时,论文应表述为经验方法。
- Q: 稀有答案的 log-ratio 噪声主导,而 plurality-wrong 子集正是稀有答案集中的地方。 A: 同意这是一个真实问题。缓解手段是 Dirichlet 收缩、增大 N 和 K、只对候选簇打分,但收缩可能使方法退化为多数票。因此要报告收缩强度扫描,并在 plurality-wrong 子集上单独验证。若噪声不可控,这就是 kill 条件。
- Q: λ 用另一个模型家族的一致性来调,引入了跨模型共识信号,而该信号已知优于 PRM,却没有对应基线。 A: 该批评有效。我们将移除跨模型 λ 选择,主结果用固定 λ 或带标签的小开发集,并加入匹配 token 的 cross-model jury 基线;跨模型调参只作为单独标注的消融。
- Q: “首次使用 SP/BTS 与 LLM meta-prediction”的新颖性声明可能是假的,已有使用二阶预测的 LLM 聚合工作。 A: 很可能成立:area chair 提到 'Beyond Majority Voting'(2025)等。我没有在提供的 closest papers 中看到它们,也未核实全文。应删除“首次”的说法,把贡献收窄为单模型 prefix 节点价值、correlated-error 预算路由和无标签校准协议,并在投稿前做完整文献检索。
- Q: AIME 只有 30 题,plurality-wrong 子集极小,10–25% 的回收率落在 seed 方差内。 A: 同意。结果必须跨数据集汇总、多 seed 并给出 bootstrap CI,主指标为总体等 token 准确率,而不是 AIME 子集上的百分比。
- Q: 更强的模型更校准,信号会消失,且 plurality-wrong 题更少,实际影响有限。 A: 很可能如此。我们会在多个规模上测试并如实报告趋势。若信号随规模消失,贡献只限于小/中型模型和诊断价值,这需要在论文定位中明说。
- Q: SP requires respondents with different private signals. Single-model samples lack this heterogeneity, and s(a) measures meta-prediction calibration error rather than truth. A: This objection is largely valid and is the central risk of this idea. We claim no theoretical guarantee. The claim is only that an empirical signal exists if meta-prediction under particular prompts carries knowledge of "common traps" that the sampler lacks. We test this directly through the go/no-go experiment and introduce heterogeneity through personas, partial information, different models, and similar approaches. If the design cannot be formally mapped to SP conditions, the paper should present it as an empirical method.
- Q: Log-ratio noise for rare answers dominates, and the plurality-wrong subset is precisely where rare answers are concentrated. A: We agree that this is a real problem. Mitigations include Dirichlet shrinkage, increasing N and K, and scoring only candidate clusters, but shrinkage may make the method collapse into majority voting. Report a sweep over shrinkage strengths and validate separately on the plurality-wrong subset. Uncontrollable noise is a kill criterion.
- Q: λ is tuned using agreement with another model family, introducing a cross-model consensus signal that is already known to outperform PRMs, without a corresponding baseline. A: This criticism is valid. We will remove cross-model λ selection, use a fixed λ or a small labeled development set for the primary results, and add a token-matched cross-model jury baseline; cross-model tuning will appear only as a separately labeled ablation.
- Q: The novelty claim of "the first use of SP/BTS with LLM meta-prediction" may be false; LLM aggregation work using second-order predictions already exists. A: This is likely valid: the area chair mentions 'Beyond Majority Voting' (2025) and related work. I did not see these in the supplied closest papers and have not verified their full texts. Remove the "first" claim, narrow the contribution to single-model prefix-node value, correlated-error budget routing, and a label-free calibration protocol, and conduct a full literature search before submission.
- Q: AIME has only 30 problems, its plurality-wrong subset is tiny, and a 10–25% recovery rate falls within seed variance. A: Agreed. Results must be pooled across datasets, use multiple seeds, and include bootstrap CI; the primary metric is overall equal-token accuracy rather than percentages on the AIME subset.
- Q: Stronger models are better calibrated, so the signal will disappear, and they have fewer plurality-wrong problems, limiting practical impact. A: This is quite possible. We will test multiple scales and report the trends faithfully. If the signal disappears with scale, the contribution is limited to small/medium models and diagnostic value, which must be stated explicitly in the paper's positioning.
Conference fit
NeurIPS/ICLR 属于 test-time compute 与 peer prediction 的交叉,题材契合,但 area chair 给出的结论是 reject(evidence_plan 2、differentiation 2、realism 2,feasibility 4)。若 go/no-go 通过且加入 jury 基线与置信区间,可争取主会。若信号为负,更现实的是 workshop 或诊断/负结果型论文。建议两周内的 pilot 结果决定是否投入数月,实验成本低(约 150 GPU 小时,10 天),适合先试。
The intersection of test-time compute and peer prediction fits the subject matter of NeurIPS/ICLR, but the area chair's verdict is reject (evidence_plan 2, differentiation 2, realism 2, feasibility 4). If the go/no-go study passes and jury baselines and confidence intervals are added, a main-conference submission could be pursued. If the signal is negative, a workshop or diagnostic/negative-results paper is more realistic. Use the pilot results within two weeks to decide whether to invest several months. The low experiment cost (approximately 150 GPU-hours, 10 days) makes this suitable for an initial trial.