A research proposal generated by Ariadne. The experiments below are planned, and the review scores describe this internal selection.
核心洞见: 有廉价部分验证器时,额外样本只有在能揭示新行为时才有用;执行签名簇上的 Good-Turing 缺失质量可在无标签条件下估计这一概率。即停止规则应从“leader 稳不稳”(投票问题)改为“还有没有未见质量”(物种丰度问题)。
本轮 top-k 排名稳定性 1.00 · BT 强度 1.59 · AC 加权分 3.0/5(borderline)· 查新 DISTINCT(风险 low,证据等级 corpus)· 扛过红队 1 轮 · direction: Execution-based verification without learning
在只有公开样例测试这类“部分验证器”的 best-of-N 代码生成中,用执行签名簇上的 Good-Turing/Chao 未见质量(而非 leader 一致性)跨题分配采样预算;核心待证命题是:M_t 在控制了“公开测试通过状态”之后,是否仍有增量预测价值。
Core insight: When an inexpensive partial verifier is available, additional samples are useful only if they reveal new behavior; Good-Turing missing mass over execution-signature clusters can estimate this probability without labels. The stopping rule should therefore change from 'is the leader stable?' (a voting problem) to 'does unseen mass remain?' (a species-abundance problem).
Top-k ranking stability 1.00 · BT strength 1.59 · AC weighted score 3.0/5 (borderline) · Novelty search DISTINCT (risk low, evidence level corpus) · Survived 1 red-team round · direction: Execution-based verification without learning
In best-of-N code generation with only a 'partial verifier' such as public example tests, use Good-Turing/Chao unseen mass over execution-signature clusters, rather than leader agreement, to allocate the sampling budget across problems. The central proposition to establish is whether M_t retains incremental predictive value after controlling for 'public-test pass status.'
Why the contribution could be memorable
模式为 conventional_core_atypical_pairing:常规核心(best-of-N + 测试过滤 + 执行聚类)搭配非常规的物种估计工具。评审若会记住它,靠的是一个可检验的诊断结论:覆盖统计能否区分“卡在错误吸引子”与“尚未探索”,且在控制公开测试状态后仍有预测力。即使端到端提升只有约 2 个点,这个诊断也有引用价值;若不成立,则是一个清晰的负结果。目前仅是假设,并非已有结论。
The pattern is conventional_core_atypical_pairing: a conventional core (best-of-N + test filtering + execution clustering) paired with an unconventional species-estimation tool. If reviewers remember it, this will be because of a testable diagnostic conclusion: whether coverage statistics can distinguish 'being trapped in an incorrect attractor' from 'not yet explored,' while retaining predictive power after controlling for public-test status. Even if the end-to-end improvement is only approximately 2 percentage points, this diagnostic could merit citation; if it does not hold, the result is a clear negative finding. At present, this is only a hypothesis, not an established conclusion.
Abstract
在 best-of-N 加测试过滤的代码生成中,预算分配通常依赖答案一致性或难度代理,无法区分“因已解出而一致”“卡在单一错误吸引子”与“分布平坦、尚未探索”。我们把采样视为对程序行为的捕获事件:让每个样本在约 20 个由同一 LLM 生成的输入上执行,取输出签名哈希得到行为簇;用 Good-Turing 缺失质量 M_t(单例簇数/t)与 Chao1 丰度估计“下一个样本揭示新行为”的概率,再结合已见簇通过公开样例测试的比例(Beta 先验),估计“出现未见且能通过测试的簇”的概率,并贪心地把每批 8 个样本分给边际收益最大的题。对 M_t≈0 且无通过簇的“错误吸引子”题,触发强制多样化(升温或换提示)。我们刻意不预设该方法优于简单基线:将“首次通过公开测试即停止 + 对未解题 round-robin”作为必须击败的强基线,并做条件分析(只在尚无通过簇的题上比较 M_t、leader margin 与难度探针对后续增益的预测力)。评估在 LiveCodeBench 上使用 Pareto 曲线与 rank correlation;若 M_t 无增量价值则放弃。目前没有任何实验证据,以上均为待检验假设。
In code generation using best-of-N plus test filtering, budget allocation usually relies on answer agreement or difficulty proxies, which cannot distinguish 'agreement because the problem has been solved,' 'being trapped in a single incorrect attractor,' and 'a flat, insufficiently explored distribution.' We treat sampling as capture events for program behavior: execute each sample on approximately 20 inputs generated by the same LLM, then hash its output signature to obtain a behavioral cluster. Good-Turing missing mass M_t (singleton-cluster count/t) and Chao1 abundance are used to estimate the probability that 'the next sample reveals new behavior.' Combined with the proportion of observed clusters that pass public example tests (with a Beta prior), these estimate the probability that 'an unseen cluster that passes the tests will appear.' Each batch of 8 samples is then greedily allocated to the problem with the largest marginal gain. For 'incorrect attractor' problems with M_t≈0 and no passing cluster, trigger forced diversification by increasing temperature or changing the prompt. We deliberately do not assume that this method outperforms simple baselines: 'stop at the first public-test pass + round-robin sampling of unsolved problems' is a strong baseline that must be beaten. We also conduct conditional analysis, comparing the predictive power of M_t, leader margin, and difficulty probes for subsequent gains only on problems with no passing cluster yet. LiveCodeBench evaluation uses Pareto curves and rank correlation; abandon the method if M_t adds no value. There is currently no experimental evidence; all of the above are hypotheses to be tested.
Motivation
Brief 指出测试时算力均匀分配会浪费在简单题上、困难题得不到足够服务。现有自适应 self-consistency 继承了投票的“leader 是否稳定”停止规则,而在有部分验证器的代码场景中,额外样本只有在可能暴露尚未见过的行为时才有价值。代码的执行签名提供了数学答案聚类所没有的精确且廉价的语义聚类。Qwen3-32B、DeepSeek-R1 distill 配合 vLLM 使每题 64–256 样本可行,LiveCodeBench 提供时间切分的抗污染数据。但需坦白:这个动机成立的前提是 M_t 提供了公开测试状态之外的信息,这一点尚未被证明。
The brief notes that uniform allocation of test-time compute wastes resources on easy problems while hard problems receive insufficient attention. Existing adaptive self-consistency inherits the voting-based stopping rule of 'whether the leader is stable.' In coding settings with partial verifiers, however, additional samples have value only when they can expose behavior not previously seen. Code execution signatures provide exact, inexpensive semantic clustering unavailable in mathematical answer clustering. Qwen3-32B and DeepSeek-R1 distill, combined with vLLM, make 64–256 samples per problem feasible, and LiveCodeBench provides temporally split data designed to resist contamination. We must be candid, however: this motivation depends on M_t providing information beyond public-test status, which has not yet been demonstrated.
The proposed gap in prior work
领域惯例是把一致性当作停止/分配信号;把样本看作对行为的“捕获事件”并借用生态学/语言学的未见物种估计并不常见。但诚实地说,专家完全可能预期“公开测试通过状态”已解释大部分增益,且困难题上的未见质量多是垃圾单例(崩溃、超时、异常输出)。因此其非显然性尚未被证实,恰是实验要检验的内容。
The convention in the field is to use agreement as a stopping/allocation signal. Treating samples as 'capture events' for behavior and borrowing unseen-species estimators from ecology/linguistics is uncommon. Candidly, however, an expert could reasonably expect 'public-test pass status' to explain most gains, while unseen mass on difficult problems largely consists of junk singletons (crashes, timeouts, and anomalous outputs). The non-obviousness of the approach has therefore not yet been established; this is precisely what the experiments must test.
Proposed method
- 行为簇构建:每题由同一 LLM 一次性生成约 20 个输入(经 differential testing 过滤),对每个样本程序执行得到输出签名哈希;崩溃、超时、异常统一折叠为单一“无效”簇,避免垃圾单例污染(针对红队的 junk singleton 质疑),并将该选择作为敏感性变量。
- 覆盖统计:维护簇计数 n_k,t 个样本后仅在有效簇上计算 Good-Turing 缺失质量 M_t = 单例簇数 / t,并计算 Chao1 丰度估计。
- 部分验证器信号:统计已见簇中通过公开样例测试的比例,用 Beta 先验得到“未见簇可通过测试”的后验概率 P_pass。
- 分配策略:全局预算 B,每轮把下一批 8 个样本分给 M_t·P_pass 最大的题;已有通过簇且其计数超过基于 Chao 的置信下界的题被冻结。
- 最终答案:选通过公开测试且计数最高的簇;同时报告 CodeT 式双重一致性排序作对照。
- 强制多样化(默认作为独立可开关模块,主方法先不含):对 M_t≈0 且无通过簇的题切换温度或提示。为保护 exchangeability,每个采样配置各自重新计数并作为分层(strata),避免跨配置混合计数。
- 必备基线与诊断:实现“首次通过公开测试即停 + 未解题 round-robin”以及对每题公开通过率的 Thompson/UCB 分配;并做条件分析,在无通过簇的题上比较 M_t、leader margin、难度探针对后续 k 个样本真实增益的 rank correlation,再对“有通过簇但隐藏测试失败”的题(公开测试弱)单独分析。
- 离线回放框架:缓存全部样本与执行矩阵,在随机排序上重放各分配策略,使消融和多 seed 置信区间几乎零成本。
- Construct behavioral clusters: for each problem, have the same LLM generate approximately 20 inputs in one batch (filtered through differential testing). Execute each sample program to obtain an output-signature hash. Collapse crashes, timeouts, and exceptions into a single 'invalid' cluster to prevent contamination by junk singletons (addressing the red team's junk-singleton criticism), and treat this choice as a sensitivity variable.
- Coverage statistics: maintain cluster counts n_k. After t samples, compute Good-Turing missing mass M_t = singleton-cluster count / t over valid clusters only, and calculate the Chao1 abundance estimate.
- Partial-verifier signal: measure the proportion of observed clusters that pass public example tests, and use a Beta prior to obtain the posterior probability P_pass that 'an unseen cluster can pass the tests.'
- Allocation policy: under a global budget B, assign the next batch of 8 samples each round to the problem with the largest M_t·P_pass. Freeze problems that already have a passing cluster whose count exceeds a Chao-based confidence lower bound.
- Final answer: select the cluster with the highest count among those passing public tests; also report CodeT-style dual-consistency ranking as a comparison.
- Forced diversification (by default, an independent module that can be switched on or off; initially excluded from the primary method): switch temperature or prompt for problems with M_t≈0 and no passing cluster. To preserve exchangeability, restart counts separately for each sampling configuration and treat configurations as strata, avoiding pooled counts across configurations.
- Required baselines and diagnostics: implement 'stop at the first public-test pass + round-robin sampling of unsolved problems' and Thompson/UCB allocation based on each problem's public-test pass rate. Conduct conditional analysis: on problems with no passing cluster, compare the rank correlations of M_t, leader margin, and difficulty probes with the actual gain from the next k samples. Then separately analyze problems with a passing cluster that nevertheless fails hidden tests (weak public tests).
- Offline replay framework: cache all samples and the execution matrix, then replay each allocation policy over random orderings, making ablations and multi-seed confidence intervals almost cost-free.
Distinction from nearest work
| paper | difference |
|---|---|
| Semantic Voting: Execution-Grounded Consensus for LLM Code Generation (2026) | 该工作按执行行为分组候选并做共识投票,机制与本文的行为聚类相似;本文用同类聚类来估计未见质量并做跨题预算分配,而非选答案。证据仅来自摘要级引文,未核实其是否涉及预算分配,需精读。 |
| Learning to Represent Programs with Property Signatures (2020) | 该工作用在输入输出上评估性质得到程序签名用于表示学习;本文的执行签名仅用作行为等价的聚类键,不涉及学习表示。二者只在“签名”概念上相关,差异主要在目的。 |
| Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative Models (2025) | 该工作把 missing mass 用于生成模型的不确定性量化(保形预测);本文把 Good-Turing/Chao 用于测试时采样预算分配并结合部分验证器。这是最接近的“工具层”先例,本文的新意主要在应用场景与分配规则,属适度而非根本性的差异。 |
| Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (2024) | 该工作按提示难度做 compute-optimal 的逐 prompt 分配;本文的分配信号来自在线、无标签的覆盖统计,而不是预估难度。需与难度探针基线直接比较。 |
| Efficient Test-Time Scaling via Self-Calibration (2025) | 该工作用校准后的置信度决定采样量并在 Best-of-N/self-consistency 中早停;本文使用未见行为质量并显式利用部分验证器。置信度类方法是必须对比的相关基线。 |
| 总体证据说明 | 提供的 5 篇论文并未覆盖 Damani et al. 与 adaptive self-consistency 等重要相关工作,且未检索 2025–26 是否已有把 Good-Turing 用于采样预算的论文,新颖性判断证据偏薄,需要在启动前做补充文献检索。 |
| paper | difference |
|---|---|
| Semantic Voting: Execution-Grounded Consensus for LLM Code Generation (2026) | This work groups candidates by execution behavior and applies consensus voting, a mechanism similar to the present work's behavioral clustering. Here, the same type of clustering is used to estimate unseen mass and allocate budgets across problems, rather than to select answers. The evidence is only an abstract-level citation; whether that paper addresses budget allocation has not been verified and requires a close reading. |
| Learning to Represent Programs with Property Signatures (2020) | This work derives program signatures by evaluating properties over inputs and outputs for representation learning. The present work uses execution signatures only as clustering keys for behavioral equivalence and does not learn representations. The two are related only through the concept of a 'signature'; the difference is primarily their purpose. |
| Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative Models (2025) | This work uses missing mass for uncertainty quantification in generative models (conformal prediction). The present work uses Good-Turing/Chao for test-time sampling-budget allocation and combines them with a partial verifier. This is the nearest precedent at the 'tool level'; the novelty here lies mainly in the application setting and allocation rule, constituting a moderate rather than fundamental difference. |
| Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (2024) | This work makes compute-optimal allocations per prompt based on prompt difficulty. Here, the allocation signal comes from online coverage statistics without labels, rather than estimated difficulty. Direct comparison against difficulty-probe baselines is required. |
| Efficient Test-Time Scaling via Self-Calibration (2025) | This work uses calibrated confidence to determine the number of samples and stops early in Best-of-N/self-consistency. The present work uses unseen behavioral mass and explicitly incorporates a partial verifier. Confidence-based methods are relevant baselines that must be compared. |
| Overall evidence note | The 5 supplied papers do not cover important related work such as Damani et al. and adaptive self-consistency. Nor has a search established whether 2025–26 papers already use Good-Turing for sampling budgets. The evidence for the novelty judgment is thin, and a supplementary literature search is required before starting. |
Planned experiments
以 Qwen3-8B 做主实验,扩展到 Qwen3-32B 与另一模型族(如 DeepSeek-R1 distill)。每题缓存 128–256 个样本及其在生成输入与公开/隐藏测试上的执行矩阵,所有策略在缓存上离线重放,随机排序重复多次。所有方法在总生成 token 相同下比较,探针开销计入。至少 3 个采样 seed,使用 bootstrap 置信区间(LiveCodeBench 后截止切分仅数百题,2–4 点差异需谨慎)。
- datasets: LiveCodeBench(后截止切分,easy/medium/hard); HumanEval+/MBPP+(易题对照); 可选扩展:带 answer-equivalence 聚类的数学集(如 MATH-500)以检验泛化
- baselines: 每题均匀 N + 测试过滤; 首次通过公开测试即停 + 对未解题 round-robin(最强平凡基线,必须击败); 对公开通过率的 Thompson/UCB 分配; 基于解答一致性的自适应早停 self-consistency(Aggarwal et al. 风格); 逐题难度预测/探针样本的 bandit 分配(Damani et al. 风格); CodeT 式双重一致性排序(均匀 N); 基于真实通过率的 oracle 分配(上界)
- metrics: pass@1 对总生成 token 的 Pareto 曲线(含探针开销); 等精度下节省的样本比例; M_t 与真实后续增益的 Spearman 相关(分层:无通过簇 / 有通过簇但隐藏测试失败); 预算在难题与易题间的分配比例
- ablations: 用 leader margin 替换 Good-Turing; 去掉通过比例的 Beta 先验; 用代码文本/embedding 聚类替换执行签名; 生成输入数 5/20/50; 是否折叠崩溃/超时的无效簇; 去掉强制多样化,或按配置分层计数 vs 混合计数; 仅 P_pass(无 M_t)与仅 M_t(无 P_pass)的拆分,以隔离 M_t 的增量贡献
- expected: 原假设为:在同等总样本下,相对一致性自适应停止与均匀分配提升至少 3 个绝对点(LiveCodeBench hard),或在等精度下节省 ≥40% 样本;M_t 对后续 32 样本增益的 rank correlation 高于 leader margin。基于红队意见,我的保守预期是:大部分节省来自“首次通过即停”的平凡基线,M_t 的增量在“无通过簇”的中等难度题上可能较小(约 +1–2 点或 10–20% 额外节省),强推理模型下通过率趋于两极化会进一步压缩收益。以上均为预期而非结果。
- 否证条件: 若在已控制公开测试状态的条件分析中,M_t 对后续增益的 rank correlation 不高于 leader margin 加难度探针,或其 Pareto 曲线相对“首次通过即停 + round-robin”与 Thompson/UCB 基线在 3 个 seed 上未达 p<0.05 的改进,则放弃该方法;可退而把覆盖诊断作为分析型短文,但前提是诊断本身成立。
Use Qwen3-8B for the primary experiments, extending to Qwen3-32B and another model family (such as DeepSeek-R1 distill). Cache 128–256 samples per problem and their execution matrices on generated inputs and public/hidden tests. Replay all policies offline on the cache, repeating with random orderings. Compare all methods at the same total generation-token cost, including probe overhead. Use at least 3 sampling seeds and bootstrap confidence intervals (the post-cutoff LiveCodeBench split contains only a few hundred problems, so differences of 2–4 percentage points require caution).
- datasets: LiveCodeBench (post-cutoff split, easy/medium/hard); HumanEval+/MBPP+ (easy-problem controls); optional extension: mathematical datasets with answer-equivalence clustering (such as MATH-500) to test generalization
- baselines: uniform N per problem + test filtering; stop at the first public-test pass + round-robin sampling of unsolved problems (the strongest trivial baseline, which must be beaten); Thompson/UCB allocation based on public-test pass rates; adaptive early-stopping self-consistency based on solution agreement (in the style of Aggarwal et al.); bandit allocation based on per-problem difficulty prediction/probe samples (in the style of Damani et al.); CodeT-style dual-consistency ranking (uniform N); oracle allocation based on actual pass rates (upper bound)
- metrics: pass@1 versus total generation tokens as a Pareto curve (including probe overhead); proportion of samples saved at equal accuracy; Spearman correlation between M_t and actual subsequent gain (strata: no passing cluster / a passing cluster that fails hidden tests); budget allocation proportions between hard and easy problems
- ablations: replace Good-Turing with leader margin; remove the Beta prior on the passing proportion; replace execution signatures with code-text/embedding clustering; use 5/20/50 generated inputs; whether to collapse crashes/timeouts into an invalid cluster; remove forced diversification, or count separately by configuration versus pooling counts; separate P_pass-only (without M_t) and M_t-only (without P_pass) variants to isolate the incremental contribution of M_t
- expected: the original hypothesis is an improvement of at least 3 absolute percentage points (LiveCodeBench hard) over agreement-based adaptive stopping and uniform allocation at the same total sample count, or savings of ≥40% of samples at equal accuracy; M_t should have a higher rank correlation with gain from the next 32 samples than leader margin. Based on red-team feedback, my conservative expectation is that most savings will come from the trivial 'stop at the first pass' baseline. The incremental value of M_t on moderately difficult problems with 'no passing cluster' may be small (approximately +1–2 percentage points or 10–20% additional savings), and polarization of pass rates under strong reasoning models may further compress the gains. These are all expectations, not results.
- falsification criteria: abandon the method if, in conditional analysis controlling for public-test status, M_t's rank correlation with subsequent gain is no higher than that of leader margin plus difficulty probes, or if its Pareto curve does not achieve an improvement at p<0.05 over the 'stop at the first pass + round-robin' and Thompson/UCB baselines across 3 seeds. Coverage diagnostics could instead support a short analytical paper, but only if the diagnostics themselves hold.
Two-week pilot plan
- 第 1–2 天:在 200 道 LiveCodeBench 后截止题上用 vLLM 为 Qwen3-8B 每题采样 128 个解(约 25k 生成,1–2 块 A100 不到 1 天),缓存全部输出。
- 第 2–3 天:让同一 LLM 为每题生成约 20 个输入并做 differential testing 过滤;在 CPU 上沙箱执行,得到签名矩阵和公开/隐藏测试通过矩阵,标注崩溃/超时簇。
- 第 4 天:先实现最强平凡基线(首次通过即停 + round-robin)与均匀分配、Thompson/UCB,并在回放框架中验证其合理性。
- 第 5–6 天:实现 M_t、Chao1、P_pass 与贪心分配,在随机排序上回放;同时计算 leader margin 与难度探针信号。
- 第 7 天:做核心条件分析——在无通过簇的题上比较各信号对后续 k 个样本真实增益的 Spearman 相关;并检查垃圾单例比例及折叠无效簇前后的差异。
- 第 8–10 天:读出成功/失败信号(成功:M_t 的相关性比 agreement margin 高 ≥0.1,且回放 Pareto 相对强基线节省 ≥20%),同时补做 2–3 小时的文献检索,确认 2025–26 无已有 Good-Turing 采样预算论文;据此决定是否继续。
- Days 1–2: use vLLM to sample 128 solutions per problem from Qwen3-8B on 200 post-cutoff LiveCodeBench problems (approximately 25k generations, less than 1 day on 1–2 A100 GPUs), caching all outputs.
- Days 2–3: have the same LLM generate approximately 20 inputs per problem and filter them through differential testing. Execute in a sandbox on CPUs to obtain signature matrices and public/hidden-test pass matrices; label crash/timeout clusters.
- Day 4: first implement the strongest trivial baseline (stop at the first pass + round-robin), uniform allocation, and Thompson/UCB, then verify that they behave reasonably in the replay framework.
- Days 5–6: implement M_t, Chao1, P_pass, and greedy allocation; replay on random orderings while also calculating leader-margin and difficulty-probe signals.
- Day 7: perform the core conditional analysis—on problems with no passing cluster, compare each signal's Spearman correlation with the actual gain from the next k samples. Also examine the proportion of junk singletons and the difference before and after collapsing invalid clusters.
- Days 8–10: assess success/failure signals (success: M_t's correlation exceeds agreement margin by ≥0.1, and the replay Pareto comparison saves ≥20% relative to strong baselines). Also conduct a supplementary 2–3-hour literature search to confirm that no existing 2025–26 paper uses Good-Turing for sampling budgets; decide whether to continue on this basis.
Risks and responses
- M_t 没有超出公开测试状态的增量价值,平凡基线即可拿到大部分收益(审稿人最可能抓住的致命点)。 → 把平凡基线和 P_pass-only 拆分作为主要对照,先做条件分析;若不成立按 kill criterion 终止。
- 困难/无望题上的单例簇多为垃圾(崩溃、超时、边界情形差异),使 M_t 虚高并浪费预算。 → 折叠无效簇、按有效性加权单例、调整探针输入数并报告敏感性;检查生成输入对欠规约边界的拆分问题。
- 强制多样化(改温度/提示)破坏 exchangeability,使 Good-Turing 估计失效。 → 主方法不含该模块;作为独立消融,并按配置分层重新计数。
- 执行签名在模型生成输入上把正确程序拆开(欠规约边界),或公开测试太弱使“通过”无信息。 → 对比文本/embedding 聚类和不同输入数;对弱公开测试题单独分析;必要时加入隐藏测试一致性作为诊断(仅评估用)。
- 强推理模型使通过率两极化(要么解出要么无望),收益很小;且仅适用于代码,泛化受限。 → 跨模型规模检验并按通过率分层汇报;可选扩展到带答案等价聚类的数学任务;若仍小,转为分析型贡献。
- LiveCodeBench 题量小导致噪声大;相关工作拥挤存在被抢发风险;文献覆盖不全。 → 多 seed + bootstrap CI;尽早补做文献检索并快速出预印本。
- M_t has no incremental value beyond public-test status, and trivial baselines capture most of the gains (the potentially fatal issue reviewers are most likely to identify). → Use the trivial baseline and the P_pass-only variant as the main controls; perform conditional analysis first, and terminate according to the kill criterion if the premise does not hold.
- Singleton clusters on hard/hopeless problems are mostly junk (crashes, timeouts, and differences on boundary cases), inflating M_t and wasting the budget. → Collapse invalid clusters, weight singletons by validity, adjust the number of probe inputs, and report sensitivity; check whether generated inputs split programs on underspecified boundaries.
- Forced diversification (changing temperature/prompt) violates exchangeability, invalidating the Good-Turing estimate. → Exclude this module from the primary method; include it as an independent ablation and restart counts separately by configuration.
- Execution signatures on model-generated inputs split correct programs (underspecified boundaries), or public tests are so weak that 'passing' is uninformative. → Compare text/embedding clustering and different input counts; analyze weak-public-test problems separately; if necessary, add hidden-test consistency as a diagnostic (for evaluation only).
- Strong reasoning models polarize pass rates (either solved or hopeless), leaving little gain; applicability is limited to code, restricting generalization. → Test across model scales and report by pass-rate strata; optionally extend to mathematical tasks with answer-equivalence clustering; if gains remain small, reposition the contribution as analytical.
- LiveCodeBench's small problem count produces substantial noise; related work is crowded, creating a risk of being scooped; the literature coverage is incomplete. → Use multiple seeds + bootstrap CIs; conduct the supplementary literature search early and release a preprint promptly.
Reviewer questions and responses
- Q: “首次通过公开测试即停 + 对未解题 round-robin”这种平凡规则已完成大部分重分配,你的方法增益可能只是来自使用了验证器。 A: 这一点成立,且原计划确实缺该基线。修订后它是主对照,同时加入 Thompson/UCB 与 P_pass-only 消融来隔离 M_t 的增量。若 M_t 在无通过簇题上的预测力不超过 round-robin 或难度探针,则按 kill criterion 放弃。我不预设会赢。
- Q: 输出签名簇上的缺失质量被垃圾单例主导,在困难或无望题上会错误地提示“还需要预算”。 A: 这是真实风险。缓解为折叠崩溃/超时/异常、按有效性加权单例,并报告对输入数和折叠选择的敏感性,同时直接统计单例中无效输出的占比。但对‘语义上错误但有效’的多样输出,这些措施无法完全消除问题,需要实验判断。
- Q: 强制多样化切换破坏 Good-Turing 所需的 exchangeability,方法不自洽。 A: 同意。主方法不含该切换;作为单独消融时按采样配置分层并重新计数,把各配置视为不同‘群落’。若分层后仍无收益,则删去该模块,不影响核心估计器的主张。
- Q: 适用面窄(仅代码、需可执行输出),与 brief 的数学/PRM/步级引导关系较弱。 A: 确实只覆盖 brief 中‘自适应预算’一环,没有步级信号。可选的数学答案等价聚类扩展能部分缓解,但数学聚类较粗,效果不保证;应把论文定位为部分验证器场景下的分配与诊断研究,而非通用 PRM 替代。
- Q: 新颖性有限:执行聚类、missing mass 和自适应分配都各有先例,只是拼接。 A: 同意差异是适度的,且我提供的相关论文覆盖不全,尚未核实 2025–26 的最新工作。价值主要取决于条件诊断(能否区分‘卡住’与‘未探索’)是否成立;若否,不建议作为主会方法论文投出。
- Q: A trivial rule such as 'stop at the first public-test pass + round-robin sampling of unsolved problems' already performs most of the reallocation. Your gains may come solely from using the verifier. A: This is valid, and the original plan did indeed omit that baseline. In the revised plan it is the main control, alongside Thompson/UCB and a P_pass-only ablation to isolate M_t's incremental contribution. If M_t's predictive power on problems with no passing cluster does not exceed round-robin or difficulty probes, the method will be abandoned according to the kill criterion. I do not assume it will win.
- Q: Missing mass over output-signature clusters is dominated by junk singletons and will incorrectly indicate that hard or hopeless problems 'still need budget.' A: This is a real risk. Mitigations include collapsing crashes/timeouts/exceptions, weighting singletons by validity, reporting sensitivity to input counts and the decision to collapse clusters, and directly measuring the proportion of invalid outputs among singletons. However, these measures cannot fully eliminate the problem of diverse outputs that are 'semantically incorrect but valid'; experiments must determine the outcome.
- Q: Forced diversification violates the exchangeability required by Good-Turing, making the method inconsistent. A: Agreed. The primary method excludes this switch; when it is used as a separate ablation, counts are restarted and stratified by sampling configuration, treating each configuration as a different 'community.' If stratification still yields no gains, remove the module; this does not affect the claim about the core estimator.
- Q: The scope is narrow (code only, requiring executable outputs), and the connection to the brief's mathematics/PRM/step-level guidance is weak. A: It indeed covers only the brief's 'adaptive budget' component and provides no step-level signal. The optional extension to mathematical answer-equivalence clustering can partly address this, but mathematical clustering is coarser and effectiveness is not guaranteed. The paper should be positioned as a study of allocation and diagnostics in partial-verifier settings, rather than a general replacement for PRMs.
- Q: Novelty is limited: execution clustering, missing mass, and adaptive allocation each have precedents, and this merely combines them. A: Agreed that the distinction is moderate. The related papers I supplied do not provide full coverage, and the latest 2025–26 work has not yet been verified. The value depends primarily on whether the conditional diagnostic can distinguish 'stuck' from 'unexplored.' If it cannot, I would not recommend submitting it as a main-conference methods paper.
Conference fit
适合 NeurIPS/ICLR 的测试时算力方向,但前提是论文以有原则的分析为核心(覆盖统计作为诊断与分配信号,并含强基线与条件分析)。若 M_t 增量只有 1–2 点,更适合作为 workshop 或分析型论文;对 ACL 契合度较弱。综合评估为 borderline:可行性高(离线回放、约 8 天可得出 go/no-go),但意义与泛化性中等偏低,建议先做两周 pilot 再决定是否投入数月。
The test-time-compute tracks of NeurIPS/ICLR are suitable, provided the paper centers on principled analysis: coverage statistics as diagnostic and allocation signals, with strong baselines and conditional analysis. If M_t adds only 1–2 percentage points, a workshop or analytical paper would be more appropriate; the fit with ACL is weaker. The overall assessment is borderline: feasibility is high (offline replay and a go/no-go decision in approximately 8 days), but significance and generalizability are moderately low. A two-week pilot is recommended before deciding whether to invest several months.