← Research record
I039 / ARIADNE 7.5Not retained

Guidance Headroom: Measure the Room for Step Guidance Before Investing in Reward Models

Observed search gains combine policy headroom, scorer quality and task difficulty, obscuring why reward models fail.

01Nested rollouts嵌套推理轨迹02Guidance headroom引导潜力上界03Resolution over time随时间解析
Conceptual research hypothesis · No experimental result is depicted.

A research proposal generated by Ariadne. The experiments below are planned, and the review scores describe this internal selection.

核心洞见: step-level guidance 能否起作用,首先是 policy 的性质:取决于各深度处不同前缀之间续写成功率的方差。任何只依赖前缀的 scorer 至多解释这部分方差,因此 oracle MC 价值函数引导的搜索是 value-guided search 的参考上界,而且这个上界不需要 PRM 就能计算。

本轮 top-k 排名稳定性 — · BT 强度 — · AC 加权分 2.85/5(borderline)· 查新 DISTINCT(风险 low,证据等级 corpus)· 扛过红队 1 轮 · direction: Guidance headroom and search capability

AC 指出的致命问题: The main hypothesis test is close to circular. Oracle headroom H is the gain from selecting prefixes by MC p_d, and V(d) is the between-prefix variance of that same p_d. V 'explaining' H beyond p(1-p) is therefore largely implied by construction, and pre-registered kill criterion (a) can barely fail. The non-tautological claims are (i) that cheap headroom estimates predict realised gain from PRMs or cheap signals on held-out models, and (ii) that RL training causally shifts resolution later at matched difficulty. Neither is the primary outcome, and the 'Gain = H × capture' factorisation has no derivation.

AC 的改进建议:

  • Make the primary outcome non-circular: use the cheap estimate (R=4, m=4, one rollout set) to predict realised guidance gain (tuned PRM beam search, and agreement/consistency-based label-free guidance) on held-out models and problems, with leave-one-model-out evaluation and independent rollouts for predictor and outcome.
  • Replace 'Gain_PRM = H × capture' with an explicit regression (gain ~ H + capture + H:capture), or derive a bound linking selection gain to AUROC under a stated model, such as a Gaussian latent value. Do not present it as an identity.
  • Specify a nested, survivors-only rollout tree for the multi-stage H (branch only from the top-k at each depth). Give a token budget per model and dataset, and cut the grid to about 6 models × 3 datasets with problem subsets so it fits 8×A100.
  • Run the RL vs SFT vs base atlas on pass-rate-matched problem subsets (so p(1-p) is equal across variants), using same-base model pairs (e.g., Qwen2.5-Math base → SFT-distill → RL). That separates late resolution from difficulty and length.
  • Define depth for long CoT without peeking at final length (e.g., absolute token offsets or step counts), and report how sensitive the resolution profiles are to segmentation.

用嵌套 Monte Carlo rollout 测量 policy 自身在中间前缀处的结果可判定性(between-prefix variance 与 oracle 多阶段搜索增益),不依赖任何 PRM 就给出 step-level guidance 的参考上界,并检验这个廉价估计能否在 held-out 模型上预测真实 PRM / beam search 的增益。

Core insight: Whether step-level guidance can work is first a property of the policy: it depends on the variance in continuation success rates between different prefixes at each depth. Any scorer that depends only on the prefix can explain at most this component of the variance. Search guided by an oracle MC value function therefore provides a reference upper bound for value-guided search, and this bound can be computed without a PRM.

top-k ranking stability — · BT strength — · AC weighted score 2.85/5 (borderline) · Novelty search DISTINCT (risk low, evidence level corpus) · Survived 1 red-team round · direction: Guidance headroom and search capability

Fatal issue identified by the AC: The main hypothesis test is close to circular. Oracle headroom H is the gain from selecting prefixes by MC p_d, and V(d) is the between-prefix variance of that same p_d. V 'explaining' H beyond p(1-p) is therefore largely implied by construction, and pre-registered kill criterion (a) can barely fail. The non-tautological claims are (i) that cheap headroom estimates predict realised gain from PRMs or cheap signals on held-out models, and (ii) that RL training causally shifts resolution later at matched difficulty. Neither is the primary outcome, and the 'Gain = H × capture' factorisation has no derivation.

AC recommendations for improvement:

  • Make the primary outcome non-circular: use the cheap estimate (R=4, m=4, one rollout set) to predict realised guidance gain (tuned PRM beam search, and agreement/consistency-based label-free guidance) on held-out models and problems, with leave-one-model-out evaluation and independent rollouts for predictor and outcome.
  • Replace 'Gain_PRM = H × capture' with an explicit regression (gain ~ H + capture + H:capture), or derive a bound linking selection gain to AUROC under a stated model, such as a Gaussian latent value. Do not present it as an identity.
  • Specify a nested, survivors-only rollout tree for the multi-stage H (branch only from the top-k at each depth). Give a token budget per model and dataset, and cut the grid to about 6 models × 3 datasets with problem subsets so it fits 8×A100.
  • Run the RL vs SFT vs base atlas on pass-rate-matched problem subsets (so p(1-p) is equal across variants), using same-base model pairs (e.g., Qwen2.5-Math base → SFT-distill → RL). That separates late resolution from difficulty and length.
  • Define depth for long CoT without peeking at final length (e.g., absolute token offsets or step counts), and report how sensitive the resolution profiles are to segmentation.

Use nested Monte Carlo rollouts to measure how distinguishable outcomes are at intermediate prefixes of the policy itself (between-prefix variance and oracle multi-stage search gain), provide a reference upper bound for step-level guidance without relying on any PRM, and test whether this cheap estimate can predict actual PRM / beam search gains on held-out models.

Why the contribution could be memorable

研究视角:新测量工具。评审会记住的是一个具体、可复用的东西:一个 30 题、R=m=4 的"PRM 投资前体检",以及一张展示同基座模型经 SFT-distill → RL 后结果可判定点向后移动的图。前提是它在 held-out 模型上预测了真实 PRM 增益。如果只展示 V 能解释 H,评审会视之为恒等式,不会记住。

Research perspective: new measurement instrument. What reviewers would remember is something concrete and reusable: a 30-problem, R=m=4 "checkup before investing in a PRM," and a plot showing that the point at which outcomes become distinguishable moves later when a model with the same base undergoes SFT-distill → RL. The prerequisite is that it predicts actual PRM gains on held-out models. If the work only shows that V can explain H, reviewers will regard it as an identity and will not remember it.

Abstract

团队在训练 PRM 或做 step-level tree search 之前,通常不知道某个 model×benchmark 组合的结果是否在中间步骤就已可判定。已报告的 PRM 增益混杂了三个因素:policy headroom、scorer 质量和题目难度。本文提出一个不依赖 scorer 的测量工具:每题在深度 d 抽取 R 个前缀根,每个根做 m 条续写,用 benchmark checker 打分,并把续写拆成估计半与 held-out 评估半以消除 winner's curse。我们用单因素 ANOVA 去除根内二项噪声,得到 between-root 方差分量 sigma_d^2 及归一化形式 V(d);给出仅限单阶段剪枝的 oracle 增益上界(基于 Aven 的期望最大值界),并在缓存的嵌套树上定义多阶段 oracle headroom H。核心检验不是 V 能否"解释"H(这近乎恒等式),而是廉价估计(R=m=4,约 30 题)能否在 leave-one-model-out、使用独立 rollout 的条件下,预测真实 PRM-guided beam search 与 best-of-N 的增益。我们还在同基座、通过率匹配的题目子集上,构建 base / SFT-distill / RL / instruct 的结果解析时机图谱。本文不提出新的搜索算法。

Before training a PRM or performing step-level tree search, teams usually do not know whether outcomes for a particular model×benchmark combination are already distinguishable at intermediate steps. Reported PRM gains confound three factors: policy headroom, scorer quality, and problem difficulty. This proposal introduces a scorer-independent measurement tool: for each problem, sample R prefix roots at depth d, generate m continuations from each root, score them with the benchmark checker, and split the continuations into an estimation half and a held-out evaluation half to remove the winner's curse. We use one-way ANOVA to remove within-root binomial noise, obtaining the between-root variance component sigma_d^2 and its normalized form V(d); provide an upper bound on oracle gain restricted to single-stage pruning (based on Aven's expected-maximum bound); and define multi-stage oracle headroom H on a cached nested tree. The central test is not whether V can "explain" H (which is close to an identity), but whether a cheap estimate (R=m=4, about 30 problems) can predict actual gains from PRM-guided beam search and best-of-N under leave-one-model-out evaluation using independent rollouts. We also construct an atlas of outcome-resolution timing for base / SFT-distill / RL / instruct variants sharing the same base, on pass-rate-matched problem subsets. This proposal introduces no new search algorithm.

Motivation

PRM 引导的 best-of-N 与 beam search 在不同模型上的收益差异很大,在长 CoT 的 RL 模型上常常收益很小。现在无法区分:是 PRM 质量差(capture 低),是 policy 在早期前缀处本来就没有可利用的信号(headroom 低),还是难度所致。训练 PRM 成本高(step 标签或 MC rollout),所以需要一个事前的 go/no-go 判断。在 8×A100、4 个月、2 名博士生的预算下,这个测量必须便宜,而且必须能被独立的真实增益检验,而不是自洽的定义。

The gains from PRM-guided best-of-N and beam search differ substantially between models, and are often very small for RL models with long CoT. It is currently impossible to distinguish whether this reflects poor PRM quality (low capture), a policy that inherently offers no exploitable signal at early prefixes (low headroom), or problem difficulty. Training a PRM is expensive (requiring step labels or MC rollouts), so an advance go/no-go judgment is needed. With a budget of 8×A100, 4 months, and 2 PhD students, this measurement must be cheap and must be testable against independent actual gains, rather than merely being a self-consistent definition.

The proposed gap in prior work

常规做法是用 ProcessBench AUROC 或搜索增益比较 PRM,两者都把 scorer 和 policy 缠在一起。Math-Shepherd / OmegaPRM 使用同样的 MC rollout,但只把它当训练标签。把它重新解释为 policy 层面的 headroom,需要一个观念转变:PRM 至多等于 oracle 价值函数。需要坦率承认:"RL 长 CoT 模型晚解析"这个方向本身很多专家已经会预期,所以真正的新意只能来自精确的量化、有明确范围的上界,以及对真实增益的非循环预测,而不是方向本身。

The usual practice is to compare PRMs using ProcessBench AUROC or search gains; both entangle the scorer and the policy. Math-Shepherd / OmegaPRM use the same MC rollouts, but only as training labels. Reinterpreting them as policy-level headroom requires a conceptual shift: a PRM can at best match the oracle value function. It must be acknowledged candidly that many experts would already expect "late resolution in RL models with long CoT." The real novelty can therefore come only from precise quantification, an upper bound with an explicit scope, and non-circular prediction of actual gains, rather than from the direction of the result itself.

Proposed method

  1. 嵌套 rollout 树与深度定义:每题在 4 个深度抽取 R 个前缀根(完整版 R=8,廉价版 R=4),每根 m 条续写(8 或 4),用 checker 打分。深度不能依赖未知的最终长度,主设定使用绝对 token 偏移或步数网格,百分比深度只作为在完成的参考链上事后归一化的对照。步骤切分对比 newline、double-newline 与 token 固定窗口,并报告 profile 对切分的敏感度。代码与 GPQA 的边界单独定义(代码按语句/函数块,GPQA 按段落),这部分目前规格不足,要在 pilot 中先确定。
  2. ANOVA 修正的 resolution profile:续写拆成估计半与评估半;用单因素 ANOVA 去除根内二项方差,得到无偏的 between-root 方差分量 sigma_d^2,以及归一化的 V(d)=sigma_d^2/(p(1-p)),并与 naive 估计器对比。注意 m 拆半后每半仅 4 条,单根估计很噪,所以主要在题目/cell 层面汇总,不依赖单根值。
  3. 单阶段剪枝上界定理:n 个前缀、oracle 保留 top-k,在等 token 下相对无引导采样的精度增益至多为 sigma_d*(n-1)/sqrt(2n-1)(k=1,Aven 界),一般 k 有类似形式。只对单阶段剪枝成立;完整解的 best-of-N 重排不在 V(d) 覆盖范围内。需要在论文中写清楚假设(续写条件 i.i.d.、等 token 的换算),并报告该界与实测 oracle 增益的比值。
  4. 多阶段 oracle headroom H:在 survivors-only 的嵌套树上定义。每个深度只从上一层的 top-k 存活前缀继续分支,避免指数增长,并给出每个模型×数据集的 token 预算。H 等于 oracle 引导的 successive halving / beam 搜索在固定预算下的精度,减去等 token 的 majority vote 或通过率基线,用 held-out 半评估选中的叶子。H 只是有限树上的经验参考值,不是对所有搜索算法的上界;最终 verifier 的 pass@N headroom 单独报告,以分离早期引导部分。
  5. 非循环的主检验:用廉价估计(R=m=4,一套 rollout)作为预测变量,用独立采样的 rollout 做结果变量,预测真实的引导增益(调优后的 Qwen2.5-Math-PRM / Math-Shepherd 的 best-of-N 与 beam search,以及一个 label-free 的一致性引导基线),在 held-out 模型和题目上做 leave-one-model-out 评估。V 对 H 的混合效应回归降级为一致性检查,不再作为主假设。
  6. PRM 增益的回归式分解:放弃'Gain_PRM = H × capture'的恒等式表述。改为显式回归 gain ~ H + capture + H:capture,其中 capture 是逐题的 PRM 前缀分数与 held-out MC p_d 的组内秩相关(Kendall tau / AUROC)。可选地在 Gaussian latent value 的假设模型下推导选择增益与 AUROC 的关系,作为补充而非主张;并在同一递归搜索中把 oracle-MC 换成 PRM 分数,隔离 capture 的作用。
  7. RL vs SFT vs base 图谱:使用同基座模型对(如 Qwen2.5 base → SFT-distill → RL → instruct 变体),并在通过率匹配的题目子集上比较,使 p(1-p) 在各变体间相等,以分离'晚解析'与难度和链长。cell 层面不少于约 30 个,8–10 个模型,1.5B–32B,32B 最多 2 个数据集各 60 题。
  8. 廉价 go/no-go 版本:评估 R=m=4、30 题的估计是否足以对模型排序、并与完整估计和真实增益一致,给出稳定性与保真度曲线。
  1. Nested rollout tree and depth definition: for each problem, sample R prefix roots at 4 depths (R=8 for the full version, R=4 for the cheap version), with m continuations per root (8 or 4), scored by the checker. Depth must not depend on an unknown final length. The primary setting uses absolute token offsets or a step-count grid; percentage depths serve only as a comparison normalized retrospectively on completed reference chains. Compare newline, double-newline, and fixed token windows for step segmentation, and report the profile's sensitivity to segmentation. Define boundaries separately for code and GPQA (statements/function blocks for code, paragraphs for GPQA). This component is currently underspecified and must first be settled in the pilot.
  2. ANOVA-corrected resolution profile: split continuations into estimation and evaluation halves; use one-way ANOVA to remove within-root binomial variance, obtaining an unbiased between-root variance component sigma_d^2 and the normalized V(d)=sigma_d^2/(p(1-p)); compare these with the naive estimator. Note that after splitting m in half, each half contains only 4 continuations, making individual-root estimates very noisy. Aggregate mainly at the problem/cell level rather than relying on individual-root values.
  3. Upper-bound theorem for single-stage pruning: for n prefixes with the oracle retaining the top-k, the accuracy gain over unguided sampling at equal token budgets is at most sigma_d*(n-1)/sqrt(2n-1) (k=1, Aven's bound), with a similar form for general k. This holds only for single-stage pruning; best-of-N reranking of complete solutions is outside the scope of V(d). State the assumptions explicitly in the paper (conditionally i.i.d. continuations and the conversion to equal token budgets), and report the ratio of the bound to the measured oracle gain.
  4. Multi-stage oracle headroom H: define it on a survivors-only nested tree. At each depth, branch only from the preceding layer's top-k surviving prefixes to avoid exponential growth, and specify a token budget for each model×dataset. H is the accuracy of oracle-guided successive halving / beam search under a fixed budget minus the equal-token majority-vote or pass-rate baseline, with selected leaves evaluated using the held-out half. H is only an empirical reference value on a finite tree, not an upper bound for all search algorithms; report the final verifier's pass@N headroom separately to distinguish the early-guidance component.
  5. Non-circular primary test: use the cheap estimate (R=m=4, one rollout set) as the predictor, and independently sampled rollouts for the outcome, to predict actual guidance gains (tuned Qwen2.5-Math-PRM / Math-Shepherd best-of-N and beam search, plus a label-free consistency-guided baseline). Conduct leave-one-model-out evaluation on held-out models and problems. Downgrade the mixed-effects regression of H on V to a consistency check; it is no longer the primary hypothesis.
  6. Regression-based decomposition of PRM gain: abandon the presentation of 'Gain_PRM = H × capture' as an identity. Replace it with the explicit regression gain ~ H + capture + H:capture, where capture is the within-group rank correlation, computed per problem, between PRM prefix scores and held-out MC p_d (Kendall tau / AUROC). Optionally derive the relationship between selection gain and AUROC under an assumed Gaussian latent-value model, as a supplement rather than a central claim; also replace oracle-MC with PRM scores within the same recursive search to isolate the role of capture.
  7. RL vs SFT vs base atlas: use model pairs sharing the same base (for example, Qwen2.5 base → SFT-distill → RL → instruct variants), and compare them on pass-rate-matched problem subsets so that p(1-p) is equal across variants, separating 'late resolution' from difficulty and chain length. Include no fewer than approximately 30 cells, 8–10 models, and sizes of 1.5B–32B; for 32B, use at most 2 datasets with 60 problems each.
  8. Cheap go/no-go version: evaluate whether estimates using R=m=4 and 30 problems suffice to rank models and agree with full estimates and actual gains; provide stability and fidelity curves.

Distinction from nearest work

paper difference
The Critical Horizon: Inspection Design Principles for Multi-Stage Operations and Deep Reasoning (2026) 该文同样用对状态的 W 条独立续写的经验成功率作为 V(s) 的无偏估计(机制相似)。本文区别在于:用 held-out 半 + ANOVA 修正估计 between-prefix 方差,并把它变成 policy 层面的 headroom 与上界,再与真实 PRM 增益做对照。但框架相近,我无法仅凭所给证据(只有一句摘录)判断其理论结果的重合程度,需要通读后再确认差异是否足够。
Process Supervision for Chain-of-Thought Reasoning via Monte Carlo Net Information Gain (2026) 该文用 MC rollout 上的信息论量自动生成 PRM 训练标签(目的是训练 PRM)。本文不训练 scorer,而是把 rollout 当作 policy 层面的 scorer-free headroom 测量。这与 Math-Shepherd/OmegaPRM 的差异同类,属于'同一量、不同用途'。
Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning (2026) 该文建立共享前缀的树,并声称方差降低等于前缀可解释的回报方差,与本文的方差分解思路在数学上相近;但它面向 RL 训练中的 advantage 估计,领域为 agent。本文是用于推理时搜索的诊断,不涉及训练。
Enhancing LLM Reasoning with Reward-guided Tree Search (2024) 该文是基于 reward model 的树搜索方法(机制不同)。本文是它所依赖的引导信号是否有可利用空间的前置测量,而不是新搜索方法;可作为被诊断的真实搜索基线之一。所给证据仅有标题,其细节没有被我核实。
Dynamic and Generalizable Process Reward Modeling (2025) 该文改进 PRM 本身的泛化性。本文与之互补:在不涉及 PRM 的情况下衡量 policy 的 headroom,并把 PRM 增益拆成 headroom 与 capture。所给证据仅有标题,细节未核实。
总体证据说明 提供的 5 篇近邻论文都没有直接做 scorer-free 的 policy 级 headroom 加对真实 PRM 增益的预测检验,但证据偏薄(只有 3 条摘录);对 early-answer probing、point-of-no-return、ICC 分析等相关文献,仅在 closest_known 中提及,未经检索核实,不能据此主张其为空白。
paper difference
The Critical Horizon: Inspection Design Principles for Multi-Stage Operations and Deep Reasoning (2026) This paper likewise uses the empirical success rate of W independent continuations from a state as an unbiased estimate of V(s) (a similar mechanism). The distinction here is the use of a held-out half + ANOVA correction to estimate between-prefix variance, turn it into policy-level headroom and an upper bound, and then compare it with actual PRM gains. However, the frameworks are close. I cannot determine the extent of overlap in their theoretical results from the supplied evidence alone (only one excerpted sentence); the full paper must be read before confirming whether the distinction is sufficient.
Process Supervision for Chain-of-Thought Reasoning via Monte Carlo Net Information Gain (2026) This paper uses information-theoretic quantities from MC rollouts to generate PRM training labels automatically (with the aim of training a PRM). This proposal does not train a scorer, instead treating rollouts as a scorer-free measurement of policy-level headroom. This is the same type of distinction as that from Math-Shepherd/OmegaPRM: 'the same quantity, a different use.'
Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning (2026) This paper builds a tree with shared prefixes and claims that the variance reduction equals the return variance explained by the prefix, which is mathematically close to this proposal's variance-decomposition idea. However, it addresses advantage estimation in RL training, in the agent domain. This proposal is a diagnostic for inference-time search and does not involve training.
Enhancing LLM Reasoning with Reward-guided Tree Search (2024) This paper proposes a reward-model-based tree-search method (a different mechanism). This proposal provides an advance measurement of whether the guidance signal on which such a method relies offers exploitable headroom, rather than a new search method; the paper's method could serve as one of the actual search baselines being diagnosed. The supplied evidence includes only the title; I have not verified its details.
Dynamic and Generalizable Process Reward Modeling (2025) This paper improves the generalization of the PRM itself. This proposal is complementary: it measures policy headroom without involving a PRM and decomposes PRM gains into headroom and capture. The supplied evidence includes only the title; the details have not been verified.
Overall evidence note None of the 5 supplied neighboring papers directly combines scorer-free policy-level headroom with a test of its ability to predict actual PRM gains. However, the evidence is thin (only 3 excerpts). Related literature on early-answer probing, points of no return, ICC analysis, and similar topics is mentioned only in closest_known and has not been verified through a literature search; this does not support claiming that the area is a gap.

Planned experiments

每个 cell 为 model×dataset。缓存嵌套 rollout(survivors-only 的多阶段树)并拆成估计/评估半。预测变量与结果变量使用独立 rollout。真实引导用调优的 PRM best-of-N 与 beam search(调 beam 宽和步骤切分),在匹配 token 下比较。预先做 power analysis(用 pilot 的方差分量模拟),据此确定每个 cell 的题数。通过率匹配子集用于图谱。计算:8×A100,pilot 约 150 GPU-hours,完整约 700 GPU-hours(1.5B–14B 为主,32B 至多 2 个数据集各 60 题)。

  • datasets: MATH-500(对强 RL 模型接近饱和,p(1-p)→0,需通过率匹配或只取中等难度题); OlympiadBench 子集; AIME24+25 合并(60 题,置信区间很宽); GPQA-Diamond; LiveCodeBench(测试执行,子集); GSM8K 仅作饱和的 sanity check,不进入主检验
  • baselines: PRM 增益的预测变量基线:PRM 的 ProcessBench F1;policy 通过率/最终答案熵(难度);平均链长;p(1-p);及其组合; 等 token majority voting 与通过率匹配的无引导采样,作为增益参考; 真实引导:Qwen2.5-Math-PRM 与 Math-Shepherd 的 best-of-N 与 beam search(调优 beam 宽与步骤切分,含长 CoT 的 newline 与 double-newline); label-free 一致性/agreement 引导基线(与 brief 的种子想法对应); oracle-MC 引导搜索与最终 verifier pass@N headroom 作为参考; naive(未修正)方差估计器
  • metrics: 主要:廉价估计(R=m=4)在 leave-one-model-out 下对真实引导增益(PRM beam/BoN 与 label-free 引导)的预测误差与 Spearman,与上述难度/长度/ProcessBench 基线用 Steiger/bootstrap 比较相关差异; 次要一致性检查:题目层面混合效应模型(模型、数据集随机截距,题目嵌套在 cell 中),报告加入 V(25%) 后相对 p(1-p)+长度+难度的似然比检验与 partial R^2,明确说明这近乎构造性结果; gain ~ H + capture + H:capture 的回归系数与交互项; 最优真实 PRM 方法捕获的 oracle H 比例; 单阶段上界与实测 oracle k=1 剪枝增益的比值(紧致度); RL vs SFT 在通过率匹配后的 V(25%) 差异及 bootstrap CI(先按题目再按模型重采样); 对 m、R、深度定义和切分方式的敏感性
  • ablations: 深度网格、R 与 m(估计稳定性、廉价版保真度); naive vs ANOVA 修正估计器; 步骤边界 vs token 固定窗口 vs 绝对 token 偏移的深度定义; held-out 评估半 vs 同半选择(量化 winner's curse); 逐题 vs 汇总 profile; 同一家族的 base / SFT-distill / RL / instruct 变体; 在同一递归搜索中用 PRM 分数替换 oracle-MC 以隔离 capture; 通过率匹配 vs 未匹配
  • expected: 预期 RL 长 CoT 变体的 V(25%) 更低(晚解析),早期剪枝的 H 更小;廉价估计在 held-out 模型上对真实 PRM/beam 增益的排序优于难度/链长/ProcessBench 基线;长 CoT 上 capture 进一步下降,使两个原因可分离。这些是预期而非结论:最可能的失败是 V 与 p(1-p) 高度共线,或廉价估计在 held-out 模型上不能泛化。
  • 否证条件: 预先登记,且按红队意见改为非循环形式:(a) 若廉价估计(R=m=4,独立 rollout)在 leave-one-model-out 下对真实 PRM/beam 增益的预测不优于 p(1-p)+链长+ProcessBench F1 基线(Steiger/bootstrap p>0.05,且误差降低不足一个预设的最小幅度),则诊断工具的主张被放弃,论文缩减为图谱。(b) 若通过率匹配后 RL 与 SFT 同基座变体的 V(25%) 差异不显著,则撤销'晚解析'主张。(c) 若单阶段界在大多数 cell 中松于实测 oracle 增益的 5 倍,则报告其为无信息。原先对 V 解释 H 的似然比检验(p>0.05 且 partial R^2<0.03)仅保留为一致性检查,因为它近乎构造性成立,不能作为有效的 kill 条件。

Each cell is a model×dataset combination. Cache nested rollouts (a survivors-only multi-stage tree) and split them into estimation/evaluation halves. Use independent rollouts for predictors and outcomes. For actual guidance, use tuned PRM best-of-N and beam search (tuning beam width and step segmentation), compared at matched token budgets. Conduct power analysis in advance (simulating with variance components from the pilot) to determine the number of problems per cell. Use pass-rate-matched subsets for the atlas. Compute: 8×A100, approximately 150 GPU-hours for the pilot and approximately 700 GPU-hours for the full study (primarily 1.5B–14B; for 32B, at most 2 datasets with 60 problems each).

  • datasets: MATH-500 (near saturation for strong RL models, p(1-p)→0; requires pass-rate matching or selecting only moderately difficult problems); an OlympiadBench subset; combined AIME24+25 (60 problems, very wide confidence intervals); GPQA-Diamond; LiveCodeBench (test execution, a subset); GSM8K only as a saturation sanity check, excluded from the primary test
  • baselines: Predictor baselines for PRM gains: PRM ProcessBench F1; policy pass rate/final-answer entropy (difficulty); mean chain length; p(1-p); and combinations of these. Equal-token majority voting and pass-rate-matched unguided sampling as references for gains. Actual guidance: Qwen2.5-Math-PRM and Math-Shepherd best-of-N and beam search (tuned beam width and step segmentation, including newline and double-newline for long CoT). A label-free consistency/agreement-guided baseline (corresponding to the seed idea in the brief). Oracle-MC-guided search and final-verifier pass@N headroom as references. A naive (uncorrected) variance estimator.
  • metrics: Primary: prediction error and Spearman correlation for the cheap estimate's (R=m=4) predictions of actual guidance gains (PRM beam/BoN and label-free guidance) under leave-one-model-out evaluation; compare differences in correlation against the difficulty/length/ProcessBench baselines above using Steiger/bootstrap tests. Secondary consistency check: a problem-level mixed-effects model (random intercepts for model and dataset, with problems nested within cells); report the likelihood-ratio test and partial R^2 for adding V(25%) to p(1-p)+length+difficulty, explicitly noting that this result is close to being implied by construction. Regression coefficients and the interaction term in gain ~ H + capture + H:capture. The proportion of oracle H captured by the best actual PRM method. The ratio of the single-stage upper bound to measured oracle k=1 pruning gain (tightness). The RL vs SFT difference in V(25%) after pass-rate matching, with bootstrap CI (resampling problems first, then models). Sensitivity to m, R, depth definitions, and segmentation.
  • ablations: Depth grid, R, and m (estimation stability and fidelity of the cheap version); naive vs ANOVA-corrected estimators; depth definitions based on step boundaries vs fixed token windows vs absolute token offsets; held-out evaluation half vs selection on the same half (quantifying the winner's curse); per-problem vs aggregated profiles; base / SFT-distill / RL / instruct variants within the same family; replacing oracle-MC with PRM scores in the same recursive search to isolate capture; pass-rate-matched vs unmatched comparisons
  • expected: RL variants with long CoT are expected to have lower V(25%) (late resolution) and smaller H for early pruning; the cheap estimate is expected to rank actual PRM/beam gains on held-out models better than difficulty/chain-length/ProcessBench baselines; capture is expected to decrease further on long CoT, allowing the two causes to be separated. These are expectations, not conclusions: the most likely failures are high collinearity between V and p(1-p), or failure of the cheap estimate to generalize to held-out models.
  • falsification criteria: Preregister these criteria, revised into a non-circular form in response to the red team: (a) if the cheap estimate (R=m=4, independent rollouts) does not predict actual PRM/beam gains better than the p(1-p)+chain-length+ProcessBench F1 baseline under leave-one-model-out evaluation (Steiger/bootstrap p>0.05 and error reduction below a prespecified minimum magnitude), abandon the diagnostic-tool claim and reduce the paper to an atlas. (b) If the V(25%) difference between RL and SFT variants sharing the same base is not significant after pass-rate matching, withdraw the 'late resolution' claim. (c) If the single-stage bound is more than 5 times the measured oracle gain in most cells, report it as uninformative. Retain the original likelihood-ratio test of V explaining H (p>0.05 and partial R^2<0.03) only as a consistency check, because it holds almost by construction and cannot serve as an effective kill criterion.

Two-week pilot plan

  1. 第 1–2 天:搭建 vLLM + prefix caching 的嵌套 rollout 管线,确定深度定义(绝对 token 偏移/步数)与步骤切分;在 20 道 MATH 题上做冒烟测试,验证 checker、拆半与缓存正确。同时通读 Critical Horizon 与 Branching Policy Optimization,确认定理与估计器的重叠程度,并决定是否调整定位。
  2. 第 3–6 天:用 Qwen2.5-7B-Instruct、DeepSeek-R1-Distill-Qwen-7B 和一个 7B/14B RL 变体,在 100 题 MATH-500 与 60 题 AIME 合并集上缓存 R=8 × 4 深度 × m=8 的 rollout(拆半);用通过率先筛出中等难度题,避免 p(1-p)→0。
  3. 第 7–9 天:计算 ANOVA 修正的 sigma_d^2 与 V(d);比较单阶段 oracle 增益与 Aven 界;在 survivors-only 树上计算多阶段 H(含 held-out 评估)。检查 naive 与修正估计器的差异。
  4. 第 10–12 天:在匹配 token 下运行 Qwen2.5-Math-PRM best-of-N/beam 得到真实增益,算出逐题 capture;在小规模上做 gain ~ H + capture 的回归草图,并检查廉价版(R=m=4,30 题)与完整版的一致性。
  5. 第 13–14 天:用 pilot 的方差分量做 power analysis,确定完整实验每 cell 的题数;依据预设的 kill 判据做一次中期 go/no-go,并写下是否收缩为图谱的决定。注意:3 个模型的 pilot 不足以检验 leave-one-model-out 预测,只能检验各量能否稳定估计、界是否成立。
  1. Days 1–2: build a nested-rollout pipeline using vLLM + prefix caching; establish the depth definition (absolute token offsets/step counts) and step segmentation; run a smoke test on 20 MATH problems to verify the checker, splitting into halves, and caching. In parallel, read Critical Horizon and Branching Policy Optimization in full to establish the degree of overlap in theorems and estimators and decide whether to adjust the positioning.
  2. Days 3–6: using Qwen2.5-7B-Instruct, DeepSeek-R1-Distill-Qwen-7B, and one 7B/14B RL variant, cache R=8 × 4 depths × m=8 rollouts (split into halves) on 100 MATH-500 problems and the combined 60-problem AIME set. First use pass rates to select moderately difficult problems, avoiding p(1-p)→0.
  3. Days 7–9: compute ANOVA-corrected sigma_d^2 and V(d); compare single-stage oracle gain with Aven's bound; compute multi-stage H on the survivors-only tree (including held-out evaluation). Examine differences between naive and corrected estimators.
  4. Days 10–12: run Qwen2.5-Math-PRM best-of-N/beam at matched token budgets to obtain actual gains and compute per-problem capture; sketch a small-scale gain ~ H + capture regression, and check agreement between the cheap version (R=m=4, 30 problems) and the full version.
  5. Days 13–14: conduct power analysis using the pilot's variance components to determine the problem count per cell for the full experiment; perform an interim go/no-go review against the prespecified kill criteria and record the decision on whether to reduce the scope to an atlas. Note: a 3-model pilot is insufficient to test leave-one-model-out prediction; it can only test whether the quantities can be estimated stably and whether the bound holds.

Risks and responses

  • 主检验近乎循环:H 由 MC p_d 选择定义,V 是同一 p_d 的方差,V 解释 H 近乎构造性成立(红队的致命问题) → 把主结果改为用独立 rollout 的廉价估计预测真实 PRM/beam/label-free 引导增益,leave-one-model-out;V→H 回归只作一致性检查。
  • V(d) 与难度 p(1-p) 共线,诊断冗余 → 通过率匹配的子集与对照基线;如果仍冗余,则按 kill 判据缩减为图谱。
  • 深度定义与步骤切分改变 profile(长 CoT 的百分比深度需要未知最终长度) → 主设定用绝对 token/步数网格,报告多种切分的敏感度;代码与 GPQA 的边界在 pilot 中先确定,必要时缩小范围到数学。
  • 多阶段 oracle 树的计算量与有限树上的乐观偏差 → survivors-only 分支、held-out 半评估、缩减到约 6 个模型×3 个数据集的子集、32B 只做少数 cell;对 m 做敏感性分析。
  • m=8 拆半后每半仅 4 条续写,单根估计噪声大;AIME 只有 60 题,置信区间宽;MATH-500 对强模型饱和 → 在 cell/题目层面汇总,用 power analysis 定题数,必要时增加 m 或题数;饱和数据集只取中等难度题。
  • PRM 在长 CoT 上分布外失效,使真实增益检验被 capture 主导 → 显式建模 capture,并加入 label-free 一致性引导作为不依赖 PRM 的结果变量。
  • 模型家族、训练方式和难度相互混淆;被他人抢先发表(图谱部分容易复制) → 同基座模型对 + 通过率匹配;尽快发布廉价 go/no-go 工具和预注册结果,把贡献重心放在非循环的预测检验上。
  • '晚解析'方向本身多数专家已预期,新意有限 → 强调定量、有范围的上界与对真实增益的预测;若预测检验失败,诚实地把论文定位为小型分析而不是方法论文。
  • The primary test is nearly circular: H is defined by selection using MC p_d, and V is the variance of the same p_d; V explaining H holds almost by construction (the red team's fatal issue). → Change the primary outcome to prediction of actual PRM/beam/label-free guidance gains using a cheap estimate from independent rollouts, with leave-one-model-out evaluation; use the V→H regression only as a consistency check.
  • V(d) is collinear with difficulty p(1-p), making the diagnostic redundant. → Use pass-rate-matched subsets and comparison baselines; if it remains redundant, reduce the scope to an atlas under the kill criteria.
  • Depth definitions and step segmentation alter the profile (percentage depth for long CoT requires an unknown final length). → Use an absolute token/step-count grid as the primary setting and report sensitivity to multiple segmentation methods; settle code and GPQA boundaries first in the pilot, narrowing the scope to mathematics if necessary.
  • Compute costs of the multi-stage oracle tree and optimistic bias on a finite tree. → Use survivors-only branching and held-out-half evaluation; reduce the grid to subsets of approximately 6 models×3 datasets; use 32B only in a few cells; analyze sensitivity to m.
  • Splitting m=8 leaves only 4 continuations in each half, so individual-root estimates are noisy; AIME has only 60 problems and wide confidence intervals; MATH-500 is saturated for strong models. → Aggregate at the cell/problem level, determine problem counts through power analysis, and increase m or problem counts if necessary; select only moderately difficult problems from saturated datasets.
  • PRMs fail out of distribution on long CoT, allowing capture to dominate the actual-gain test. → Model capture explicitly and include label-free consistency guidance as an outcome that does not depend on a PRM.
  • Model family, training regime, and difficulty are confounded; others may publish first (the atlas is easy to reproduce). → Use same-base model pairs + pass-rate matching; release the cheap go/no-go tool and preregistered results promptly, focusing the contribution on the non-circular prediction test.
  • Most experts already expect the 'late resolution' direction, limiting novelty. → Emphasize quantification, a scoped upper bound, and prediction of actual gains; if the prediction test fails, position the paper honestly as a small analysis rather than a methods paper.

Reviewer questions and responses

  • Q: 主假设近乎同义反复:H 是按 MC p_d 选择前缀的增益,V(d) 是同一 p_d 的 between-prefix 方差,用 V 在 p(1-p) 之外解释 H 主要是在验证一个恒等式。 A: 这个批评基本成立。我们的回应是改动设计,而不是辩护:主结果改为用独立 rollout 采集的廉价估计去预测真实 PRM beam/best-of-N 和 label-free 引导的增益(leave-one-model-out),V→H 仅作一致性检查,kill 条件也相应改写。若廉价估计预测不了真实增益,论文就缩为图谱。
  • Q: 所述 rollout 设计无法得到 10→25→50→75→100% 的 held-out 多阶段 oracle H,完整做需要指数分支的树,超出计算预算。 A: 原表述确实不充分。修订为 survivors-only 的树:每个深度只从上一层 top-k 存活前缀分支,并按模型和数据集给定 token 预算,把网格缩到约 6 个模型×3 个数据集的子集。H 仅是有限树上的经验参考值。具体的 token 成本还需要在 pilot 中实测,目前的 700 GPU-hours 估算未经验证。
  • Q: 'Gain_PRM = H × capture' 不是恒等式,秩相关也不会乘性地映射到选择增益,没有推导。 A: 同意。我们不再把它当作恒等式,改为显式回归 gain ~ H + capture + H:capture 并检验交互项;如果能在 Gaussian latent value 之类的明确模型下推出选择增益与 AUROC 的关系,会作为补充给出,否则只保留经验回归。
  • Q: 深度定义为最终长度的百分比,但长 CoT 的最终长度事先未知;分段方式也会大幅改变 profile。 A: 主设定改为绝对 token 偏移/步数,百分比深度只作对照,并报告多种切分的敏感度。代码与 GPQA 的边界定义目前规格不足,是已知缺口,要在 pilot 中解决,否则范围缩减到数学。
  • Q: 'RL 模型晚解析'是已被广泛预期的结论,新意有限;且 RL、SFT、难度、链长相混淆。 A: 方向确实不令人意外,贡献在于量化、同基座对照和通过率匹配,以及它能否预测真实增益。如果通过率匹配后差异消失,我们会撤销这一主张(kill 条件 b)。
  • Q: 与 Critical Horizon 等 2026 年工作过于接近,且他人可能很快做出图谱。 A: 我们只能依据所给的摘录判断机制相似,差异在于 held-out 估计、有范围的上界以及对真实 PRM 增益的预测检验。需要在 pilot 第一天通读后确认;图谱部分确实容易被复制,所以重心放在预测检验上。
  • Q: The primary hypothesis is nearly tautological: H is the gain from selecting prefixes using MC p_d, and V(d) is the between-prefix variance of that same p_d; using V to explain H beyond p(1-p) largely amounts to verifying an identity. A: This criticism is largely valid. Our response is to change the design rather than defend it: make the primary outcome prediction of actual PRM beam/best-of-N and label-free guidance gains using a cheap estimate collected from independent rollouts (leave-one-model-out), retain V→H only as a consistency check, and rewrite the kill criteria accordingly. If the cheap estimate cannot predict actual gains, reduce the paper to an atlas.
  • Q: The stated rollout design cannot produce held-out multi-stage oracle H for 10→25→50→75→100%; a complete implementation requires an exponentially branching tree, exceeding the compute budget. A: The original description was indeed insufficient. Revise it to a survivors-only tree: at each depth, branch only from the preceding layer's top-k surviving prefixes, assign token budgets by model and dataset, and reduce the grid to subsets of approximately 6 models×3 datasets. H is only an empirical reference value on a finite tree. The concrete token cost still needs to be measured in the pilot; the current estimate of 700 GPU-hours has not been verified.
  • Q: 'Gain_PRM = H × capture' is not an identity, and rank correlation does not map multiplicatively to selection gain; there is no derivation. A: Agreed. We will no longer treat it as an identity, replacing it with the explicit regression gain ~ H + capture + H:capture and testing the interaction term. If the relationship between selection gain and AUROC can be derived under an explicit model such as a Gaussian latent-value model, we will provide it as a supplement; otherwise, retain only the empirical regression.
  • Q: Depth is defined as a percentage of final length, but the final length of long CoT is unknown in advance; segmentation can also change the profile substantially. A: Change the primary setting to absolute token offsets/step counts, use percentage depth only as a comparison, and report sensitivity to multiple segmentation methods. The boundary definitions for code and GPQA are currently underspecified, a known gap to resolve in the pilot; otherwise, narrow the scope to mathematics.
  • Q: 'Late resolution in RL models' is already widely expected and offers limited novelty; RL, SFT, difficulty, and chain length are also confounded. A: The direction is indeed unsurprising. The contribution lies in quantification, same-base comparisons and pass-rate matching, and whether it can predict actual gains. If the difference disappears after pass-rate matching, we will withdraw this claim (kill criterion b).
  • Q: The work is too close to Critical Horizon and other 2026 papers, and others may produce an atlas quickly. A: We can judge the mechanisms to be similar only from the supplied excerpts; the distinctions are held-out estimation, a scoped upper bound, and a test of prediction of actual PRM gains. These must be confirmed by reading the full papers on the pilot's first day. The atlas is indeed easy to reproduce, so the focus is on the prediction test.

Conference fit

ICLR(也可投 NeurIPS 的分析/评测方向)。这是理解类/工具类论文,只有在统计设计严谨、并且廉价估计能在 held-out 模型上预测真实 PRM/搜索增益时才符合。若只剩下 V 解释 H 的结果,会被视为在验证恒等式,匹配度较低;此时应缩为图谱或投 workshop。内部评审的整体判断是 borderline(各项评分 2–3 分,evidence_plan 最低)。

ICLR (or the analysis/evaluation track of NeurIPS). This is a paper focused on understanding or providing a tool, and it fits only if the statistical design is rigorous and the cheap estimate can predict actual PRM/search gains on held-out models. If only the result that V explains H remains, it will be regarded as verifying an identity and the fit will be weak; in that case, reduce the work to an atlas or submit to a workshop. The overall internal-review judgment is borderline (individual scores of 2–3, with evidence_plan the lowest).

All research directionsAriadne v7.5