A research proposal generated by Ariadne. The experiments below are planned, and the review scores describe this internal selection.
核心洞见: 瓶颈是标注策略而不是学习器:argmin 选择使 propensity 为 0 或 1,重要性加权无意义;只有引入有界、受安全约束的随机化,标签才成为 off-policy 有效的训练数据。
本轮 top-k 排名稳定性 1.00 · BT 强度 1.72 · AC 加权分 3.2/5(borderline)· 查新 NEAR(风险 high,证据等级 websearch)· 扛过红队 1 轮 · direction: Label allocation by plan-regret dynamics
优化器只会在自己 argmin 选出的计划上拿到运行时基数标签,导致 TiCard/FLAIR/LEO 类反馈学习器"锁死"盲区;本文把计划选择视为带已知 propensity 的 logging policy,但需先解决"安全集够不到盲区"这一致命缺陷。
Core insight: The bottleneck is the labeling policy rather than the learner: argmin selection makes propensities 0 or 1, rendering importance weighting meaningless. Only bounded, safety-constrained randomization makes the labels valid training data for off-policy learning.
Top-k ranking stability 1.00 · BT strength 1.72 · AC weighted score 3.2/5 (borderline) · Novelty assessment NEAR (risk high; evidence level websearch) · Survived 1 red-team round · direction: Label allocation by plan-regret dynamics
The optimizer obtains runtime cardinality labels only for plans selected by its own argmin policy, causing TiCard/FLAIR/LEO-style feedback learners to lock in their blind spots. This proposal treats plan selection as a logging policy with known propensities, but must first resolve the fatal flaw that the safety set may not reach those blind spots.
Why the contribution could be memorable
属于 anomaly_reframe:把"反馈学习为什么在漂移后不恢复"重新解释为自选标签导致的 censoring/lock-in,并给出可测量指标(未执行替代子计划的误差)。评审会记住的是"优化器即 logging policy"这一框架和 lock-in 的量化;即使修复方案效果有限,这一度量与分析本身也有独立价值(AC 也认为这是最持久的贡献)。
This is an anomaly_reframe: it reinterprets “why feedback learning fails to recover after drift” as censoring/lock-in caused by self-selected labels, and supplies a measurable quantity—the error on unexecuted alternative subplans. Reviewers may remember the “optimizer as logging policy” framework and the quantification of lock-in. Even if the correction mechanism has limited effectiveness, the measurement and analysis have independent value; the AC also considers this the most durable contribution.
Abstract
基于代价的优化器利用运行时反馈(LEO、POP、FLAIR)或残差修正(TiCard)改进基数估计,通常把反馈视为"免费的监督标签"。但被标注的子计划是由优化器自身 argmin 选出的,标签集合因此被审查(censored):被低估的坏子计划会被执行并纠正,被高估的好替代方案则永远不会执行,也永远得不到纠正,漂移之后这种闭环依然持续。我们提出把优化器视为 off-policy 评估中的 logging policy:在 conformal 代价界"近似平局"的安全集内按 softmax 随机选择计划,记录 propensity,用 clipped doubly-robust 权重训练冻结基础估计器之上的残差模型,并用 ACI/weighted conformal 给出计划代价级别的重规划触发条件。我们计划在 STATS-CEB、JOB/CEB、DSB、TPC-DS 的漂移流上,度量未执行替代子计划的误差(lock-in 指标)、P-Error、总延迟与 P99 回归。需要强调:红队指出安全集可能无法覆盖"高估但实际很好"的子计划,这是当前设计的核心风险;因此我们把 lock-in 的度量与表征作为最稳妥的贡献,把传播性修复作为需要实验验证的次要贡献。
Cost-based optimizers improve cardinality estimation through runtime feedback (LEO, POP, FLAIR) or residual correction (TiCard), usually treating feedback as “free supervised labels.” However, the labeled subplans are chosen by the optimizer’s own argmin policy, censoring the label set: underestimated bad subplans are executed and corrected, while overestimated good alternatives are never executed and never corrected. This closed loop can persist after drift. We propose treating the optimizer as a logging policy in off-policy evaluation: select plans randomly by softmax within a near-tie safety set defined by conformal cost bounds, record propensities, train a residual model on top of a frozen base estimator using clipped doubly robust weights, and use ACI/weighted conformal methods to provide plan-cost-level replanning triggers. Planned experiments on drifting STATS-CEB, JOB/CEB, DSB and TPC-DS streams measure the error on unexecuted alternative subplans (the lock-in metric), P-Error, total latency and P99 regressions. The red team has identified a central risk: the safety set may fail to cover “overestimated but actually good” subplans. We therefore regard measuring and characterizing lock-in as the safest contribution, and propensity-based correction as a secondary contribution requiring experimental validation.
Motivation
反馈驱动的基数修正已成常态(TiCard、FLAIR 让在线修正很廉价),真正的瓶颈从"学习器"转向"标签从哪来"。标签来自优化器自己选择的计划,存在选择偏差,但文献中尚未针对残差/in-context 反馈学习器系统地命名、度量这一问题。若偏差确实导致漂移后的持续次优,则单纯增加反馈或换更强学习器都无法解决。
Feedback-driven cardinality correction is now common; TiCard and FLAIR make online correction inexpensive. The real bottleneck has shifted from “the learner” to “where labels come from.” Labels come from plans chosen by the optimizer itself and are subject to selection bias, yet the literature has not systematically named and measured this problem for residual/in-context feedback learners. If this bias does cause persistent suboptimality after drift, simply adding more feedback or switching to a stronger learner cannot solve it.
The proposed gap in prior work
一般认为"标签越多越好、学习器是瓶颈"。把优化器本身当作 OPE 中的 logging policy 在反馈文献里没有先例;而且它与学习器非加性交互(同一个残差模型在随机化日志上显著更好)。但要诚实指出:该论断目前只是假设,且"随机化 vs. 加权"各自的贡献尚未分离;选择偏差在 Balsa/Bao 的探索动机中已有一般性认识。
The usual assumptions are “more labels are better” and “the learner is the bottleneck.” Treating the optimizer itself as the logging policy in off-policy evaluation (OPE) has no precedent in the feedback literature, and it would interact non-additively with the learner: the same residual model would perform significantly better on randomized logs. However, this is currently only a hypothesis, and the respective contributions of randomization and weighting have not yet been separated. Selection bias is already generally recognized in the motivation for exploration in Balsa/Bao.
Proposed method
- Safe-set 随机 logging policy:候选计划中,conformal-upper 代价 ≤ (1+eps)×最佳 conformal-lower 代价者构成安全集,按估计代价的 softmax(温度 tau)抽样,其余情况取 argmin;记录 propensity,安全集内下界 pmin,全局预算限制探索查询比例(如 <5%)。
- 子计划 propensity:对包含该子计划的所有计划的 propensity 求和,作为 checkpoint 标签的采样概率。
- 残差学习器:冻结基础估计器,在 log-error 上训练残差模型(GBR/kNN/TabPFN 类),使用 clipped doubly-robust 权重(outcome model 即残差本身,权重来自记录的 pi);pi=0 的子计划回退到基础估计并放宽区间。
- 不确定性:对每个子计划 log-error 使用 ACI,并以已知 pi 作为似然比做 weighted conformal;漂移通过 ACI 在线调整 alpha 处理。
- 重规划触发:仅当替代方案的 conformal-upper 剩余代价 < 当前计划的 conformal-lower 剩余代价 + 切换代价时切换,对比较中的 k 个子计划做 union bound(alpha/k);保证为计划代价级别、有界 propensity 协变量偏移下的近似保证,漂移下形式化结果仅为 ACI 长期覆盖。
- 【必须新增,回应致命缺陷】盲区可达性修正:用 optimistic(下界)代价构造安全集,或对 pi=0 的高估子计划使用采样/仅基数/LIMIT 受限的低成本 probe,确保盲区子计划获得正 propensity;并报告每个 benchmark 上非平凡近似平局集合的比例与盲区覆盖率。
- 工程实现:在 PostgreSQL(pg_hint_plan)与 DuckDB 的优化器 hook 中包装计划选择,并记录 propensity 日志。
- Safe-set randomized logging policy: among candidate plans, those whose conformal-upper cost is ≤ (1+eps) × the best conformal-lower cost form the safety set. Sample by softmax over estimated costs, with temperature tau; otherwise choose argmin. Record propensities, enforce a lower bound pmin within the safety set, and use a global budget to limit the fraction of exploratory queries, for example <5%.
- Subplan propensity: sum the propensities of all plans containing the subplan to obtain the sampling probability of its checkpoint label.
- Residual learner: freeze the base estimator and train a residual model on log-error, such as GBR/kNN/TabPFN, using clipped doubly robust weights. The outcome model is the residual model itself, and weights come from the recorded pi. For subplans with pi=0, fall back to the base estimate and widen the interval.
- Uncertainty: apply ACI to each subplan’s log-error and use known pi as the likelihood ratio for weighted conformal prediction. Handle drift through ACI’s online adjustment of alpha.
- Replanning trigger: switch only when the alternative’s conformal-upper remaining cost < the current plan’s conformal-lower remaining cost + switching cost. Apply a union bound (alpha/k) over the k subplans in the comparison. The guarantee is an approximate guarantee at plan-cost level under covariate shift with bounded propensities; under drift, the formal result is limited to ACI’s long-run coverage.
- [Required addition in response to the fatal flaw] Correct blind-spot reachability: construct the safety set using optimistic (lower-bound) costs, or use low-cost sampling/cardinality-only/LIMIT-bounded probes for overestimated subplans with pi=0, ensuring that blind-spot subplans receive positive propensity. Report the proportion of nontrivial near-tie sets and blind-spot coverage on every benchmark.
- Engineering implementation: wrap plan selection through PostgreSQL’s pg_hint_plan and DuckDB optimizer hooks, and record propensity logs.
Distinction from nearest work
| paper | difference |
|---|---|
| Balsa: Learning a Query Optimizer Without Expert Demonstrations | Balsa 通过 safe execution 与 safe exploration 增加计划覆盖,但未记录 propensity,也不用于基数标签的 off-policy 纠偏。本文的差异在于记录 propensity 并做 DR 加权;但如果无加权的随机化即可达到同样效果,这个差异就会消失(AC 认为差异主要是组合)。 |
| Unbiased Learning to Rank with Unbiased Propensity Estimation | 该文在点击日志上用 IPW 做无偏 LTR,属于领域外方法。本文把相同的 propensity 逻辑迁移到优化器的计划选择,并且 propensity 是已知而非估计的;新意在迁移与安全约束,而非 IPW 本身。 |
| TiCard: Deployable EXPLAIN-only Residual Learning for Cardinality Estimation | TiCard 在原生估计器之上学习乘性残差,使用离线标签,不考虑标签选择偏差。本文把 TiCard 作为主要残差基线和学习器,贡献在标签策略而非残差模型。 |
| Conformal Prediction for Verifiable Learned Query Optimization | 该文用代价预测与实际延迟的 non-conformity 分数对延迟做界并验证计划质量。本文使用 ACI/weighted conformal 作用于子计划 log-error 并聚合成计划代价级别的重规划触发;证据有限(该文的 evaluation 维度未知),需读全文确认重叠程度。 |
| LEO – DB2's Learning Optimizer (Stillger et al.) | LEO 从已执行计划观测到的基数学习并反馈给优化器(精确匹配缓存),隐含地忽略选择偏差。本文把 LEO 缓存作为受 lock-in 影响的学习器之一进行度量并修复。提供的证据未覆盖 query-feedback 中'收集计划外表达式反馈'的工作,红队指出该类工作存在,需另行文献核查。 |
| paper | difference |
|---|---|
| Balsa: Learning a Query Optimizer Without Expert Demonstrations | Balsa increases plan coverage through safe execution and safe exploration, but does not record propensities or use them for off-policy correction of cardinality labels. This proposal differs by recording propensities and applying DR weighting. If unweighted randomization achieves the same effect, however, that distinction disappears; the AC considers the differentiation mainly a combination. |
| Unbiased Learning to Rank with Unbiased Propensity Estimation | This paper uses IPW on click logs for unbiased learning to rank (LTR), a method from another domain. This proposal transfers the same propensity logic to optimizer plan selection, with propensities known rather than estimated. The novelty lies in the transfer and safety constraints, not IPW itself. |
| TiCard: Deployable EXPLAIN-only Residual Learning for Cardinality Estimation | TiCard learns multiplicative residuals over a native estimator using offline labels, without considering label-selection bias. This proposal uses TiCard as the main residual baseline and learner; its contribution concerns the labeling policy rather than the residual model. |
| Conformal Prediction for Verifiable Learned Query Optimization | This paper bounds latency and verifies plan quality using nonconformity scores between predicted cost and actual latency. This proposal applies ACI/weighted conformal methods to subplan log-error and aggregates them into a plan-cost-level replanning trigger. Evidence is limited—the paper’s evaluation facet is unknown—and the full text must be read to establish the extent of overlap. |
| LEO – DB2's Learning Optimizer (Stillger et al.) | LEO learns from cardinalities observed in executed plans and feeds them back to the optimizer through an exact-match cache, implicitly ignoring selection bias. This proposal measures and corrects LEO’s cache as one of the learners affected by lock-in. The supplied evidence does not cover query-feedback work that collects feedback on expressions outside the selected plan; the red team notes that such work exists, requiring a separate literature check. |
Planned experiments
在 PostgreSQL 与 DuckDB 上实现优化器 hook。先做反事实模拟器:对每个查询枚举候选计划并离线执行得到真实子计划基数,使 logging policy 可回放。再做端到端在线流实验:带数据/负载漂移(NeurBench 风格更新)的重复模板流,多 seed、多查询流。所有基线调参预算一致。
- datasets: STATS-CEB(含 NeurBench 风格数据与负载漂移); JOB 与 CEB(数据更新、未见模板); DSB(规模化刷新); TPC-DS(倾斜刷新)
- baselines: 冻结的 PostgreSQL/DuckDB 统计信息; TiCard(使用相同 checkpoint 标签,主要残差基线); FLAIR; LEO 风格精确匹配反馈缓存; POP/Perron 重优化(无跨查询学习); 重训练的 NeuroCard/FactorJoin/MSCN(上界); Bao; 无安全集、无 propensity 加权的 epsilon-greedy 随机计划选择; Oracle 基数; (新增)Balsa/Bao 式 bandit 探索 + TiCard 无加权训练; (新增)SafeBound 等悲观界基础估计器作为更强基础模型
- metrics: 未执行替代子计划的误差(lock-in 度量); P-Error 与 Q-error 分位数; 总延迟与 P99 延迟; 相对原生优化器的回归数与探索 regret; 漂移下计划代价区间覆盖率; 更新时间、所需标签数、probe 比例; 非平凡近似平局集合占比、盲区子计划覆盖率
- ablations: 【关键新增】随机化探索 + 无加权训练 vs. 随机化 + IPW vs. DR(分离 off-policy 有效性与标签多样性); logging policy:仅 argmin、epsilon-greedy、safe-set softmax(eps、tau 扫描); 安全集大小 vs. 盲区覆盖率扫描; 基于 conformal-upper 的安全集 vs. 基于 optimistic 下界的安全集 vs. 低成本 probe; IPW vs DR vs 无加权;clipping 水平; 残差学习器类型(GBR、kNN、TabPFN 类); weighted conformal vs ACI vs 普通 split conformal; 触发器:无 union bound、无回退; 漂移类型与强度;已见 vs 未见模板; 达到给定误差所需探索查询数(样本复杂度曲线)
- expected: 假设(未验证):有偏标签学习器出现可见 lock-in;随机化 + 传播性加权的学习器相对 TiCard/FLAIR(checkpoint 标签)使漂移后 P-Error 至少降低 25%,探索使总延迟开销 <5%;延迟在重训练估计器的 15% 以内,更新成本低 50 倍以上;相对未校准触发器 P99 回归减少 20–40%。H2:eps 5–10%、探索查询 <5% 时,DR 残差弥合有偏标签学习器与 oracle 之间至少一半的 P-Error 差距。诚实预期:相当一部分增益可能来自标签多样性而非加权,需实验区分。
- 否证条件: 出现以下任一情形则停止或降级为分析论文:(a) 有偏 checkpoint 标签下的 TiCard/FLAIR 在漂移后 P-Error 和延迟与完整方法相差 5% 以内(无 lock-in);(b) 带加权的随机化日志不优于无加权的随机化日志(propensity 无增益);(c) 在四个 benchmark 中任意两个上,探索成本超过其节省的总延迟。另加:(d) 盲区子计划在任何可行 eps 下 propensity 仍为 0 且 probe 方案不能解决。
Implement optimizer hooks in PostgreSQL and DuckDB. First construct a counterfactual simulator: enumerate candidate plans for each query and execute them offline to obtain true subplan cardinalities, allowing logging policies to be replayed. Then run end-to-end online-stream experiments with repeated-template streams under data/workload drift, using NeurBench-style updates, multiple seeds and multiple query streams. Give all baselines the same tuning budget.
- datasets: STATS-CEB, including NeurBench-style data and workload drift; JOB and CEB, with data updates and unseen templates; DSB, with scaled refreshes; TPC-DS, with skewed refreshes.
- baselines: Frozen PostgreSQL/DuckDB statistics; TiCard using the same checkpoint labels, the main residual baseline; FLAIR; LEO-style exact-match feedback caching; POP/Perron reoptimization without cross-query learning; retrained NeuroCard/FactorJoin/MSCN as an upper reference; Bao; epsilon-greedy random plan selection without a safety set or propensity weighting; oracle cardinalities; [Added] Balsa/Bao-style bandit exploration + unweighted TiCard training; [Added] stronger base estimators using pessimistic bounds, such as SafeBound.
- metrics: Error on unexecuted alternative subplans, the lock-in measure; P-Error and Q-error quantiles; total latency and P99 latency; regressions relative to the native optimizer and exploration regret; plan-cost-interval coverage under drift; update time, number of labels required and probe fraction; proportion of nontrivial near-tie sets and coverage of blind-spot subplans.
- ablations: [Critical addition] Randomized exploration + unweighted training vs. randomization + IPW vs. DR, separating off-policy validity from label diversity; logging policy: argmin only, epsilon-greedy, safe-set softmax, with eps and tau sweeps; safety-set size vs. blind-spot-coverage sweeps; conformal-upper safety sets vs. optimistic-lower-bound safety sets vs. low-cost probes; IPW vs. DR vs. no weighting; clipping level; residual-learner type, including GBR, kNN and TabPFN; weighted conformal vs. ACI vs. ordinary split conformal; triggers without a union bound or fallback; drift type and strength; seen vs. unseen templates; exploratory queries needed to achieve a given error, yielding sample-complexity curves.
- expected: Unverified hypotheses: biased-label learners exhibit visible lock-in; a learner using randomization + propensity weighting reduces post-drift P-Error by at least 25% relative to TiCard/FLAIR using checkpoint labels, with exploration adding <5% total-latency overhead; latency is within 15% of that of a retrained estimator, with update cost more than 50 times lower; P99 regressions decrease by 20–40% relative to an uncalibrated trigger. H2: with eps at 5–10% and exploratory queries <5%, DR residual learning closes at least half the P-Error gap between a biased-label learner and the oracle. A substantial share of the gain may come from label diversity rather than weighting; experiments must distinguish these contributions.
- Falsification criteria: Stop or reduce the contribution to an analysis paper if any of the following occurs: (a) TiCard/FLAIR using biased checkpoint labels is within 5% of the full method in post-drift P-Error and latency, indicating no lock-in; (b) weighted randomized logs do not outperform unweighted randomized logs, indicating no propensity benefit; (c) exploration cost exceeds the total latency saved on any two of the four benchmarks. [Additional criterion] (d) blind-spot subplans retain propensity 0 at every feasible eps, and the probe approach cannot resolve this.
Two-week pilot plan
- 第1–4天:在 DuckDB/PostgreSQL 上搭建 JOB 反事实模拟器,对每个查询枚举候选计划并离线执行一次得到真实子计划基数;同时统计每个查询的近似平局集合大小。
- 第5–6天:先做诊断性实验——在带漂移的数据变体上统计'高估但实际好'的子计划有多少落入安全集(propensity>0),直接检验红队致命缺陷;若覆盖率极低,立即转向 optimistic 下界安全集或低成本 probe。
- 第7–9天:在在线流上比较 TiCard(GBR)分别用 argmin 日志、epsilon-greedy 日志、safe-set 随机日志训练,度量未执行子计划误差与 P-Error。
- 第10–12天:加入 IPW 与 DR,以及'随机化+无加权'关键消融,按 kill 标准 (a)(b) 判断。
- 第13–14天:将修正注入优化器,测量延迟与探索 regret,跑多 seed 方差;汇总 go/no-go。成功信号:500 个查询后有偏标签学习器在未执行替代子计划上的误差比随机标签学习器高至少 30%,且加权恢复至少一半,探索查询 <5%。计划说明为 21 天,前两周完成上述内容,算力用 64 核,A100 仅用于 TabPFN/FLAIR。
- Days 1–4: construct a JOB counterfactual simulator in DuckDB/PostgreSQL. Enumerate candidate plans for each query and execute each offline once to obtain true subplan cardinalities. Also measure the size of each query’s near-tie set.
- Days 5–6: first run a diagnostic experiment on data variants with drift: count how many “overestimated but actually good” subplans enter the safety set (propensity>0), directly testing the red team’s fatal concern. If coverage is extremely low, immediately switch to an optimistic-lower-bound safety set or low-cost probes.
- Days 7–9: on online streams, compare TiCard (GBR) trained on argmin logs, epsilon-greedy logs and safe-set randomized logs. Measure unexecuted-subplan error and P-Error.
- Days 10–12: add IPW and DR and the critical “randomization + no weighting” ablation. Evaluate kill criteria (a) and (b).
- Days 13–14: inject corrections into the optimizer, measure latency and exploration regret, assess variance across multiple seeds, and summarize the go/no-go decision. Success signal: after 500 queries, the biased-label learner’s error on unexecuted alternatives is at least 30% higher than that of the random-label learner, weighting recovers at least half the gap, and exploratory queries remain <5%. The plan is described as 21 days; the first two weeks complete the work above. Use 64 CPU cores, with the A100 only for TabPFN/FLAIR.
Risks and responses
- 致命风险:安全集在构造上排除了高估的好替代方案(pi=0),机制够不到动机中的盲区。 → 用 optimistic 下界构造安全集;引入采样/仅基数/LIMIT 受限的 probe;先做覆盖率诊断,若无法让盲区获得正 propensity,则转为分析论文。
- 近似平局集合太小,重叠不足,lock-in 仅弱可纠正。 → 报告各 benchmark 非平凡集合占比;扫描 eps;必要时扩大安全集并如实报告 regret。
- DR/加权估计在小样本下方差大,clipping 引入偏差。 → clipping 扫描,自归一化加权,残差模型使用正则化先验。
- 增益来自标签多样性而非 off-policy 加权。 → 必做'随机化+无加权'消融与 Bao/Balsa 式探索基线;若无加权即可匹配,则把贡献重定位为 lock-in 度量。
- union bound 区间过宽,触发器过于保守。 → 报告触发率;尝试更紧的联合方法(如 Bonferroni 替代、计划级直接校准)。
- 漂移下 conformal 覆盖只是近似(仅长期 ACI 保证)。 → 如实表述保证范围,经验报告短窗口覆盖。
- 随机化提高延迟方差;JOB 仅 113 条查询,显著性难达;lock-in 依赖人为设计的重复模板漂移流。 → 多 seed、多查询流、bootstrap 置信区间;使用 CEB/DSB 获得更长的流;公开漂移构造并给出对其不敏感的分析。
- 更强基础估计器(如 SafeBound)缩小盲区和近似平局集合,优势消失。 → 把这类基础估计器纳入实验,明确适用范围。
- 商用优化器不愿随机化计划,采用受限。 → 强调严格的探索预算与 regret 上界,提供仅用于影子/测试环境的部署模式。
- Fatal risk: the safety set excludes overestimated good alternatives by construction (pi=0), preventing the mechanism from reaching the blind spots that motivate it. → Construct the safety set with optimistic lower bounds; introduce sampling/cardinality-only/LIMIT-bounded probes; first diagnose coverage. If blind spots cannot receive positive propensity, reposition the work as an analysis paper.
- Near-tie sets are too small, overlap is insufficient, and lock-in is only weakly correctable. → Report the proportion of nontrivial sets on each benchmark, sweep eps, and widen the safety set if necessary while honestly reporting regret.
- DR/weighted estimates have high small-sample variance, while clipping introduces bias. → Sweep clipping levels, use self-normalized weighting and regularized priors in the residual model.
- Gains come from label diversity rather than off-policy weighting. → Require the “randomization + no weighting” ablation and Bao/Balsa-style exploration baselines. If no weighting matches the method, reposition the contribution as lock-in measurement.
- Union-bound intervals are too wide, making the trigger overly conservative. → Report trigger rates and try tighter joint approaches, such as alternatives to Bonferroni or direct calibration at plan level.
- Conformal coverage under drift is only approximate; only long-run ACI coverage is guaranteed. → State the guarantee’s scope accurately and report empirical short-window coverage.
- Randomization increases latency variance; JOB has only 113 queries, making significance difficult; lock-in depends on deliberately constructed repeated-template drift streams. → Use multiple seeds and query streams, bootstrap confidence intervals, CEB/DSB for longer streams, public drift constructions and analyses insensitive to those constructions.
- A stronger base estimator, such as SafeBound, reduces blind spots and near-tie sets, eliminating the advantage. → Include such base estimators and define the method’s scope.
- Commercial optimizers may resist randomized plan selection, limiting adoption. → Emphasize strict exploration budgets and regret bounds, and provide a deployment mode restricted to shadow or test environments.
Reviewer questions and responses
- Q: 安全集规则够不到论文关注的盲区:高估但实际很好的替代方案离 argmin 很远,落在近似平局安全集之外,pi=0,按方法自身规则永不被纠正。 A: 这个批评是对的,而且是目前设计的致命缺陷。只有当高估幅度落在 eps 窗口内、或 conformal 区间足够宽时,安全集才会覆盖这些子计划。我们的对策是:(1) 以 optimistic 下界构造安全集;(2) 对 pi=0 的子计划使用低成本 probe;(3) 先在模拟器中量化盲区子计划的 propensity 覆盖率。若覆盖率仍然很低,我们不会宣称纠正了盲区,论文会降级为 lock-in 的度量与表征。
- Q: 在查询反馈文献中,改变计划选择使反馈覆盖所选计划之外的表达式、并限制开销,已有人提出;learned optimizer 中的 bandit 探索也很标准,新意被削弱。 A: 部分同意。提供给我们的相关文献证据有限,不能断言已覆盖所有 query-feedback 工作,需做更完整的文献核查。我们不声称'探索'本身新颖,新意限于:对基数标签的显式 propensity 日志、DR 残差训练、以及计划代价级 conformal 触发。如果核查发现这些也已存在,则定位转向 lock-in 的度量与分析。
- Q: 缺失关键消融:随机化探索 + 无加权训练。没有它,无法把增益归因于 off-policy 有效性而非标签多样性。 A: 同意,并已加入为必做消融与 kill criterion (b)。同时加入 epsilon-greedy 和 Bao/Balsa 式探索喂给 TiCard 的无加权基线。如果加权无额外收益,我们会如实报告,并把贡献重新定位。
- Q: 这主要是 A+B 的组合(探索 + IPW + conformal),且 conformal 部分与 Conformal Prediction for Verifiable Learned Query Optimization 有重叠。 A: 组合性质成立。我们的辩护是:此组合服务于一个新的诊断(lock-in)且只有在随机化日志上 IPW 才有效。但如果实验不能展示非加性交互,这一辩护就站不住。与该 conformal 论文的重叠需要精读全文后再界定。
- Q: 评估是人为设计的:合成漂移、重复模板流,可能偏向本方法;没有真实生产轨迹。 A: 属实。我们会公开漂移构造,使用多种漂移类型/强度及未见模板,并将结论限制在这些设置内;真实 trace 若无法获得,则在局限性中明示。
- Q: 系统深度不足,读起来像 ML 论文,商用优化器也不会随机化计划。 A: 需要通过在 PostgreSQL/DuckDB 中的端到端实现、探索预算与 regret 上界来弥补。采用受限是真实局限,我们建议先面向影子执行或测试环境。
- Q: The safety-set rule cannot reach the blind spots motivating the paper. Overestimated but actually good alternatives are far from argmin, outside the near-tie safety set, with pi=0, and are never corrected under the method’s own rules. A: This criticism is correct and identifies a fatal flaw in the current design. The safety set covers these subplans only if the overestimation falls within the eps window or conformal intervals are sufficiently wide. Our responses are: (1) construct the safety set with optimistic lower bounds; (2) use low-cost probes for pi=0 subplans; (3) first quantify propensity coverage of blind-spot subplans in the simulator. If coverage remains very low, we will not claim to correct blind spots, and the paper will be reduced to measuring and characterizing lock-in.
- Q: Query-feedback literature already proposes changing plan selection so feedback covers expressions outside the selected plan while limiting overhead. Bandit exploration in learned optimizers is also standard, weakening the novelty. A: We partly agree. The supplied related-work evidence is limited, so we cannot claim coverage of all query-feedback work and need a fuller literature audit. We do not claim exploration itself is new. The claimed novelty is limited to explicit propensity logging for cardinality labels, DR residual training and plan-cost-level conformal triggers. If these also prove to have precedents, the work will shift toward measuring and analyzing lock-in.
- Q: A critical ablation is missing: randomized exploration + unweighted training. Without it, gains cannot be attributed to off-policy validity rather than label diversity. A: Agreed; it is now a required ablation and kill criterion (b). We also add epsilon-greedy and Bao/Balsa-style exploration feeding unweighted TiCard. If weighting gives no additional benefit, we will report that and reposition the contribution.
- Q: This is mainly an A+B combination of exploration, IPW and conformal prediction, and the conformal component overlaps with Conformal Prediction for Verifiable Learned Query Optimization. A: The combination characterization is valid. Our argument is that the combination serves a new diagnosis, lock-in, and IPW is valid only on randomized logs. If experiments cannot demonstrate non-additive interaction, however, that argument fails. The extent of overlap with the conformal paper requires careful full-text reading.
- Q: The evaluation is contrived: synthetic drift and repeated-template streams may favor the method, and there are no real production traces. A: Correct. We will publish the drift constructions, use multiple drift types and strengths and unseen templates, and limit conclusions to those settings. If real traces cannot be obtained, we will state that limitation.
- Q: The systems depth is insufficient; the work reads like an ML paper, and commercial optimizers would not randomize plans. A: End-to-end implementation in PostgreSQL/DuckDB, exploration budgets and regret bounds must address this. Adoption constraints are real; we recommend initially targeting shadow execution or test environments.
Conference fit
VLDB(或 SIGMOD)适合,前提是以 PostgreSQL/DuckDB 中的端到端实现、真实延迟指标和多 benchmark 漂移实验为核心。若盲区可达性问题无法解决,较稳妥的方案是调整为 VLDB 的实验与分析类论文(量化 lock-in,并以传播性日志作为次要设计)。当前 AC 评分为 borderline,风险集中在差异化、消融与盲区覆盖三点,建议以 2–3 周的 pilot 先验证这些再决定是否投入数月。
VLDB, or SIGMOD, is a fit provided the core is an end-to-end PostgreSQL/DuckDB implementation, real latency metrics and drift experiments across multiple benchmarks. If blind-spot reachability cannot be resolved, a safer direction is a VLDB experimental/analysis paper that quantifies lock-in and treats propensity-aware logging as a secondary design. The current internal AC assessment is borderline, with risks concentrated in differentiation, ablations and blind-spot coverage. A 2–3-week pilot should test these points before committing several months.