← Research record
db_I028 / ARIADNE 7.5Near miss

Cardinality Tomography with Context-Conditioned Error Atoms

Simple error decomposition misses context dependence, and corrections for unobserved plans may lack identifying information.

01Context and observations上下文与观测02Check identifiability检查可识别性03Trust justifiedcorrections采纳有依据的修正
Conceptual research hypothesis · No experimental result is depicted.

A research proposal generated by Ariadne. The experiments below are planned, and the review scores describe this internal selection.

核心洞见: 运行时观测到的中间大小不是某个子计划的缓存条目,而是对少数共享 log-error atoms 的线性测量。在单个查询内,atoms 已经以该查询自己的 filter 为条件,因此联合归因可以处理 LEO 式 predicate-keyed 因子处理不了的 filter-join 相关性。观测格(lattice)能确定、不能确定哪些子集的误差,是一个可计算、可检验的 identifiability 性质。

本轮 top-k 排名稳定性 — · BT 强度 — · AC 加权分 2.9/5(borderline)· 查新 NEAR(风险 high,证据等级 websearch)· 扛过红队 1 轮 · direction: Latent error factor models

AC 指出的致命问题: The idea's identifiability argument undermines its own intra-query claim. Mid-query, every remaining join decision involves edges whose query-local atoms have not been observed. Their posterior is therefore just the cross-query prior, which amounts to a Bayesian LEO, or it comes from probes, whose correlated-sampling assumption fails for multi-key joins. In addition, the edge-only model leaves out filter (node) errors, which breaks additivity. The headline claim that joint attribution fixes filter-join correlation for unobserved subsets is therefore likely to be small or illusory.

AC 的改进建议:

  • Add node (base-table filter) atoms alongside edge atoms, and test whether interaction terms (edge × co-joined filtered table) are needed. Report the within-query non-additive residual directly.
  • Restate the intra-query contribution precisely: show which DP decisions involve subsets whose rows lie in the observed span. Quantify this fraction on JOB/STATS-CEB plans before claiming 2x P-error gains.
  • Correct the prior-work framing with respect to ISOMER and SeqMaxEnt. Position the contribution as calibrated uncertainty, an identifiability gate and decision-theoretic re-planning, and compare head-to-head with ISOMER using the same feedback.
  • Add an equal-probe-budget baseline (exact-subplan re-optimization given the same sampled probes) and run the method on top of a strong base estimator, so that the gain can be attributed to joint attribution rather than to the probes.
  • Replace hash-correlated sampling for multi-key and cyclic joins with budgeted exact prefix execution or join-sampling (wander-join-style), and report how often censoring occurs.

把运行时观测到的中间结果大小视为对少量共享 log-error atoms 的线性测量,用贝叶斯递推更新做联合归因,并用可计算的 identifiability 判据决定何时信任对未观测子集的修正、何时回退到原生估计。

Core insight: An intermediate size observed at runtime is not a cache entry for one subplan; it is a linear measurement of a small set of shared log-error atoms. Within a single query, atoms are already conditioned on that query’s filters, so joint attribution can address filter–join correlation that LEO-style predicate-keyed factors cannot. Which subset errors the observation lattice can or cannot determine is a computable, testable identifiability property.

Top-k ranking stability — · BT strength — · AC weighted score 2.9/5 (borderline) · Novelty assessment NEAR (risk high; evidence level websearch) · Survived 1 red-team round · direction: Latent error factor models

Fatal concern identified by the AC: The idea's identifiability argument undermines its own intra-query claim. Mid-query, every remaining join decision involves edges whose query-local atoms have not been observed. Their posterior is therefore just the cross-query prior, which amounts to a Bayesian LEO, or it comes from probes, whose correlated-sampling assumption fails for multi-key joins. In addition, the edge-only model leaves out filter (node) errors, which breaks additivity. The headline claim that joint attribution fixes filter-join correlation for unobserved subsets is therefore likely to be small or illusory.

AC recommendations:

  • Add node (base-table filter) atoms alongside edge atoms, and test whether interaction terms (edge × co-joined filtered table) are needed. Report the within-query non-additive residual directly.
  • Restate the intra-query contribution precisely: show which DP decisions involve subsets whose rows lie in the observed span. Quantify this fraction on JOB/STATS-CEB plans before claiming 2x P-error gains.
  • Correct the prior-work framing with respect to ISOMER and SeqMaxEnt. Position the contribution as calibrated uncertainty, an identifiability gate and decision-theoretic re-planning, and compare head-to-head with ISOMER using the same feedback.
  • Add an equal-probe-budget baseline (exact-subplan re-optimization given the same sampled probes) and run the method on top of a strong base estimator, so that the gain can be attributed to joint attribution rather than to the probes.
  • Replace hash-correlated sampling for multi-key and cyclic joins with budgeted exact prefix execution or join-sampling (wander-join-style), and report how often censoring occurs.

Treat intermediate-result sizes observed at runtime as linear measurements of a small set of shared log-error atoms. Perform joint attribution through recursive Bayesian updates, using a computable identifiability criterion to decide when to trust corrections to unobserved subsets and when to fall back to native estimates.

Why the contribution could be memorable

模式是 跨领域结构迁移:把 network tomography(加性链路指标、路径测量、identifiability 即行空间成员关系、probe 设计)迁移到基数反馈上。评审可能记住的是"运行时反馈中哪些部分原理上可被推断"这个可检验的 identifiability 视角,以及用后验 regret gate 决定重规划。即使数值提升有限,这个框架和 gate 仍可能是可引用的贡献。但这只是可能性,需要实验证实,不应预设成立。

The pattern is cross-domain structural transfer: transfer network tomography—additive link metrics, path measurements, identifiability as row-space membership, and probe design—to cardinality feedback. Reviewers may remember the testable identifiability perspective, “which parts of runtime feedback can in principle be inferred,” and the use of a posterior regret gate for replanning. Even with limited numerical gains, this framework and gate could be citable contributions. This is only a possibility requiring experimental confirmation, not an established result.

Abstract

运行时反馈方法(POP、Perron 等人的 re-optimization、LEO)要么只修正已观测的子计划,要么用与上下文无关的 per-predicate 因子做启发式修正。它们不在重叠观测之间做联合信用分配,也不给出不确定性。我们提出 Cardinality Tomography:对查询 q 的连通子集 S,把 log(card_true/card_est) 建模为 query-local 边 atoms θ_{q,e} 之和加重尾噪声。同一查询内 filter 上下文固定,因此 filter-join 交互被吸收进 atom。跨查询则用三级层级先验(edge;edge×被过滤一侧的表;edge×相邻表的 filter signature)做 partial pooling。每次观测在 Kalman/recursive ridge 中增加一行,单次更新预期在亚毫秒级。空结果或被截断的 probe 作为 censored(Tobit)观测处理。只有当目标子集的行落在已观测行空间内(identifiable)时才信任修正,否则回退到原生估计。重规划由后验采样的 regret gate 控制,probe 按计划级 value of information 选择。评估在 PostgreSQL/DuckDB 上,使用 JOB、CEB、STATS-CEB、DSB 及漂移版本。内部评审指出的核心风险是:mid-query 时剩余决策多涉及未观测的边,收益可能很小。因此第 1 周的离线 kill 实验是决定性的。本文尚无实验结果,以上均为待检验的假设。

Runtime-feedback methods—POP, Perron et al.’s reoptimization and LEO—either correct only observed subplans or apply heuristic corrections with context-free per-predicate factors. They do not perform joint credit assignment over overlapping observations or provide uncertainty. We propose Cardinality Tomography: for a connected subset S of query q, model log(card_true/card_est) as the sum of query-local edge atoms θ_{q,e} plus heavy-tailed noise. Filter context is fixed within the same query, so filter–join interactions are absorbed into the atoms. Across queries, a three-level hierarchical prior—edge; edge × filtered-side table; edge × adjacent-table filter signature—provides partial pooling. Each observation adds one row to a Kalman/recursive-ridge update, with a target update time below one millisecond. Empty results or truncated probes are handled as censored Tobit observations. Trust a correction only when the target subset’s row lies in the observed row space and is identifiable; otherwise fall back to the native estimate. Replanning is controlled by a posterior-sampled regret gate, and probes are selected by plan-level value of information. Planned PostgreSQL/DuckDB evaluation covers JOB, CEB, STATS-CEB, DSB and their drift variants. Internal review identifies a central risk: remaining mid-query decisions mostly involve unobserved edges, so gains may be small. The Week 1 offline kill experiment is therefore decisive. There are no experimental results; everything above is a hypothesis to be tested.

Motivation

基数估计误差主导计划质量。learned CE 在数据漂移下脆弱,重训成本高。运行时反馈可以绕开"执行前必须估准"的假设,但现有做法把观测当作缓存、训练标签或 per-predicate 调整,并没有把重叠观测作为一个整体来解释。filter-join 相关性是 JOB/STATS-CEB/DSB 上的主要误差来源,会使 context-free 因子失效。目标是一条低开销(无 GPU、无重训)、对尾延迟和回退有保护的路径。需要注意:这里说"从未联合解释"并不准确,ISOMER 和 SeqMaxEnt 已经对重叠反馈做联合一致性处理,差异必须收窄到后面说明的几点。

Cardinality-estimation errors dominate plan quality. Learned cardinality estimation (CE) is fragile under data drift, and retraining is expensive. Runtime feedback can bypass the assumption that estimates must be accurate before execution, but existing approaches treat observations as caches, training labels or per-predicate adjustments rather than interpreting overlapping observations jointly. Filter–join correlation is the main error source on JOB/STATS-CEB/DSB and can invalidate context-free factors. The target is a low-overhead path, without a GPU or retraining, that protects tail latency and limits regressions. However, “never interpreted jointly” is inaccurate: ISOMER and SeqMaxEnt already impose joint consistency on overlapping feedback. Differentiation must be narrowed to the points below.

The proposed gap in prior work

专家通常认为加性 log-error 模型在 filter-join 相关下会失效,并把观测当作 cache、标签或 per-predicate 因子。这里的转变是让 atom 的上下文就是查询本身:查询内 filter 上下文是常数,可加性只需在同一查询的子集之间成立。跨查询迁移变成对 query-local atoms 的层级先验,而不是一个假设。但这一转变同时暴露了弱点:未观测边的 atom 只能靠先验或 probe 得到,否则后验就退化成贝叶斯版 LEO。内部评审把这点判为致命风险,我们同意它是最可能证伪本想法的地方。

Experts usually expect additive log-error models to fail under filter–join correlation, and treat observations as caches, labels or per-predicate factors. The shift here makes the query itself the atom’s context: filter context is constant within the query, so additivity need hold only across subsets of that query. Cross-query transfer becomes a hierarchical prior over query-local atoms rather than an assumption. However, this shift exposes a weakness: atoms for unobserved edges come only from the prior or probes; otherwise, the posterior reduces to a Bayesian version of LEO. Internal review calls this a fatal risk, and we agree it is the most likely point at which the idea will be falsified.

Proposed method

  1. 加性 atom 模型:r_q(S)=Σ_{e∈E(S)} θ_{q,e}+ε_S。按 AC 意见,必须同时加入节点(base-table filter)atoms,并检验 edge×共同连接的被过滤表 等交互项是否必要。直接报告查询内非加性残差。
  2. 层级先验:θ_{q,e} ~ N(μ_{e,c(q,e)}, τ²_type),c 取三级上下文(edge;edge×被过滤一侧的表;edge×相邻表 filter signature),未见上下文收缩到更粗层级。
  3. 在线后验更新:每个观测 (S, 真实大小) 增加关联矩阵 A 的一行,用 Kalman/recursive ridge 更新,代价 O(|E(S)|²)。对未观测 S' 输出 native×exp(A_S' mean),方差为 A_S' Σ A_S'^T。
  4. 噪声与非加性处理:per-subset 的 Student-t 噪声(Huber/IRLS 实现)。identifiability 检查:若 A_S' 不在已观测行的行空间内,则只使用先验均值部分并报告大方差,DP 回退到原生估计。
  5. Probe 观测:优先用 budget 受限的早停精确 join 或 wander-join 式 join sampling。hash 相关采样仅用于单键 join(AC 指出对多键和环形 join 不成立)。空或截断结果按 Tobit 做矩匹配截断高斯更新,绝不当作 0。
  6. Mid-query 重规划:在 breaker 处更新后验,用修正后的估计对所有子集重跑 DP;仅当后验样本下的期望收益超过切换成本加 margin 时才切换(posterior-sampled regret gate)。
  7. Probe 选择:按期望计划代价变化(value of information)贪心选择,D-optimal 作为消融对照。
  8. 跨查询更新:查询结束后用其 atoms 的后验更新 μ,并带指数遗忘以适应漂移。
  1. Additive atom model: r_q(S)=Σ_{e∈E(S)} θ_{q,e}+ε_S. Following the AC’s recommendation, node atoms for base-table filters must also be included, and the necessity of interaction terms such as edge × co-joined filtered table must be tested. Report within-query non-additive residuals directly.
  2. Hierarchical prior: θ_{q,e} ~ N(μ_{e,c(q,e)}, τ²_type), with c at three context levels: edge; edge × filtered-side table; edge × adjacent-table filter signature. Unseen contexts shrink toward coarser levels.
  3. Online posterior update: each observation (S, true size) adds a row to incidence matrix A, followed by a Kalman/recursive-ridge update costing O(|E(S)|²). For an unobserved S', output native×exp(A_S' mean), with variance A_S' Σ A_S'^T.
  4. Noise and non-additivity: use per-subset Student-t noise, implemented with Huber/IRLS. Identifiability check: if A_S' is not in the row space of observed rows, use only the prior-mean component and report large variance; DP falls back to the native estimate.
  5. Probe observations: prefer budget-limited, early-terminated exact joins or wander-join-style join sampling. Restrict hash-correlated sampling to single-key joins; the AC notes it is invalid for multi-key and cyclic joins. Handle empty or truncated results by moment-matched truncated-Gaussian Tobit updates, never as 0.
  6. Mid-query replanning: update the posterior at a breaker, then rerun DP for all subsets with corrected estimates. Switch only when expected gain under posterior samples exceeds switching cost plus a margin, using a posterior-sampled regret gate.
  7. Probe selection: greedily choose by expected change in plan cost, or value of information; use D-optimal design as an ablation baseline.
  8. Cross-query update: after a query completes, update μ using its atoms’ posteriors, with exponential forgetting for drift.

Distinction from nearest work

paper difference
Selectivity Estimation for Linear Queries via Online Learning (2026) 该文与本想法最接近(把反馈当作线性测量,facet 重叠约 0.45)。我们只掌握其摘要级证据,无法精确判断差异。我们的设想差异在于:log-error 的 query-local atoms、层级 partial pooling、identifiability gate、censored probe 和对计划决策的 regret gate。必须精读该文后再确认,若其已覆盖大部分机制则新颖性明显下降。
How I Learned to Stop Worrying and Love Re-optimization (2019) 该文用观测到的精确基数在 mid-query 重规划,只使用已观测子集。我们把观测传播到未观测子集并带不确定性,用 regret gate 替代启发式触发。该文的调优触发器是强基线,收益可能主要来自 probe 本身。
Efficient Query Re-optimization with Judicious Subquery Selections (2022) 该文选择执行哪些子查询来获得真实基数并重规划。我们的 VOI probe 选择在目标上相近,区别是以后验和 identifiability 为依据。证据只有一句摘要,具体对比需要读全文。
LEO – DB2's LEarning Optimizer (2001) LEO 按算子位置把误差启发式地、独立地分配给 predicates,并以 context-free 方式存入 catalog。我们联合求解重叠观测的信用分配,返回校准的不确定性。差异是否重要要靠 joint vs independent 的消融证明。
Learning Table Access Cardinalities with LEO (2002) 同为基于执行反馈的调整,侧重 table access。我们面向多表 join 子集,并有查询内联合归因。证据仅为摘要。
证据不足说明 ISOMER/SeqMaxEnt 不在提供的 closest papers 中,只在想法描述和 AC 意见里出现。它们已经对重叠反馈联合处理,因此'从未联合解释'的说法不能保留。真实差异收窄为贝叶斯不确定性、identifiability gate、决策级重规划,需要用相同反馈直接对比。
paper difference
Selectivity Estimation for Linear Queries via Online Learning (2026) The closest work to this idea, treating feedback as linear measurements, with facet overlap about 0.45. Only abstract-level evidence is available, so differentiation cannot be assessed precisely. The proposed differences are query-local log-error atoms, hierarchical partial pooling, an identifiability gate, censored probes and a regret gate for plan decisions. These must be confirmed through careful full-text reading; if that work covers most mechanisms, novelty falls substantially.
How I Learned to Stop Worrying and Love Re-optimization (2019) This work replans mid-query using exact observed cardinalities and only observed subsets. We propose propagating observations to unobserved subsets with uncertainty and replacing heuristic triggers with a regret gate. Its tuned trigger is a strong baseline; gains may arise mainly from the probes themselves.
Efficient Query Re-optimization with Judicious Subquery Selections (2022) This work chooses subqueries to execute for true cardinalities and replanning. Our VOI-based probe selection has a similar objective, but would rely on a posterior and identifiability. Evidence consists of one abstract sentence; a detailed comparison requires reading the full paper.
LEO – DB2's LEarning Optimizer (2001) LEO heuristically and independently attributes error to predicates according to operator position, storing it in a context-free catalog. We propose jointly solving credit assignment across overlapping observations and returning calibrated uncertainty. A joint-vs.-independent ablation must establish whether this distinction matters.
Learning Table Access Cardinalities with LEO (2002) Also uses execution-feedback adjustment, focusing on table access. We target multi-table join subsets and within-query joint attribution. Evidence is abstract-level only.
Insufficient-evidence note ISOMER/SeqMaxEnt are absent from the supplied closest-paper list, appearing only in the idea description and AC comments. They already process overlapping feedback jointly, so the claim “never interpreted jointly” must be removed. The actual differentiation is narrowed to Bayesian uncertainty, an identifiability gate and decision-level replanning, requiring direct comparisons using the same feedback.

Planned experiments

PostgreSQL 与 DuckDB,CPU 为主(64 核),无需 GPU。离线:对连通子集枚举真实基数,拟合不同 keying 方案,模拟 1-3 个观测的在线到达。在线:先通过 pg_hint_plan 回放计划,再接入 mid-query breaker hook。必须加入'同等 probe 预算'的 exact-subplan 基线,并在强基础估计器(如 LpBound 或 learned CE)之上再测一次,以把收益归因到联合归因而非 probe。

  • datasets: STATS-CEB; JOB 与 CEB(IMDB),含 CEB random-template 划分; DSB(变化参数); 漂移版 STATS-CEB(NeurBench 风格的插入与删除)
  • baselines: 原生 PostgreSQL 与 DuckDB; 忠实实现的 LEO(local 与 join 的 per-predicate 乘性因子,持久 catalog); ISOMER/max-entropy 一致性反馈(同样的反馈输入); Exact-subplan re-optimization(Perron/POP 风格),含同 probe 预算版本; 按模板索引的基数缓存; 梯度提升残差模型,以及带漂移微调的 FactorJoin; Pessimistic LpBound
  • metrics: 未观测子集上 log-error 的 held-out 解释方差(按 keying 方案); 未观测子集的 P-error 与 Q-error; 端到端延迟 P50/P99,以及每查询回退数量与幅度; 每次观测的更新时间; 后验覆盖率/校准,含 censored probes; 空/截断 probe 比例,以及 probe 成本与 probe 误差(分别报告); 被标记为 unidentifiable 的子集比例及其误差; DP 剩余决策中落在已观测行空间内的子集占比
  • ablations: Keying:仅 edge、edge×被过滤表、edge×filter signature、带层级的 query-local; 是否加入节点(filter)atoms 与交互项; 联合归因 vs 独立逐观测分配(隔离信用分配的增量); 向未观测子集传播 开/关; Tobit vs 朴素置零;均匀 1% 采样 vs 相关采样 vs 有界精确 probe; identifiability gate 开/关; VOI vs D-optimal vs 随机 probe; 后验采样切换 gate vs 固定阈值; 遗忘率;Huber vs 平方损失
  • expected: 预期(未验证)收益主要来自查询内、且集中在存在 filter-join 相关的查询上,目标是比 exact-subplan 与 LEO 低至少 2x 的 P-error。在 held-out 模板和 ad hoc 流上收益会缩小到仅查询内部分,如实报告。更新时间目标小于 1 ms。在强基础估计器上残差更小、结构更弱,收益很可能缩小。需要提醒:AC 认为 2x 偏乐观,较现实的情形是 P-error 有改善但延迟改善很小,延迟应作为主结果而非 P-error。
  • 否证条件: 第 1 周离线研究:若带层级的 query-local atoms 在 JOB/STATS-CEB 未观测子集上解释的 log-error 方差不足 50%,或相对 LEO 式 context-free 因子的优势小于 15 个百分点,则终止。此外,在模拟 native 计划已执行前缀的设定下(观测前缀子集,预测 DP 剩余决策所需子集),若后验不能优于 exact-subplan 缓存与 LEO,则终止。若在 held-out 模板划分上端到端相对 exact-subplan 反馈无统计显著收益,也终止。

Use PostgreSQL and DuckDB, primarily on a 64-core CPU machine without a GPU. Offline: enumerate true cardinalities of connected subsets, fit different keying schemes, and simulate the online arrival of 1–3 observations. Online: first replay plans through pg_hint_plan, then integrate a mid-query breaker hook. An exact-subplan baseline with the same probe budget is mandatory. Repeat the study over a strong base estimator, such as LpBound or learned CE, so gains can be attributed to joint attribution rather than probing.

  • datasets: STATS-CEB; JOB and CEB (IMDB), including CEB random-template splits; DSB with varied parameters; drifting STATS-CEB with NeurBench-style inserts and deletes.
  • baselines: Native PostgreSQL and DuckDB; faithfully implemented LEO with local/join per-predicate multiplicative factors and a persistent catalog; ISOMER/maximum-entropy consistent feedback with identical feedback inputs; exact-subplan reoptimization in the Perron/POP style, including an equal-probe-budget version; a template-indexed cardinality cache; a gradient-boosted residual model and FactorJoin fine-tuned under drift; pessimistic LpBound.
  • metrics: Held-out explained variance of log-error on unobserved subsets by keying scheme; P-error and Q-error on unobserved subsets; end-to-end P50/P99 latency and per-query regression counts and magnitudes; update time per observation; posterior coverage/calibration, including censored probes; proportion of empty/truncated probes, probe cost and probe error, reported separately; fraction and error of subsets marked unidentifiable; fraction of subsets required by remaining DP decisions that lie in the observed row space.
  • ablations: Keying: edge only, edge × filtered table, edge × filter signature, hierarchical query-local; inclusion of node/filter atoms and interactions; joint attribution vs. independent per-observation assignment, isolating the added value of credit assignment; propagation to unobserved subsets on/off; Tobit vs. naively setting results to zero; uniform 1% sampling vs. correlated sampling vs. bounded exact probes; identifiability gate on/off; VOI vs. D-optimal vs. random probes; posterior-sampled switching gate vs. a fixed threshold; forgetting rate; Huber vs. squared loss.
  • expected: Unverified expectations: gains come mainly from within-query inference and are concentrated on queries with filter–join correlation; the target is P-error at least 2x lower than exact-subplan feedback and LEO. On held-out templates and ad hoc streams, gains are expected to shrink to the within-query component, which will be reported accurately. Target update time is below 1 ms. Strong base estimators have smaller, less structured residuals, so gains are likely to diminish. The AC regards 2x as optimistic; a more realistic outcome is better P-error with very little latency improvement. Latency, rather than P-error, should be the main result.
  • Falsification criteria: In the Week 1 offline study, terminate if hierarchical query-local atoms explain less than 50% of log-error variance on unobserved JOB/STATS-CEB subsets, or improve by less than 15 percentage points over LEO-style context-free factors. Also terminate if, when simulating an already-executed prefix of a native plan—observing prefix subsets and predicting subsets needed by remaining DP decisions—the posterior cannot outperform exact-subplan caching and LEO. Terminate as well if end-to-end gains over exact-subplan feedback are not statistically significant on held-out-template splits.

Two-week pilot plan

  1. 第 1-3 天:为约 100 个 STATS-CEB 和 JOB 查询计算所有连通子集的真实基数(优先复用已有 ground truth),64 核约 1-2 天。
  2. 第 4-6 天:在四种 keying 方案加节点 atoms 与 LEO per-predicate 因子下拟合加性模型,按子集、查询、held-out 模板报告 held-out 解释方差,并直接报告查询内非加性残差。
  3. 第 7-9 天:量化 native 计划执行前缀之后,DP 剩余决策所涉子集中有多大比例落在已观测行空间内。若比例很低,先重新评估 2x P-error 的说法。
  4. 第 10-13 天:模拟在线到达(每查询 1-3 个观测),与 LEO、ISOMER 式、exact-subplan(含同 probe 预算)对比 P-error。
  5. 第 14 天:刻画 probe 质量:均匀采样与 hash 相关采样下的空结果比例(区分单键与多键 join),Tobit 更新的校准。随后应用 kill criterion,通过则开始 pg_hint_plan 回放,再接 breaker hook。
  1. Days 1–3: compute true cardinalities for all connected subsets of approximately 100 STATS-CEB and JOB queries, preferably reusing existing ground truth; this is expected to take 1–2 days on 64 cores.
  2. Days 4–6: fit additive models under the four keying schemes with node atoms, plus LEO per-predicate factors. Report held-out explained variance by subset, query and held-out template, and report within-query non-additive residuals directly.
  3. Days 7–9: after a native plan’s execution prefix, quantify the fraction of subsets needed by remaining DP decisions that lies in the observed row space. If this fraction is very low, reassess the claim of 2x P-error improvement first.
  4. Days 10–13: simulate online arrivals of 1–3 observations per query and compare P-error with LEO, ISOMER-style inference and exact-subplan methods, including equal-probe-budget comparisons.
  5. Day 14: characterize probe quality: empty-result rates under uniform and hash-correlated sampling, separating single-key from multi-key joins, and calibration of Tobit updates. Then apply the kill criterion. If it passes, begin pg_hint_plan replay and subsequently integrate a breaker hook.

Risks and responses

  • mid-query 时剩余决策涉及的 edge 都未被观测,后验退化为跨查询先验(等同贝叶斯 LEO),查询内收益很小(AC 认定的致命问题)。 → 第 1 周先量化可识别子集占比;用 VOI probe 预算主动观测;若占比过低,把贡献收缩为 identifiability 框架加 regret gate,或转向以跨查询为主的设定。
  • 仅 edge atoms 会双重计入 base-table 误差,且 edge 的有效 filter 上下文随子集变化,破坏可加性。 → 加入节点 atoms 与交互项,报告非加性残差;Student-t 噪声与 identifiability gate 限制损害,但可能限制收益上限。
  • hash 相关采样对多键和环形 join 不保持约 p 的输出,probe 大多被截断。 → 改用有界精确前缀执行或 wander-join 式采样,并报告截断比例;Tobit 只能保证无偏,不能保证信息量。
  • 收益主要来自 probe 而非归因;强基础估计器下残差缩小。 → 加入同 probe 预算基线和强基础估计器实验。
  • 与 LEO、ISOMER/SeqMaxEnt 及 2026 在线选择性论文接近,scoop 风险中到高。 → 精读 2026 论文,joint vs independent 的消融必须显示信用分配确有影响;否则不要推进为主贡献。
  • 复杂谓词的 filter signature 规范化繁琐;公开 benchmark 的跨查询复现性与真实重复负载不同。 → 先只支持常见谓词类型,其余回退到粗层级;如实声明 ad hoc 流只有查询内收益。
  • The edges involved in remaining mid-query decisions are all unobserved, so the posterior reduces to the cross-query prior—effectively Bayesian LEO—and within-query gains are small. The AC identifies this as fatal. → Quantify the fraction of identifiable subsets in Week 1. Use a VOI probe budget to obtain additional observations. If the fraction is too low, narrow the contribution to an identifiability framework plus regret gate, or redirect to a mainly cross-query setting.
  • Edge-only atoms double-count base-table error, while an edge’s effective filter context varies across subsets, violating additivity. → Add node atoms and interactions and report non-additive residuals. Student-t noise and the identifiability gate limit harm, but may also cap gains.
  • Hash-correlated sampling does not preserve an output fraction of approximately p for multi-key and cyclic joins, and most probes are truncated. → Use bounded exact-prefix execution or wander-join-style sampling and report the truncation rate. Tobit can ensure unbiasedness, not informativeness.
  • Gains arise primarily from probes rather than attribution, and stronger base estimators reduce residuals. → Add equal-probe-budget baselines and experiments with strong base estimators.
  • The work is close to LEO, ISOMER/SeqMaxEnt and the 2026 online-selectivity paper; the risk of being scooped is medium to high. → Read the 2026 paper carefully. The joint-vs.-independent ablation must show that credit assignment matters; otherwise do not pursue it as the main contribution.
  • Canonicalizing filter signatures for complex predicates is cumbersome; cross-query repeatability in public benchmarks differs from real recurring workloads. → Initially support common predicate types only and fall back to coarser levels for the rest. State accurately that ad hoc streams yield only within-query gains.

Reviewer questions and responses

  • Q: query-local edge atoms 吸收不了 filter-join 相关性,因为同一查询中 edge 的有效 filter 上下文随子集不同而变化。 A: 部分成立。可加性只在上下文固定时才合理,目前没有证据表明它在查询内成立。因此第 1 周直接测量非加性残差,并加入交互项与节点 atoms;若残差大,则 kill criterion 触发。
  • Q: 仅 edge atoms 会重复计入 base-table 估计误差,这些误差进入每个包含该表的子集,与度数无关。 A: 同意,这是模型设定缺陷。修正是加入节点 atoms,并作为必做消融(edge-only vs edge+node)报告。
  • Q: mid-query 剩余决策涉及的都是未观测的边,查询内联合归因帮不上忙。 A: 这是最严重的质疑,我们无法在理论上反驳。只能经验检验:量化剩余决策中落在已观测行空间内的子集比例,并用 VOI probe 主动补充观测。若比例低,查询内收益确实可能很小或不存在,需要缩小主张。
  • Q: 说重叠观测从未被联合解释是错的,ISOMER 已强制联合一致,SeqMaxEnt 也把反馈当线性测量。 A: 该批评正确,该表述将删除。真实差异收窄为:对 log-error 做贝叶斯后验与校准不确定性、identifiability gate、censored probe、计划级 VOI 与 regret gate。需要用相同反馈与 ISOMER 正面对比,来证明这些差异带来实际增益。
  • Q: hash 相关采样在多键 join 上不保持约 p 的输出,JOB/STATS 的多路 join probe 多数会被截断而无信息。 A: 成立。因此 probe 改用有界精确前缀执行或 join sampling,hash 相关采样仅限单键;并单独报告截断比例和 probe 成本。若有效观测过少,probe 路线就不可行。
  • Q: 收益可能来自 probe 而非归因(AC kill question)。 A: 通过同 probe 预算的 exact-subplan 基线和 joint vs independent 消融来回答;结果可能是收益确实较小,我们会如实报告。
  • Q: Query-local edge atoms cannot absorb filter–join correlation, because an edge’s effective filter context varies across subsets within the same query. A: Partly correct. Additivity is reasonable only with fixed context, and there is currently no evidence that it holds within a query. Therefore, Week 1 directly measures non-additive residuals and adds interactions and node atoms. Large residuals trigger the kill criterion.
  • Q: Edge-only atoms repeatedly count base-table estimation errors. Those errors enter every subset containing that table, independently of its degree. A: Agreed; this is a model-specification flaw. Add node atoms and report the required edge-only vs. edge+node ablation.
  • Q: All remaining mid-query decisions involve unobserved edges, so within-query joint attribution does not help. A: This is the most serious objection, and we cannot refute it theoretically. It can only be tested empirically: quantify the fraction of subsets for remaining decisions that lies in the observed row space, and actively supplement observations through VOI probes. If the fraction is low, within-query gains may indeed be small or nonexistent, requiring a narrower claim.
  • Q: It is wrong to say overlapping observations have never been interpreted jointly. ISOMER already enforces joint consistency, and SeqMaxEnt also treats feedback as linear measurements. A: Correct; that statement will be removed. Actual differentiation is narrowed to a Bayesian log-error posterior and calibrated uncertainty, an identifiability gate, censored probes, and plan-level VOI and regret gates. Direct comparison with ISOMER using the same feedback must show that these differences produce real gains.
  • Q: Hash-correlated sampling does not preserve an output fraction of approximately p in multi-key joins. Most multiway-join probes on JOB/STATS will be truncated and uninformative. A: Correct. Replace probes with bounded exact-prefix execution or join sampling, restricting hash-correlated sampling to single-key joins. Report truncation rates and probe cost separately. If too few observations are informative, the probing approach is infeasible.
  • Q: Gains may come from probes rather than attribution—the AC’s kill question. A: Answer through an equal-probe-budget exact-subplan baseline and the joint-vs.-independent ablation. Gains may indeed be small, and we will report them accurately.

Conference fit

适合 VLDB/SIGMOD 的 query optimization 方向(AC 给 venue_fit 4 分),前提是交付 PostgreSQL/DuckDB 的端到端集成,并以延迟(而非只是 P-error)作为主结果。但 AC 的总体判定为 borderline(differentiation 2、realism 2),受众真实但较小。建议只在第 1 周 kill 实验通过、且'可识别子集占比'足够高时,才投入后续数月。

Suitable for query-optimization research at VLDB/SIGMOD; the AC gives venue_fit 4. This requires end-to-end PostgreSQL/DuckDB integration and latency, rather than P-error alone, as the primary result. However, the overall internal AC verdict is borderline, with differentiation 2 and realism 2; the audience is real but fairly small. Commit the subsequent months only if the Week 1 kill experiment passes and the fraction of identifiable subsets is sufficiently high.

All research directionsAriadne v7.5