A complete loop, an unfinished research question.
Test-time compute for LLM math and code reasoning: reduce expensive step supervision, improve robustness and allocate compute where further sampling is worthwhile.
Run quality: degraded · 2 of 3 requested directions retained.
Two of three requested ideas passed all internal gates. The quality label records the shortfall; the two selected proposals remain unexecuted.
SEVEN LIVE RUNS
A development history with its costs intact.
| Run | Topic | Generated | Selected | New / cached calls | Reported cost | Minutes |
|---|---|---|---|---|---|---|
| micro1 | KV cache in long-context reasoning services | 6 | 2 | 53 / 0 | $4.14 | 17.5 |
| v1 | Test-time compute for LLM reasoning | 35 | 0 | 164 / 0 | $12.22 | 24.8 |
| v2 | Test-time compute for LLM reasoning | 34 | 2 | 267 / 0 | $20.07 | 26.5 |
| v3 | Test-time compute for LLM reasoning | 17 | 1 | 89 / 113 | ≈ $15.20 | 28 |
| db | Learned query optimization; Chinese input | 27 | 1 | 293 / 3 | $20.75 | 37 |
| corpus | Test-time compute for LLM reasoning | 27 | 1 | 229 / 0 | $14.15 | 21.8 |
| release | Test-time compute for LLM reasoning | 37 | 2 | 175 / 173 | $20.84 | ≈ 30 |
Author-reported development runs, including cached continuation; settings and cost accounting differ. The database row follows its detailed report (27 candidates, 37 minutes).
The author records seven live CLI runs. micro1 is a KV-cache prototype; later runs mainly examine reasoning-time compute plus a Chinese database case. “Two directions” describes the complete conference-targeted validation, not seven independent repetitions with identical settings.
Finalists, repairs and near misses.
Release review began with five original finalists and two subsequent repairs. Here are the four full reasoning dossiers delivered by the report: two selected ideas and two near misses. Each opens into its complete research argument.
Surprisingly-Popular Search: Peer-Prediction Scoring to Break Confident-but-Wrong Consensus
Majority voting among samples from one model can reinforce shared mistakes.
Coverage-Guided Sampling: Allocate More Code Samples Where Behavior Remains Unseen
Uniform sampling wastes compute on solved coding problems; consensus can confuse a solved problem with a wrong attractor.
Informed-vs-Naive Surprise: Construct Testable Information Differences to Challenge Wrong Consensus
Text reasoning chains from one model can confidently agree on a wrong answer, while peer prediction alone lacks independent information.
Guidance Headroom: Measure the Room for Step Guidance Before Investing in Reward Models
Observed search gains combine policy headroom, scorer quality and task difficulty, obscuring why reward models fail.
A SECOND DOMAIN
The same research ambition, a database question.
A Chinese-input run explored learned query optimization with SIGMOD and VLDB as target venues. Its detailed report records 27 generated directions, one selected dossier, two near misses, $20.75 equivalent cost and 37 minutes.
Optimizer Lock-In: Correct Self-Selected Runtime Feedback with Bounded Exploration
Runtime feedback covers only selected plans; good but overestimated alternatives may remain invisible.
Cardinality Tomography: Treat Runtime Observations as an Error-Attribution Problem
Intermediate result sizes provide local feedback; how can it transfer to unobserved subplans?
Cardinality Tomography with Context-Conditioned Error Atoms
Simple error decomposition misses context dependence, and corrections for unobserved plans may lack identifying information.
COMPUTE & COST
What the run cost.
175 new model-role calls · 17 reported minutes
173 cached calls · approximately 30 minutes all-in
Costs are equivalent dollar amounts reported by Claude CLI. Cache replay adds no new model charge; its prior cost remains in the all-in accounting. Model-role calls are not a measured count of underlying API requests. The supplied report does not contain run-level token totals.
How to read the result.
Selected means retained by this automated run, before executing the proposed research. The release AC scores are borderline, not conference decisions.
Top-k stability measures sensitivity of internal ranking to resampling; it is not an acceptance probability.
Literature search helps test differentiation, but retrieval gaps remain. Expected accuracy, compute or latency gains are proposed outcomes, not measurements.
