← Ariadne v7.5

LIVE RESEARCH / 2026-10-03

The research record

What survived, what changed, and what the run actually cost.

Research folios in a quiet atelier

A complete loop, an unfinished research question.

Test-time compute for LLM math and code reasoning: reduce expensive step supervision, improve robustness and allocate compute where further sampling is worthwhile.

37Generated directions
5Original finalists
2Selected dossiers
$11.04New-work equivalent cost
24Alive
6Rejected
7Superseded by repairs
Candidate state accounting: 24 + 6 + 7 = 37. Alive does not mean selected.

Run quality: degraded · 2 of 3 requested directions retained.

Two of three requested ideas passed all internal gates. The quality label records the shortfall; the two selected proposals remain unexecuted.

SEVEN LIVE RUNS

A development history with its costs intact.

RunTopicGeneratedSelectedNew / cached callsReported costMinutes
micro1KV cache in long-context reasoning services6253 / 0$4.1417.5
v1Test-time compute for LLM reasoning350164 / 0$12.2224.8
v2Test-time compute for LLM reasoning342267 / 0$20.0726.5
v3Test-time compute for LLM reasoning17189 / 113≈ $15.2028
dbLearned query optimization; Chinese input271293 / 3$20.7537
corpusTest-time compute for LLM reasoning271229 / 0$14.1521.8
releaseTest-time compute for LLM reasoning372175 / 173$20.84≈ 30

Author-reported development runs, including cached continuation; settings and cost accounting differ. The database row follows its detailed report (27 candidates, 37 minutes).

The author records seven live CLI runs. micro1 is a KV-cache prototype; later runs mainly examine reasoning-time compute plus a Chinese database case. “Two directions” describes the complete conference-targeted validation, not seven independent repetitions with identical settings.

Finalists, repairs and near misses.

Release review began with five original finalists and two subsequent repairs. Here are the four full reasoning dossiers delivered by the report: two selected ideas and two near misses. Each opens into its complete research argument.

A SECOND DOMAIN

The same research ambition, a database question.

A Chinese-input run explored learned query optimization with SIGMOD and VLDB as target venues. Its detailed report records 27 generated directions, one selected dossier, two near misses, $20.75 equivalent cost and 37 minutes.

COMPUTE & COST

What the run cost.

New release work$11.04

175 new model-role calls · 17 reported minutes

Reused historical work$9.80
Reported all-in work$20.84

173 cached calls · approximately 30 minutes all-in

Costs are equivalent dollar amounts reported by Claude CLI. Cache replay adds no new model charge; its prior cost remains in the all-in accounting. Model-role calls are not a measured count of underlying API requests. The supplied report does not contain run-level token totals.

How to read the result.

01

Selected means retained by this automated run, before executing the proposed research. The release AC scores are borderline, not conference decisions.

02

Top-k stability measures sensitivity of internal ranking to resampling; it is not an acceptance probability.

03

Literature search helps test differentiation, but retrieval gaps remain. Expected accuracy, compute or latency gains are proposed outcomes, not measurements.