A FIRST STEP
From a possibility to a research direction.
A research machine needs more than a stream of proposals. It needs to stay with a question: discover what is already known, listen to meaningful criticism, revise its assumptions and decide which experiment deserves the next unit of effort.
Ariadne moves this ambition into a running discovery loop. Its dossiers are aimed at rigorous research expectations, including CCF A-class venues such as NeurIPS, ICLR and ACL. The release’s two selected ideas remain pre-experiment directions; the supplied internal reviews are borderline.
THE RESEARCH LOOP
A thread through the search.
A conceptual view of the research responsibilities. Select a stage to explore its role.
Ground the question
Start with the bottleneck, constraints and closest work. A useful idea should explain what progress would change.
Open the search
Explore different explanatory and methodological directions, preserving alternatives before settling on a favorite.
Ask for evidence
Compare with literature and concrete objections. A contested judgment deserves a second look.
Make a real revision
Turn a supported objection into a change of hypothesis, scope or test. Useful criticism can become the next discovery.
Compare and calibrate
Compare directions and inspect sensitivity to judgment noise, ordering and incomplete evidence.
Test the strongest claim
Challenge the nearest-work distinction, practical feasibility and the observation that could overturn the idea.
Leave a research dossier
Deliver a research argument, remaining concerns and a falsifiable pilot. A dossier is the start of the next experiment.
FROM DAEDALUS TO ARIADNE
The process becomes a decision.
Traceable exploration
Structural thinking, tools and an accountable record of exploration.
Directions worth testing
Literature-grounded challenges, substantive repairs, comparison and a research dossier with a falsifiable next step.
The upgrade is demonstrated here as a change in screening and delivery capability. The supplied records do not constitute a compute-matched v7.4-versus-v7.5 quality benchmark.
THE RELEASE DOSSIERS
Two directions to take into the lab.
Informed-vs-Naive Surprise: Construct Testable Information Differences to Challenge Wrong Consensus
Text reasoning chains from one model can confidently agree on a wrong answer, while peer prediction alone lacks independent information.
Coverage-Guided Sampling: Allocate More Code Samples Where Behavior Remains Unseen
Uniform sampling wastes compute on solved coding problems; consensus can confuse a solved problem with a wrong attractor.
A REAL RUN / 03 OCT 2026
Search. Revise. Leave a record.
The release validation studied test-time compute for LLM math and code reasoning. It generated 37 directions, examined five original finalists and delivered two selected dossiers plus two near misses.

The release identified both planted controls (2/2). It recorded 78 pairwise comparisons and an 83% order-consistency rate. These describe this automated review, rather than human-review accuracy.
The requested three-direction target was not reached. The run is marked degraded and exposes remaining concerns. That disclosure is part of the research record, rather than replacing uncertainty with a polished score.
COMPUTE & COST
What the run cost.
175 new model-role calls · 17 reported minutes
173 cached calls · approximately 30 minutes all-in
Costs are equivalent dollar amounts reported by Claude CLI. Cache replay adds no new model charge; its prior cost remains in the all-in accounting. Model-role calls are not a measured count of underlying API requests. The supplied report does not contain run-level token totals.
MODELS & EVIDENCE
Choose the model. Keep the question.
The release supports CLI and API model access, research-budget controls, Chinese or English reports, and continuation from saved work. Lite, standard and deep offer different search scopes.
Literature inputs can use Semantic Scholar, OpenAlex, arXiv, CLI web search or a supplied corpus. Release validation used Claude CLI and collected literature; the adapter list is not a vendor-by-vendor live benchmark. These describe the private release, not an upgraded public web runner.
