AutoResearch/Innovation Generator

INTRODUCING / 7.5.0

Ariadnev7.5

Let evidence carry the next idea.

Begin with a question. Let evidence guide exploration and criticism drive evolution. Ariadne v7.5 is a first step toward an autonomous research machine, developing directions worth testing.

A red thread through a vellum labyrinth
37Generated directions
5Original finalists
2Selected dossiers
$11.04New-work equivalent cost

A FIRST STEP

From a possibility to a research direction.

A research machine needs more than a stream of proposals. It needs to stay with a question: discover what is already known, listen to meaningful criticism, revise its assumptions and decide which experiment deserves the next unit of effort.

Ariadne moves this ambition into a running discovery loop. Its dossiers are aimed at rigorous research expectations, including CCF A-class venues such as NeurIPS, ICLR and ACL. The release’s two selected ideas remain pre-experiment directions; the supplied internal reviews are borderline.

THE RESEARCH LOOP

A thread through the search.

A conceptual view of the research responsibilities. Select a stage to explore its role.

01Ground the question找到真正的问题02Open the search展开不同可能03Ask for evidence让批评有依据04Make a real revision让想法实质改变05Compare and calibrate比较,也校准06Test the strongest claim追问最强的主张07Leave a research dossier交付可检验的方向
01

Ground the question

Start with the bottleneck, constraints and closest work. A useful idea should explain what progress would change.

FROM DAEDALUS TO ARIADNE

The process becomes a decision.

7.4 / DAEDALUS

Traceable exploration

Structural thinking, tools and an accountable record of exploration.

7.5 / ARIADNE

Directions worth testing

Literature-grounded challenges, substantive repairs, comparison and a research dossier with a falsifiable next step.

The upgrade is demonstrated here as a change in screening and delivery capability. The supplied records do not constitute a compute-matched v7.4-versus-v7.5 quality benchmark.

THE RELEASE DOSSIERS

Two directions to take into the lab.

Explore the full finalist and repair record →

A REAL RUN / 03 OCT 2026

Search. Revise. Leave a record.

The release validation studied test-time compute for LLM math and code reasoning. It generated 37 directions, examined five original finalists and delivered two selected dossiers plus two near misses.

24Alive
6Rejected
7Superseded by repairs
Candidate state accounting: 24 + 6 + 7 = 37. Alive does not mean selected.
Research folios linked by a fine thread

The release identified both planted controls (2/2). It recorded 78 pairwise comparisons and an 83% order-consistency rate. These describe this automated review, rather than human-review accuracy.

The requested three-direction target was not reached. The run is marked degraded and exposes remaining concerns. That disclosure is part of the research record, rather than replacing uncertainty with a polished score.

Read all seven development runs →

COMPUTE & COST

What the run cost.

New release work$11.04

175 new model-role calls · 17 reported minutes

Reused historical work$9.80
Reported all-in work$20.84

173 cached calls · approximately 30 minutes all-in

Costs are equivalent dollar amounts reported by Claude CLI. Cache replay adds no new model charge; its prior cost remains in the all-in accounting. Model-role calls are not a measured count of underlying API requests. The supplied report does not contain run-level token totals.

MODELS & EVIDENCE

Choose the model. Keep the question.

The release supports CLI and API model access, research-budget controls, Chinese or English reports, and continuation from saved work. Lite, standard and deep offer different search scopes.

Claude Code CLIDocumented live release runs
Codex CLI / mixed CLIImplemented adapters
Anthropic / OpenAI APIImplemented native transports
OpenAI-compatible APICompatible endpoint adapter

Literature inputs can use Semantic Scholar, OpenAlex, arXiv, CLI web search or a supplied corpus. Release validation used Claude CLI and collected literature; the adapter list is not a vendor-by-vendor live benchmark. These describe the private release, not an upgraded public web runner.

THE NEXT EXPERIMENT

A better question deserves a longer thread.