This stage organizes research ideation around twelve contribution types, including methods, theory, perspectives and measurement. It fixes the subfield and essential prior work, combines first-principles kernels with transferable anchors, and develops candidates with a clear difference, risk and decisive test.
V4 expands the one-shot protocol into seven search spaces: domains, contribution lenses, kernels and anchors, mathematics, physics, experiments and reviewer attacks. The scope and budget are fixed first; candidates and refutations become structured state before review, repair and termination.
Hill focuses on engineering reliability. A model-call bridge and structural and semantic validation turn missing responses, placeholders and empty refutation ledgers into pauses or explicit invalid states, with review records saved separately.
Ocean assigns tree expansion to specialized model roles for field splitting, kernel mapping, anchor transfer, mechanism construction, experiment design and refutation. Python wraps and records their responses; the mathematics–physics bridge becomes a staged model-call pipeline.
Stars repairs advancement after a pause, frontier construction without a literature response and incomplete final-report fields. Operators are scheduled by node type, with visits and attempted actions recorded to make process state explicit.
Jupiter separates draft, checked, refuted and internally accepted candidate states. Ledger decisions return to the tree and block unsuitable candidates; state changes and events are saved so the tree and final report remain consistent.
Galaxy accepts both single-node and list model responses, injects kernel and anchor seeds and uses provenance flags to keep seed-only candidates out of the internal acceptance path. Reviewer rejection, candidate–ledger binding and call budgets become runtime constraints.
Cosmos uses stage-bucket scheduling to connect mechanism, experiment, attack and candidate stages, reducing the starvation of later-stage nodes by exploration scores. It also tightens model-field coverage, ledger identity binding and renewed refutation after repairs.
Nova replaces random mathematics–physics triggering with a deterministic depth condition, prioritizes the main candidate path until a first candidate exists and limits bridge branches. Terminal progress and repair-budget guards reduce branch starvation and unchecked repair growth.
Makes verification a separate execution layer. Candidate artifacts start unverified, and explicit checks govern promotion, rejection, unresolved evidence and honest failure.
Retains Eureka’s discovery layer while repairing its verification interface: translates discovery objects into Aurora payloads and reconciles every output to one terminal verdict.
Routes harvested equation payloads into verification and extends the Equation Lab with canonical forms, a prior-equation library, small numerical probes and physics diagnostics.
Adds explicit viability and bridge firewalls before idea harvesting. Failed candidates remain legible in structured failure records with repair reasons instead of being packaged as discoveries.
Wires Forge’s standalone components into the live discovery chain: resolved units, viability records, repair traces, CoMath state and prior caps contribute to one package.
Expands from equation-term mutation to the choice of mathematical descriptors. Domain, math and physics search trees produce first-class mappings and structured equation ASTs.
Adds compatibility, execution and quality constraints to Origin: prune mappings, select diverse elites, execute supported AST terms and route proof strategies by conjecture type.
Closes the ontology-to-idea evidence chain around one candidate ID. The release focuses on linkage, property provenance and replay rather than adding mappings or solvers.
Extends idea generation into mathematics, algorithms and architectures, organizing candidates through evidence graphs, strongest-prior checks and route-specific gates.
Turns candidates into compilable objects: equations use registered operator semantics, algorithms face subprocess tests and fixed benchmarks, and architectures produce executable forward paths, all bound to evidence and prior deltas.
Turns criticism into executable counterexamples and discriminating experiments, and organizes candidates as research-programme sequences through provenance, re-instantiation, null models, paired bets, anomalies and question genomes.
Organizes research around challenge responses, preregistered predictions and explanatory obligations, adding world-evidence accounting, learning-progress curiosity, prequential meters, explanation checks and world witnesses.
Treats reports as evidence indexes: true audits, answer-free sealed challenges, adjudicator-minted settlements, stable hash identities and prosecutor replay must agree, with demonstrations isolated as PASS_DEMO.
Replaces one-shot prose with ResearchDelta objects that record questions, representations, assumptions, mechanisms, falsifiers and lineage, organizing multi-generation search through isolated populations, multi-niche archives and independent event replay.
Upgrades the default path to an observation-driven controller: HTTP model providers propose constrained candidates, the operator locks tasks and baselines, execution yields concrete counterevidence and revision obligations, and evidence authorizes representation changes.
Inherit v6.8 engineering mechanisms and add primary-text evidence, rival explanations, typed actions, explicit stopping, Fusion contracts and local receipts in a standalone Aletheia module. Synthetic demos and fixed-weight meta evaluation check execution paths.
Add an authenticated Codex proposal provider and signed paired pilots. The model selected different correct next measurements under two observations. After same-case field clarification, both proposals passed revision and host contracts. Initial failures, usage and execution-gate repairs remain recorded.
Record the authorized literature-development loop and frozen eight-state A/B/C/D/E comparison, repairing source fields, query context, acquisition receipts, stop closure and observed-action registration. Feedback control is supported; ordinary active design matches the paired result, leaving separate harness gains unestablished.
The introduction draws on bundled notes, changelogs and central code. Software tests, synthetic diagnostics and scientific effectiveness use different evidence standards.