Freeze the facts. Compare the choice.
Both conditions received the same author-invented evidence, numbers, terminology and task scope. The guided condition additionally received a frozen InkTide guide; the baseline had no InkTide guidance or published-reference access. Separate native reviewers saw randomized A/B packages. Five development pairs also received position-reversed review.
Five published main-paper PDFs calibrated the reviewers; they were not a same-facts third arm. No human preference labels were provided, and actual token/time budgets were not equalized. Later repairs reused development materials, so those rounds are not unseen tests.
Keep each round visible.
| Group | Pairs | Prefer InkTide | Tie | Prefer baseline | Strict wins |
|---|---|---|---|---|---|
| R1 | 5 | 2 | 1 | 2 | 1 |
| R2 | 5 | 0 | 3 | 2 | 0 |
| R3 · repair subset | 3 | 1 | 1 | 1 | 1 |
| Held-out | 4 | 3 | 1 | 0 | 2 |
The 17 primary reviews comprise 13 development/repair reviews and four held-out reviews. Reused baseline drafts and repeat reviews are counted separately from the 26 distinct generated manuscripts. A changed score across rounds includes reviewer variation as well as draft variation.
Preference and quality are different judgments.
The common-input failure swapped first-step roles in an ordered numeric ledger, even while both drafts followed the equations and trajectories. The blocked task had incompatible rules for when a dependent read could be issued. Reporting those outcomes prevents a preference count from hiding input validity.
Four pairs. Four English ties.
Current advantages, when present, are in evidence responsibilities, cross-section logic or repairs. The study has not established stable professional-English superiority. The reviewer calibration and the fixed synthetic workload do not replace an external human study.
Paragraph transfer is a separate snapshot.
The paragraph audit reviewed 249 pairs: 86 favored the guided version, 107 tied and 56 favored the baseline. Its 191 effective pairs contain 85 better and 106 similar outcomes; effectiveness is not a win count. Those records are distinct from whole-manuscript evaluation.
A figure can be checked against its input.
An earlier synthetic role demo used five latency samples per configuration, with means of 140 and 105 ms. The resulting 25% arithmetic difference tested figure/caption preservation; it was not an InkTide runtime benchmark. The checker rejected a deliberately altered 80 ms mean, invented significance, deployment claims and a missing source card.