← InkTide

WRITING UNDER COMPARISON

The writing study

Original drafts, blind comparison and the cases that changed the workflow.

Freeze the facts. Compare the choice.

Both conditions received the same author-invented evidence, numbers, terminology and task scope. The guided condition additionally received a frozen InkTide guide; the baseline had no InkTide guidance or published-reference access. Separate native reviewers saw randomized A/B packages. Five development pairs also received position-reversed review.

Five published main-paper PDFs calibrated the reviewers; they were not a same-facts third arm. No human preference labels were provided, and actual token/time budgets were not equalized. Later repairs reused development materials, so those rounds are not unseen tests.

Keep each round visible.

GroupPairsPrefer InkTideTiePrefer baselineStrict wins
R152121
R250320
R3 · repair subset31111
Held-out43102

The 17 primary reviews comprise 13 development/repair reviews and four held-out reviews. Reused baseline drafts and repeat reviews are counted separately from the 26 distinct generated manuscripts. A changed score across rounds includes reviewer variation as well as draft variation.

Preference and quality are different judgments.

Two strict quality wins, one preferred/shared-input failure, one tie, one blocked.两组严格胜出、一组偏好但共同输入失败、一组平局、一组阻塞。
Two strict quality wins, one preferred/shared-input failure, one tie, one blocked.

The common-input failure swapped first-step roles in an ordered numeric ledger, even while both drafts followed the equations and trajectories. The blocked task had incompatible rules for when a dependent read could be issued. Reporting those outcomes prevents a preference count from hiding input validity.

Four pairs. Four English ties.

Model-only scientific-English scores; all four pairs tied.模型专业英语评分,四组配对全部相同。
Model-only scientific-English scores; all four pairs tied.

Current advantages, when present, are in evidence responsibilities, cross-section logic or repairs. The study has not established stable professional-English superiority. The reviewer calibration and the fixed synthetic workload do not replace an external human study.

Paragraph transfer is a separate snapshot.

The paragraph audit reviewed 249 pairs: 86 favored the guided version, 107 tied and 56 favored the baseline. Its 191 effective pairs contain 85 better and 106 similar outcomes; effectiveness is not a win count. Those records are distinct from whole-manuscript evaluation.

A figure can be checked against its input.

An earlier synthetic role demo used five latency samples per configuration, with means of 140 and 105 ms. The resulting 25% arithmetic difference tested figure/caption preservation; it was not an InkTide runtime benchmark. The checker rejected a deliberately altered 80 ms mean, invented significance, deployment claims and a missing source card.

InkTide overview