← InkTide cases

DEV-01 / FROZEN WRITING STUDY

Fluent English can still make the wrong inference

A grounded case from fixed synthetic author facts and model-only review.

01Fluent wording
02Check the control
03Repair the inference

What happened

In the regression-router replay, both round-1 drafts scored 5/5 for English. The guided conclusion nevertheless attributed improved accuracy to more expensive calls and treated an early-versus-late diagnostic as a dependence. The supplied ablation had the same 35% call rate but a different error, and the temporal cause remained unknown. The baseline preserved those limits and was preferred; the guided draft failed causality and claim-scope gates. After two real repair rounds it reached a tie, not a win.

What the comparison tells us

The numbers and vocabulary survived, but causal roles did not. This is a concrete reason for checking conclusions against controls rather than scoring language fluency alone.

What remains open

The unchanged baseline was judged by fresh reviewers across rounds; score shifts combine draft/rater variation. R3 still had local repetition and attachment ambiguities.

Read the comparison recordInkTide