The Rewrite, and What It Couldn't Fix

I fixed the framing error I could see. The trouble is the mistakes I can't see — and fluent output is not evidence either way.

Share
The Rewrite, and What It Couldn't Fix

Third in a side series. The first two are here and here.

In the last post I found that my summaries had a habit I had almost published as a success: given a paper that was not a clinical trial, the model would take the trial-shaped template it was working from and fill in fields that did not belong — a "primary outcome" for a basic science experiment that has no such thing. Not a wrong fact. A wrong frame. This post is about fixing that, and about the part of the problem a fix does not reach.

The cause was in the form

The template I had borrowed named a fixed set of headings, and they were all drawn from clinical research: study design, population, exclusion criteria, primary outcome, statistical methods. Every paper, whatever it was, got asked to fill them.

For a trial, fine. For a mouse tumour-evolution study, or a paper on foreign affairs, or a piece of literary criticism, those headings are the wrong questions. And a model asked the wrong question does not usually refuse it. It answers as best it can, which means it reaches into the paper and labels something as the "primary outcome" whether or not the paper has one. The rigidity I had blamed on the model was in the form I gave it.

The fix

So I changed what the form asks.

Instead of naming clinical-trial fields, the prompt now asks the model to first decide what kind of study the paper is — trial, basic experiment, theoretical work, review, or something from the humanities entirely — and then describe the methods in that field's own terms. It is told, in as many words, not to force a concept where it does not fit, and to reframe a heading into a question the paper can actually answer.

The headings themselves are now field-neutral. Not "primary outcome" but "approach." Not "exclusion criteria" but "what the study can and cannot claim." Language that means something whether the paper is a trial or a poem.

Testing it

I ran a basic science paper through the new version — another Nature oncology study, this one on chemotherapy-induced mutations in relapsed childhood cancers.

Under "approach," the old prompt would have manufactured a primary outcome. The new one wrote what the research actually did: mutational signature analysis across a whole-genome-sequenced cohort, linking treatment history to mutation patterns. No trial vocabulary. The framing fit the paper. The numbers — a near-tripling of private mutational signatures after treatment, platinum signatures appearing within a year in over a third of treated tumours — matched the abstract.

The specific category error I could recognize was gone.

What the rewrite couldn't fix

I want to be exact about that last sentence, because it is doing careful work.

I checked two things: that the figures matched the abstract, and that the framing was no longer wrong on its face. Both held. What I did not check — what I cannot check — is whether the summary captures what actually matters about the paper. Whether a cancer geneticist reading it would nod, or wince at an emphasis placed wrong, a caveat dropped, a secondary finding promoted over the real result.

That verification is beyond me, and rewriting the prompt did nothing to bring it closer. I fixed the mistake I was equipped to see. The mistakes I am not equipped to see are exactly the ones a better prompt is least likely to announce — because the better the output reads, the less it trips whatever expertise I do have, and outside my own field I have almost none to trip.

So the honest ledger is narrow. The tool is better than it was: it no longer makes one specific error I can catch. Whether it makes subtler errors I can't catch is a question I have no instrument for, and the fluency of the output is not evidence either way. If anything, fluency is the thing that should make me more cautious, not less.

Why this is still worth doing

None of which is an argument to stop.

Before any of this, the papers outside my daily work simply went unread. The system did not teach me cancer genomics — I still cannot judge that childhood-cancer paper the way its authors could. What it changed is that the paper is in front of me at all, in a form I will actually open, with a summary accurate on every point I am able to verify and a frame that no longer misrepresents the kind of study it is. For anything that matters, the original is still one click away, and now I know it exists.

That is the whole claim. Not that the summaries are the truth — that they are a door I will actually walk through, to papers I would otherwise never have opened. I built this to stop losing them, not to stop reading them, and on that narrow promise it delivers.


Next: the running cost, once there's enough of it to report honestly.