5W1H as a prompt structure: an open question
Does constraining a model to who / what / when / where / why / how produce more usable story skeletons than a free-form prompt? We do not know yet. This is what it would take to find out.
A note about a question, not a result. Nothing here has been built or measured.
The question
Ask a model for a story premise in free form and you get prose — often fluent, often specific in the wrong places, and usually hard to revise because there is no seam to cut along.
The alternative worth testing is to force the output through the six journalistic axes: who, what, when, where, why, how. Six slots, filled one at a time, before any of it becomes a sentence.
The hypothesis is that the constraint does two things a free-form prompt does not:
- It makes gaps visible. An empty why is obvious in a slot and invisible in a paragraph.
- It makes revision local. Changing where should not require regenerating the premise.
It is a hypothesis. It could easily be wrong in the other direction — six slots may push the model toward generic, separable answers precisely because they are separable, and a good premise may be exactly the thing whose parts do not come apart cleanly.
What “more usable” would have to mean
The word doing the most work here is usable, and it currently means nothing measurable. Before this is worth running, it needs an operational definition. Candidates:
- Revision distance — how much of the output survives to a draft a writer keeps, versus how much gets thrown away.
- Gap detection — whether a reader can identify what is missing from the premise faster in one format than the other.
- Independence — whether changing one axis leaves the other five intact, or cascades.
- Sameness across runs — whether the structured version collapses toward the same handful of answers when prompted repeatedly.
That last one is the risk I would want measured first, because it is the failure mode that would make the whole idea not worth pursuing.
What would have to be done
A comparison needs the same seeds through both paths, blind rating by someone who did not write the prompts, and enough runs that a difference is not one lucky sample. None of that exists yet. This note is here so the question is written down before it gets answered, rather than after.