A striking result from an informal test

The claim that people prefer AI-written stories comes from an online experiment run in August 2025 by fantasy author Mark Lawrence. Readers were shown eight short fantasy pieces, some written by published authors and others generated with AI, then asked to rate their quality and judge whether each came from a human or a machine.

Lawrence reported that the aggregate audience result correctly identified three stories, misidentified three and produced no clear verdict on two. AI-written entries also performed better on average in the quality ratings, with one AI story receiving the highest score. Those findings are notable because they challenge a widespread assumption that machine-written fiction is immediately recognisable or necessarily less enjoyable.

But the experiment’s own framing matters. Lawrence described it as not scientifically rigorous. It was a voluntary quiz promoted to his online readership, rather than a representative, controlled study of the general public. The pieces were also very short, in a genre familiar to the host and likely to many participants. The headline conclusion is therefore better understood as evidence that AI can be persuasive in a specific form of short fiction, not proof that most readers everywhere prefer machine-authored literature.

Why flash fiction suits the comparison

Flash fiction is an unusually favourable format for contemporary language models. A story of a few hundred words can establish a premise, create atmosphere and deliver a twist without requiring the sustained consistency of a novel. It limits the model’s need to track character development, chronology, setting details and themes across many chapters.

That does not make the result meaningless. Readers evaluate what is on the page, and a short form still requires control of language, pace and expectation. Yet the format makes it difficult to extend the conclusion to longer works. A reader’s response to a polished vignette may differ substantially from their response to a full-length book, where voice, structure, emotional development and continuity carry more weight.

The human comparison also deserves care. Accomplished novelists are not automatically specialists in flash fiction, and a one-off prompt-based test cannot capture the revision, research and editorial processes involved in professional publishing. It measures the appeal of individual texts under blind conditions, rather than the broader contribution of an author to a work.

Research offers a more mixed picture

Academic evidence does not support a single universal verdict on AI fiction. A 2025 study published in Humanities and Social Sciences Communications compared stories produced by ChatGPT with stories written by non-professional human participants from the same prompts. It found no meaningful overall difference in enjoyment or appreciation. However, readers reported greater narrative transportation — the sense of becoming absorbed in a story world — for the human-written texts.

The researchers linked some of that difference to linguistic features. Human stories contained more personal pronouns, which were associated with stronger transportation. This suggests that surface fluency and immediate enjoyment are not the only dimensions of storytelling. A text can read smoothly and still be less effective at creating intimacy or sustained immersion.

Other research shows that labels themselves influence reactions. In an incentivised experiment involving 654 participants, a story described as AI-generated received lower subjective assessments than the same content presented as human-written. Yet participants’ reading time, willingness to pay and willingness to work to continue reading did not significantly differ. The result points to a gap between stated preferences about authorship and behaviour when people encounter a story.

Together, these studies suggest two separate questions. One is whether readers like a text when its origin is hidden. The other is whether they value it in the same way once they know how it was made. The answers may diverge.

Detection is a skill, not a universal instinct

The Lawrence test is also consistent with a broader problem: identifying generated text is hard when the material is short, edited or written to a familiar genre convention. Readers can mistake polished human prose for AI and mistake generic AI prose for human writing. Confidence in a judgement does not guarantee accuracy.

However, it would be inaccurate to say that humans cannot detect AI writing at all. Research presented at the Association for Computational Linguistics found that frequent users of ChatGPT for writing tasks were highly effective at distinguishing AI-generated from human-written non-fiction articles. A group of five experienced annotators made only one error across 300 articles in the study.

That finding does not directly settle the question of fiction, and it should not be treated as a licence to make high-stakes authorship accusations from stylistic impressions alone. It does show that familiarity with language models can improve judgement, while also underlining that detection performance depends on the material, the models involved, the amount of text available and the experience of the reader.

What the test means for publishing

The most important implication is not that human writing has become obsolete. Rather, AI is increasingly capable of producing short passages that satisfy common reader expectations for genre fiction: clear prose, an arresting setup, recognisable emotional cues and a complete ending.

That raises practical questions for publishers, platforms and readers. If quality ratings alone cannot reliably reveal origin, disclosure policies become more important. Readers may reasonably care about whether a work was generated, assisted, edited or authored by a person, even when they enjoy the result. Clear labelling can allow audiences to make that choice without asking them to solve an unreliable detection puzzle.

For writers, the pressure is likely to fall most heavily on easily commodified forms of text: blurbs, short promotional fiction, rapid serial content and formula-driven genre material. Longer works may remain a more demanding test because they require coherent planning, intentional revision and a distinctive relationship with readers over time.

The online quiz does not establish that AI is better than authors, nor that readers have stopped valuing human creativity. It does establish something narrower and more consequential: in a blind test of short fantasy stories, machine-generated writing could compete for attention and approval. That makes authorship less visible at the point of reading — and makes transparency, editorial judgement and informed audience choice more important.

Sources