What the pilot actually tested

A recent UK pilot has been easy to misread as another example of artificial intelligence marking children’s work. It was not that. The Standards and Testing Agency (STA), the body responsible for national curriculum assessments in England, used large language models to help create material for one exercise that assessed the judgement of human moderators.

These moderators have an important quality-assurance role. At the end of Key Stage 2, when most pupils are aged 10 or 11, teachers make statutory judgements on English writing. Local authorities then moderate a sample of schools to check that teachers are applying the national framework consistently. Before they can carry out this work, moderators must pass a standardisation exercise.

For the 2025–26 academic year, the STA ran three such exercises. The third, available from 9 to 27 February 2026, included an optional pilot using material developed with the assistance of large language models. The agency said the purpose was to explore whether the technology could address a practical supply problem: finding enough suitable real pupil scripts while reducing the cost of producing exercises.

The distinction matters. The generated writing was used to examine whether adults could apply an established assessment framework. It was not used to make a statutory judgement about an identifiable child, nor did the published guidance describe an AI system making final decisions about pupils’ attainment.

Why authentic scripts are difficult to source

Standardisation depends on examples that sit near assessment boundaries. A straightforwardly strong or weak piece of writing may be useful for teaching, but it does not necessarily reveal whether a moderator can make difficult decisions consistently. The most valuable cases are often those with a mixed profile: a script may meet many expectations for sentence construction, vocabulary and organisation while falling short in another crucial aspect of the framework.

Building a bank of such materials from schools is labour-intensive. It requires suitable work to be collected, anonymised, reviewed and matched to detailed commentaries explaining the expected judgement. It may also be difficult to obtain examples that expose the precise areas of disagreement needed for a reliable test of moderation skill.

Generative AI can, in principle, produce many draft texts with specified characteristics. Designers could request writing with particular grammatical features, an uneven level of cohesion or evidence of a borderline attainment profile. This could make it faster to create candidate materials for human assessment specialists to select, edit and validate.

That is the strongest case for the pilot: AI as a production tool for assessment content, with the assessment authority retaining responsibility for what is ultimately used and how it is judged.

The validity challenge does not disappear

Efficiency alone cannot establish that a synthetic script is fit for assessment. A moderator’s task is rooted in real children’s writing, including its individual voice, uneven development, unexpected choices and the influence of classroom context. AI-generated prose may imitate some surface features of a pupil’s work while missing the patterns that make authentic writing educationally meaningful.

The risk is not simply that generated material could be too polished. It could instead become artificially constructed around rubric language. If a script is designed to display a set of features too neatly, moderators may learn to spot the intended signals rather than exercise the holistic professional judgement required in schools. A training item can be technically aligned with a framework yet still provide a less realistic test of moderation.

There is also a fairness issue. Standardisation exercises determine whether moderators are approved for that year’s work. Their content therefore needs to be defensible, consistent and adequately challenging. Any use of AI-generated material should be accompanied by rigorous human review, documented rationales for the expected judgement and evidence that the item performs comparably to one based on real pupil work.

The STA’s stated limitation is consequently significant: the material was to be used to evaluate moderators’ understanding of the English writing framework. It was not presented as a replacement for the framework, the commentary process or professional moderation.

A broader shift in what writing assessment measures

The pilot arrives as generative AI complicates the relationship between writing quality and writing ability. Language models can readily produce grammatical, coherent and conventionally structured prose. Research discussed by Educational Testing Service has found that AI-generated essays can perform especially well on language-related features, while human raters may identify weaker reasoning, evidence and depth of analysis than automated systems do.

This does not mean grammar and organisation have ceased to matter. They remain important components of clear communication. But it does mean that polished language is a less secure proxy for independent thought when a writer may have used AI support. Assessment systems must become clearer about whether they seek to measure composition produced unaided, a student’s capacity to develop and evaluate ideas, or the ability to use AI responsibly as part of a writing process.

For primary-school teacher assessment in England, the immediate issue is more limited. The moderator pilot concerns the preparation of training materials, not pupils’ permitted use of AI. Yet it raises the same underlying question: which human qualities must remain central when AI can generate plausible text on demand?

Keeping human accountability at the centre

A sensible approach separates two decisions that are often conflated. First, may AI assist in drafting administrative or training material? Second, may AI determine an educational outcome? The first can be explored through controlled pilots. The second requires a much higher standard of evidence, transparency and public accountability.

Current guidance from higher-education assessment specialists points in a compatible direction. It argues that criteria should focus on the learning a task is intended to evidence, rather than on speculative attempts to catch misconduct. It also advises reassessing the weighting given to activities that AI can easily automate when those activities are not the intended learning outcome.

For the STA, the next useful evidence would be practical rather than promotional: whether moderators found the AI-assisted scripts realistic; whether their results on such scripts aligned with performance on authentic examples; how much expert editing was required; and whether production savings survived the costs of validation and governance.

AI-generated writing may help relieve a genuine bottleneck in assessment administration. But the value of the experiment will depend on disciplined limits. Synthetic scripts can be useful instruments for testing human assessors only if the people who design, validate and interpret the exercise remain visibly responsible for its standards.

Sources