No Child Wrote This
Inside the Department for Education’s AI experiment that used fictional Year-6 writing portfolios in a live exercise for approving statutory writing moderators...
This story needs one correction before it needs a headline: the Department for Education did not secretly switch on an AI system this week. The Standards and Testing Agency, the DfE body responsible for national curriculum assessments, told moderation managers on 10 July 2025 that it had already spent more than 18 months researching whether large language models could help create the writing collections used in moderator standardisation. Schools Week reported the planned trial in October 2025.
What changed on 10 August 2026 is the paper trail. A newly published Algorithmic Transparency Record for STA’s “LLM - KS2 Writing Sample Generator” identifies the current model as GPT-5, describes a production process built around heavy human editing and external moderation-manager review, says the annual cost has fallen from more than £100,000 to below £5,000, and explicitly acknowledges that raw LLM output can fail to reflect atypical language patterns associated with neurodivergent and non-native-English-speaking pupils.
Read alongside the live 2025/26 standardisation materials, the result is stranger than the generic phrase “AI in education” suggests. England’s testing agency created complete fictional pupil portfolios for a live moderator-approval exercise. The portfolios were presented as handwritten schoolwork, contained realistic spelling and grammar weaknesses, and were accompanied by official commentaries evaluating the fictional pupils’ spelling, composition and handwriting. Yet the exercise itself states that no Year-6 pupil produced or handwrote any of the scripts.
This was not AI grading children. It was AI-assisted synthetic children being used to test whether adults were qualified to grade children.
The live exercise
To moderate statutory Key Stage 2 writing, prospective moderators must pass an STA standardisation exercise. The official standardisation process provides three exercises, allows a maximum of two attempts and grants STA approval to moderate once a participant successfully passes one. Exercise 3 for the 2025/26 cycle ran from 9 to 27 February 2026 and was designated as the LLM pilot.
The Exercise 3 pupil-script pack hosted by Norfolk County Council is unambiguous about the stakes. STA tells participants that the materials were developed in-house with LLM support; that the written material was generated by LLMs and substantially edited by human experts; that none of the scripts was produced or handwritten by a Year-6 pupil; and that the outcomes of the exercise would nevertheless be treated as valid in determining whether participants had the knowledge required to moderate during 2025/26.
That makes Exercise 3 an operational accreditation route, not a laboratory curiosity. Anyone passing it received the same kind of STA approval required to moderate real statutory teacher assessments.
The synthetic pupils
The exercise contains three fictional collections spanning different writing standards. Across them, the invented pupils write about subjects including Pompeii, the Benin Bronzes, climate change, Rosa Parks and the Bristol bus boycott, Hannah Hampton, mental health, Beowulf and fictional narratives. The published portfolios are rendered as convincing pieces of schoolwork rather than clean blocks of generated text.
Moderators are instructed to approach the collections as they would real evidence: assume the writing is independent; assume edits are the pupil’s own; rely on the strongest handwriting where presentation varies; and, where spelling evidence is insufficient, assume other qualifying evidence exists. The same instructions also state that the exercise contains no collection for a pupil deemed to have a “particular weakness”.
The agency’s accompanying official commentaries then evaluate the fictional artefacts against the statutory framework. In one collection the commentary discusses deliberately pupil-like spellings including “facsinating”, “attatched” and “polititians”, before judging the handwriting as legible when writing “at speed”. Another fictional pupil is likewise credited with maintaining clear joined handwriting during extended writing.
That wording creates an obvious provenance question: who or what created the handwriting? The material we found does not establish whether an adult handwrote it, whether a digital handwriting system rendered it, whether it was assembled another way, or how any claim about writing “at speed” was operationalised. The safe conclusion is not that the handwriting was AI-generated; the documentation does not prove that. The safe conclusion is that STA formally assessed handwriting characteristics in portfolios it simultaneously disclosed had not been handwritten by Year-6 children.
The policy paradox
DfE’s own 2025/26 Key Stage 2 teacher-assessment guidance is strict about the evidential status of AI-assisted pupil work. Writing edited or rewritten following AI feedback should be excluded from the collection used to establish a child’s standard, and using writing produced wholly or partly by a large language model to support a statutory teacher-assessment judgement can trigger a maladministration investigation.
That does not make DfE’s pilot automatically contradictory. A child’s portfolio is supposed to demonstrate the child’s own ability; a standardisation exercise is a deliberately engineered test instrument designed to determine whether a moderator understands the framework. Those are different functions.
But the difference sharpens the central validity question rather than removing it. If authentic authorship is essential when judging the pupil, what evidence demonstrates that a synthetic substitute is sufficiently authentic when judging the person who will judge the pupil?
AI-assisted writing cannot be used to prove what a child can write. AI-assisted fictional writing was used to prove that a moderator could judge what a child can write.
What the new GPT-5 record actually says
The 10 August transparency record calls the system “LLM - KS2 Writing Sample Generator” and says LLMs are used to produce English writing resembling Year-6 output at different assessment standards. It identifies “Chat GPT 5” as the model version and describes the model name as “Chat GPT 5 - self hosted”. It also names OpenAI in the third-party role field.
But there is an important limit to what this proves. The 2025 pilot documents refer generically to LLMs and do not identify the exact model used to generate the February 2026 Exercise 3 portfolios. The August record establishes GPT-5 as the current named system; it does not, by itself, prove that GPT-5 produced the live February material. DfE should disclose the exact model and version used for that exercise.
According to the record, an experienced teacher-assessment researcher prompts the model using Year-6 topics and genres and the statutory Teacher Assessment Framework. The resulting text is then heavily edited. Further internal researchers review it, after which around 20 local-authority moderation managers provide external quality assurance and feedback on authenticity. The current implementation is described as producing three pupil collections for one standardisation exercise.
A £100,000 process reduced to below £5,000
The economic case is stark. The government record says the previous annual production method cost more than £100,000 and that the new approach costs less than £5,000. It says the former model relied on a supplier sourcing pupil scripts and producing exercises, while the 2026/27 and 2027/28 arrangements are expected to cost around £5,000 a year for quality assurance on two full exercises, with one additional exercise using scripts originating under the previous supplier arrangement.
That is a headline reduction of more than 95 per cent. It may also be entirely rational: sourcing enough genuine pupil work at precise assessment boundaries is difficult, expensive and, as Ofqual has previously documented, capable of producing weak test material.
But the record does not provide enough accounting detail to establish whether the old and new figures are strictly like-for-like. It does not break out internal DfE staff time, model or hosting costs, infrastructure, the labour involved in substantial rewriting, or the cost basis of the roughly 20-person external QA process. The saving is therefore a documented government claim; the comparability of the two totals remains a legitimate question.
The representation problem is in DfE’s own record
The most significant risk statement comes from STA itself. In its published risk section, the agency says LLM-generated texts often exclude atypical vocabulary and sentence structures that may be used by neurodivergent pupils or pupils who are not native English speakers. Its mitigation is intensive human review and editing.
Yet the same record lists the “Impact assessments” field as “n/a”. And the live February exercise explicitly says it contains no collection involving a pupil deemed to have a particular weakness. Neither fact proves discrimination, and no evidence we found shows a pupil was wrongly assessed because of the pilot. They do, however, make the unanswered validation questions more important: how was representativeness tested, and what evidence exists for SEND, neurodivergent and EAL writing patterns before synthetic portfolios were used in a live approval route?
A transparency record that creates new questions
The record contains several entries that deserve clarification rather than speculation. It says “No” under third-party involvement, then immediately identifies “Chat GPT 5” and gives OpenAI as the third-party role. It describes the model as “self hosted”, while offering no public explanation of what that means operationally in this deployment. The system-architecture entry also contains the unexplained string “flexos.work.-5”.
The “Model performance” section is notable for what it does not contain. Instead of quantitative performance measures, it describes the human editing and review chain. The government’s own Algorithmic Transparency Recording Standard guidance says performance reporting should use measures appropriate to the tool and should describe bias or fairness evaluation over relevant subgroups where applicable. The STA entry gives process detail, but no public accuracy, agreement, equivalence or subgroup-performance figures.
Its lifecycle status is also listed as “Public Beta”, despite the technology having already gone through research, a trial and operational use in a live approval exercise. “Public Beta” may have a perfectly mundane administrative meaning inside the ATR system. But the label deserves definition when decisions carrying real moderator approval were already made.
The old system had problems too
There is a strong government defence, and any serious account has to include it: authentic pupil scripts were not a flawless gold standard.
In 2018, Ofqual recommended that STA revisit the design of the KS2 writing standardisation test after concerns about its authenticity. By 2022, the regulator recorded that the Australian Council for Educational Research was using real pupil writing to build three exercises, but fewer than half of moderators passed Exercises 1 and 2. STA attributed part of the problem to difficulty sourcing suitable pupil exemplars, which could leave materials sitting awkwardly on assessment boundaries. Exercise 3 had to be replaced with prior-year material, creating its own risk that some moderators would already have seen it.
The following year, after additional quality assurance, Ofqual reported 2023 pass rates of 87 per cent, 74 per cent and 62 per cent across the three exercises and said the strengthened process had produced more robust material.
This history matters. STA was trying to solve a real test-development problem: finding enough authentic pupil work that cleanly represents the statutory standards without ambiguity. The question is not why it looked for a new method. The question is whether the evidence demonstrating that the synthetic replacement is equally valid has been made public.
DfE promised an evaluation
The trail begins before the live exercise. In its July 2025 communication to moderation managers, STA said participant feedback on accuracy and confidence would be essential in evaluating the pilot. Its accompanying pilot FAQ said the agency had already run four research exercises, offered all local authorities opportunities to participate, distributed trial collections and surveyed moderators. Early participants had raised questions about whether some topics, vocabulary and sentence structures felt appropriate for Year-6 pupils, and STA said later collections were revised in response.
The FAQ also says findings from the research and the 2025/26 pilot would inform decisions about future use. The live standardisation timetable even scheduled a feedback-survey window immediately after Exercise 3.
Now compare that with the August 2026 record. STA is planning LLM-assisted production through 2027/28 and expects to decide in spring 2027 whether to continue or return to procurement. Yet the transparency record does not publish the basic quantitative results a reader would naturally expect from the pilot: the number of participants in Exercise 3, its pass/fail rate, the opt-out rate, comparison with Exercises 1 and 2, moderator confidence results, inter-rater agreement, equivalence measures, or subgroup validation.
As of 12 August 2026, Thom Aster could not locate a public pilot-results report containing those measures across the official records and publications reviewed for this investigation. That is not proof the analysis does not exist internally. It is precisely why the analysis should be published.
Moderators had already raised the issue
The concern is not purely theoretical. A June 2026 study in the British Educational Research Journal, based on survey data from 33 local-authority lead moderators, examined perceptions of consistency across the national moderation system. Among the qualitative concerns recorded in the study was an objection from one respondent to using AI-created material rather than real pupil work.
One respondent is not a profession-wide revolt, and it would be misleading to portray it as one. What it does establish is narrower: the authenticity issue had surfaced inside the moderation community before the new transparency record appeared.
How many moderators did this actually affect?
The new transparency record says around 2,000 moderators pass standardisation each year. That figure must not be misreported as “2,000 moderators were certified by AI”. There are three exercises and participants can pass before reaching Exercise 3. Ofqual’s 2025 annual report says the majority of moderators are normally recruited by the end of the second exercise and records 2,395 moderators recruited that year with a 70 per cent overall pass rate.
The number who specifically obtained approval through the LLM-based February 2026 Exercise 3 is therefore only a subset of the annual total. The public documents reviewed here do not reveal that number. It is one of the simplest facts DfE could release.
Those approvals still matter. Under the statutory moderation arrangements, local authorities externally moderate at least 25 per cent of maintained schools and at least 25 per cent of academies and participating independent schools. DfE also uses KS2 teacher-assessment data in school-performance measures and shares school-level information with trusts, local authorities and Ofsted.
What this investigation proves — and what it does not
Proven by published records:
STA researched LLM-generated KS2 writing for more than 18 months before announcing the live pilot.
Exercise 3 in February 2026 used LLM-generated, substantially human-edited fictional pupil portfolios.
No Year-6 child produced or handwrote the exercise scripts.
The exercise results were valid for deciding whether participants had the knowledge required for STA approval to moderate.
The current August 2026 transparency record identifies GPT-5, names OpenAI, claims annual production costs fell from more than £100,000 to below £5,000 and acknowledges representational weaknesses in raw LLM output.
The published transparency record contains no quantitative model-performance or pilot-equivalence metrics.
Not proven by the material reviewed:
That GPT-5 specifically generated the February 2026 portfolios.
That the handwriting itself was generated by AI.
That any child received an incorrect statutory assessment because of the pilot.
That the £100,000 and £5,000 figures are fully like-for-like.
That moderators broadly oppose the system.
The questions DfE should answer
Which exact model and version generated the February 2026 Exercise 3 portfolios?
What does “Chat GPT 5 - self hosted” mean in this deployment, and through which infrastructure or service is the model accessed?
Why does the transparency record state “No” for third-party involvement while identifying ChatGPT 5 and OpenAI in the adjacent third-party fields?
What is the unexplained “flexos.work.-5” text in the system-architecture field?
How many moderators attempted Exercise 3, how many passed, how many failed and how many local authorities opted out?
What were the 2025/26 pass rates for Exercises 1, 2 and 3?
What quantitative evidence established that synthetic portfolios were sufficiently valid and reliable compared with portfolios sourced from real pupil work?
What were the results of the four pre-pilot research exercises and the promised participant confidence/accuracy survey?
What validation was carried out for SEND, neurodivergent, EAL and other atypical language patterns?
Why is “Impact assessments” recorded as “n/a” despite the stated representation risk?
Who produced the handwriting used in the fictional portfolios, by what method, and how were handwriting characteristics such as writing “at speed” represented?
Were spelling and grammar weaknesses generated by the model, inserted through prompting, introduced by human editors, or a combination?
Did an LLM contribute only to the pupil scripts, or also to any official commentaries or scoring materials?
What costs are included and excluded from the claimed reduction from more than £100,000 to below £5,000?
What evidence informed the decision to continue LLM-assisted production through 2027/28 before the spring 2027 continuation/procurement decision?
The real story
This is not a story about a chatbot marking SATs. It did not. It is not evidence that children were secretly assessed by GPT-5. The documents do not prove that either.
It is more precise — and more consequential.
England’s testing agency faced a genuine problem creating reliable moderator-standardisation material from authentic pupil work. It spent more than 18 months testing an alternative, then used LLM-assisted fictional children in a live approval route. The resulting portfolios were realistic enough to contain synthetic-looking errors, handwriting and whole classroom contexts, and the official commentary treated those artefacts as evidence against the same statutory framework used in real moderation.
Afterwards, the government published a transparency record saying the current system uses GPT-5, costs a fraction of the previous process and still requires heavy human correction because raw model output can miss the language patterns of some real pupils.
What the public record still lacks is the part that should settle the argument: the measurable results showing that the synthetic test works as well as the authentic one it is beginning to replace.
Until those results are published, the key question is not whether AI can imitate a Year-6 pupil.
It is whether the state can demonstrate that its imitation is good enough to certify the people entrusted with judging the real thing.
Support This Work
If you would like to support my work and help keep me safe, all support is greatly appreciated.
Bitcoin (BTC) - bc1qevvy4y7ph5nxhsux0j6llfjepn239r52rakgcy
Ethereum (ETH) - 0x26F16D2D4d3dE1ab332deeE9d7DECA6B90654717
XRP - r4UDiUq5U8cQv5gq2zBdimtoMVxxjqzcYH
Solana (SOL) - 22SruEARKvXAKn8TzZZ1dBogQXRYyWkuYA4KYZWa4vAS
Dogecoin (DOGE) - DBjnCWtW1r7bnTorAocE5ypsgA1h7kQ1PD
Cardano (ADA) - addr1q8u5ld73yzv0ep9vguaz2zepdlv3ldpdr23ytukpk7lt9r8ef7mazgycljz2c3e6y59jzm7er76z6x4zghevrda7k2xq43ad8d
Sources
Department for Education / Standards and Testing Agency — Algorithmic Transparency Record: LLM - KS2 Writing Sample Generator (published 10 August 2026).
GOV.UK — Teacher assessment moderation, standardisation and training process (2025/26 cycle).
GOV.UK — Key Stage 2 teacher assessment guidance (independent writing, AI and maladministration rules; moderation coverage; data use).
Norfolk County Council — KS2 standardised pupil writing collections.
STA — 2025/26 KS2 Standardisation Exercise 3 pupil scripts (LLM-assisted live exercise).
Standards and Testing Agency — 10 July 2025 communication to moderation managers on the LLM pilot.
Standards and Testing Agency — Standardisation exercises LLM pilot FAQs.
Schools Week — AI questions to be trialled in SATs moderator tests (6 October 2025).
British Educational Research Journal — Moderators’ perceptions of consistency in Key Stage 2 writing moderation across local authorities (18 June 2026).
Ofqual — Observations on the consistency of moderator judgements (2018).
Ofqual — National assessments regulation annual report 2022.
Ofqual — National assessments regulation annual report 2023.
Ofqual — National assessments regulation annual report 2025.
GOV.UK — Algorithmic Transparency Recording Standard guidance for public-sector bodies.

