Earlier this month, New York City announced that around six hundred thousand children would not be allowed to use generative AI this school year.
Every child from their earliest years through to eighth grade. The largest school district in America, two thirds of its entire enrolment. Companion chatbots banned right the way up to high school. Around forty existing approved programmes having their AI components switched off. Screen time capped at thirty minutes a day for eight to eleven year olds, and no individual devices at all for the youngest.
Teachers can still use it. It is the children who cannot.
The moratorium runs for one academic year, with a coalition of teachers, parents and experts reviewing the evidence as it goes. Worth noting that the same city banned ChatGPT from its school devices in January 2023 and reversed the decision within months, so this is an institution that has already changed its mind once in public.
The line that needs thinking about came from the mayor. Mamdani said he had yet to see a study showing AI is beneficial for children in elementary and middle school, with the exception of research sponsored by the very corporations that stand to profit.
I have been reflecting on that conclusion, and I have found it harder to dismiss than I expected to.
Because Singapore has spent the same period going the other way, building AI literacy into the curriculum as national policy, on the reasoning that a generation growing up without fluency in the defining technology of their working lives will carry a disadvantage that compounds.
Two serious education systems. Same year. Opposite conclusions. Neither of them staffed by fools.
I talk often about what I call the human edge, but a great deal of what I have to say has been about adults.
Leaders and their judgement. Teachers and their expertise. What a professional can see in a piece of AI output that the machine cannot see in itself. I have argued that AI is not replacing teachers but supporting them, and I still believe that.
But what happens if we follow that argument down to a seven year old?
Let’s look at a working paper published by the Centre for Economic Policy Research in June this year. Strömberg, Lei and Wu, tracking 26,811 secondary pupils in a county in central China across thirty months of administrative data.
The headline is stark. After generative AI became widely available to those pupils, homework scores rose by eighteen per cent and homework completion time fell by about a third. Monthly closed book exam scores fell by twenty per cent within six months. High stakes entrance exam scores fell by between eighteen and twenty four per cent, with the full effect only becoming visible after roughly two years.
So, homework got better. But learning got worse. Same children, same period, same cause.
The authors are careful about the mechanism, and this is the part that matters. The losses were concentrated among roughly eighty per cent of AI users whose behaviour looked like outsourcing, identified by unusually short completion times combined with high scores. Pupils who used AI but kept spending similar time on the work showed only small losses.
So it is not the presence of the tool. It is what the tool replaces.
Two further findings worth noting:
There is a clear dose response. Pupils reporting up to an hour of AI use a week showed around a five per cent decline. Pupils using it five hours or more showed thirty.
And the losses were larger for younger pupils. Lower secondary was hit around forty per cent harder than upper secondary, a pattern the authors link to older students facing more supervised coursework and more restriction.
That is the trend in the report. And the youngest children in that study were twelve.
The explanation is not mysterious, and it is one every primary teacher will recognise.
AI does not know what is right. It produces what it predicts is most likely. A lot of the time likely and right turn out to be the same thing, which is precisely what makes the gap dangerous, because the mistakes arrive with total confidence and no change in tone.
Catching that requires already knowing what right looks like. A teacher can take something AI has produced, see immediately that the third paragraph is wrong, and fix it. That is not a technology skill. It is subject knowledge, built slowly, over years, through exactly the kind of effortful work that a chatbot removes.
Give a child the tool before they have the knowledge and you have not handed them a shortcut to expertise. You have removed the route to it. The struggle was never an obstacle sitting in the way of the learning. The struggle was the learning.
That argument is strong, and the Chinese data is the first large scale evidence that it holds at population level rather than just sounding right in a staff meeting.
It is also, I think, not quite the whole picture.
Because there is a difference between a child using AI and a child being taught with it, and a ban does not really distinguish between the two.
A child alone with a chatbot, offloading the thinking, is the thing the evidence should worry us about. A teacher using AI to generate three versions of a task pitched at different levels, deciding which child gets which, and adapting on the spot, is something else entirely. The technology is identical. The pedagogy is opposite. Only one of them removes the struggle.
And there is a second distinction, which is about what we mean by using AI at all.
Being handed a tool and told to get on with it is not the same as being taught how it works, what it does, why it is confidently wrong, and how you would know. Those are different lessons, and the second one may be among the most important things we could teach a child about the technology that will shape their working life.
Singapore’s position, at its strongest, is not that children should be left alone with chatbots. It is that AI literacy is now part of being educated, and a system which avoids the subject is postponing something rather than preventing it.
New York’s position, at its strongest, is not that AI has nothing to offer. It is that the burden of proof sits with the people selling it, and that until somebody can demonstrate a benefit to a nine year old that is not funded by the company supplying the software, the default should be no.
Both of those are defensible positions to take. What I cannot defend is the version of this debate where the only question is whether we are for it or against it.
Here is what is missing from all of it.
The Chinese study covers twelve to eighteen year olds in a different education system. New York has made a decision about six hundred thousand primary and middle school children. Singapore is building a curriculum for all ages. And nobody, anywhere, has properly tested what happens below twelve.
Primary is not a smaller version of secondary. Struggle works differently when a child is learning to read rather than revising for an exam. Memory formation is different. The child’s relationship with the adult in the room is different, and so is their relationship with a machine that talks back.
The mayor of New York said he wants to see evidence that is not funded by the people selling the software. That is a completely reasonable thing to want. It is also a request nobody has properly answered for primary-age children, because the research does not exist. Yet.
But that does raise a question. If the evidence does not exist, who is going to make it?
Perhaps we need a serious attempt at this would look like. Something along these lines, could be useful:
You would need more than one arm running at once. A group where children attempt, reason or create first, with AI introduced afterwards to explain and challenge rather than produce. A group where the tool is present from the start but guardrailed to prompt rather than hand over finished answers. And a group using none of it, because without a baseline you are describing rather than testing.
Matched classes, sitting inside normal curriculum time rather than bolted on as an extra. A term at minimum, with a review point built in.
But the part that would actually matter is the measurement, and this is where I think most claims about AI in education fall down at the moment.
You would have to measure two different things and not add them together. What a child can do unaided several days later, which is retained capability. And what they can accomplish with the tool that they could not manage alone, which is augmented capability.
Almost everything written about AI in classrooms measures the second and reports it as though it were the first. That is precisely the trap the Chinese data exposed. Homework improved while learning declined, and any single aggregate conclusion would have shown nothing but good news for six months.
And there are conditions I would want met before I would be comfortable with any new research, whoever was running it.
Written parental consent specific to the project rather than folded into something else. A safeguarding review of the actual tool, starting with whether its own terms of service permit children of that age to use it at all, which is not a detail and rules a good deal of it out immediately. An explicit stopping rule if any group shows disadvantage partway through. Ethics oversight. A named data controller. And a genuine commitment to publish a finding of no effect just as readily as a finding of harm, because a project that can only produce one answer is not research.
None of this needs to be enormous to be worth doing. A single term, a handful of matched classes, one task type in literacy and one in numeracy would be a start, and a start is what is missing. It would not settle the question. It would be the first clear look at an age group nobody has examined, and it would give the schools running it something better than opinion to work from.
That scale is within the gift of many groups of schools. A federation, a trust, a local authority, a small jurisdiction. It needs a few matched classes, teachers willing to keep a log, governors prepared to ask hard questions, and somebody willing to publish whatever comes out of it.
What it needs most is a leadership team willing to be the ones who look. That is not a resource problem. It is a nerve problem, and it is exactly the kind of moment where innovative and brave leadership earns its name.
Whether we end up doing something like that here, I honestly do not know yet. It is a considerable undertaking for three primary schools in the middle of a school year, and the conversations it would need are conversations, not formalities.
But somebody has to look at primary, and small jurisdictions are better placed for this than they usually get credit for. Small enough to coordinate properly, and large enough for the findings to mean something.
The question shouldn’t be whether children use AI. They do. It is whether they have the knowledge to judge what it gives them, and whether an adult with expertise is standing between them and the machine.
Knowledge first, always. Not because knowledge is old fashioned, but because judgement is made of it. A child who cannot tell whether an answer is right is not empowered by a tool that produces plausible answers at speed. They are exposed by it.
Supervision as a matter of design rather than filtering. A teacher choosing what the tool is for, in this lesson, for this child, with a reason they could explain.
And a distinction we are not making often enough, between learning with AI and learning about it. The first needs handling with enormous care in a primary school. The second is becoming difficult to justify leaving out.
New York has bought itself a year to think. Whatever you make of the decision, that is not a reckless way to make one.
We could use a year like that ourselves. And if somebody spent theirs producing evidence rather than waiting for it, that would be a year well spent.
