The Generative AI Learning Penalty: Evidence from Chinese Secondary Education
Using 30 months of panel data on 26,811 Chinese students in grades 7-12, we study how generative AI affects homework productivity and learning. The data combine monthly closed-book exams, high-school and college entrance exams, and homework scores and completion time across nine subjects. We exploit staggered AI adoption in a difference-in-differences design. AI adoption raises homework scores by 18% and reduces completion time by 30%, but lowers monthly exam scores by 20% within six months. High-stakes entrance-exam scores fall by 18 and 24%, with the full penalty emerging only after about two years. The losses are largest in social science subjects, followed by STEM and languages, and are especially large for junior students, high-achieving students, and boys. The learning losses are concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores. AI users who maintain similar homework completion time as non-AI users experience small learning losses.
Edit: moving my comment up here
Just in case as it’s formatted a bit weirdly
X-Axis: Homework scores
Y-Axis: Exam scores


Given OPs axis definitions.
The max homework score for nonai is 115. But there are AI users scoring 115+? I’m confused why there are no better scores above 115 for nonai. Is that data insignificant? Are students with scores higher than 115 just being flagged as AI? It makes the data look unreliable.
I’d also really like to know if the data range that isn’t comparable 115+ is at all meaningful data. Because that’s where there is a major dropoff. There is obviously a dropoff of significant before 115. But, just from looking at how the data is being presented with this massive dropoff but with zero data for nonai.
I’d be interested in knowing what fraction of the total population is actually being represented in the 115+ dropoff. The way it’s presented it could literally be like 10 students.
This is not a defense of AI use. But more a criticism of how the data is being presented.
Also, Any student could be using AI as a resource similar to how I would use solutions manuals or previous tests to study back in my education. The good students aren’t gonna copy it verbatim and get flagged. Their going to get the right answer, and use it to learn the process so they can present their own work and actually learn the material.
AI is trash. But this is nothing different than before AI when students would copy the solutions manuals blindly or their peers work. Those students have always existed. AI just makes it slightly more accessible. But, really only slightly. And, honestly, Chegg was just as easy to copy back in the day.
Edit: This graph isn’t in the paper that OP linked. So, probably why it’s presented so badly. I’m assuming it’s from some click bait article then. But the paper clarifies that this is an actual survey of students. So it’s not just from flagged AI homework.
The paper itself in section 5 supports what I hypothesized above. It’s not really about AI making students dumb. It’s about making it easier for students that already don’t want to learn to finish the homework. Same as the people that would copy solutions manuals.
Quote from section 5.
The paper doesn’t argue it from what I read. But I’d also argue their is some bias in the survey. The students scoring well in both exams and homework are less likely to consider their use of AI as meaningful enough to answer “used AI for homework”. It’s a problem with the survey. The survey is asking the students that actually learned the material to attribute their homework to AI alone in the same way the “copy and paste” AI users would.
The survey would probably benefit from having No AI, Some AI, All AI as it’s responses. Even in anonymous surveys people make these own interpretations of the questions in their head. And a student that used AI as a resource to learn is very likely to just choose “No AI” when presented with questions of how they got their final answers to homework. Because in their head they are thinking of the students that copy and pasted AI responses and think “I’m not like that”.
It’s likely why the double high scoring AI user sample is so low (paper says this). The survey is not allowing the response for this set of users to categorize themselves as using AI without feeling like they are like the copy and paste students. So those people are likely just self categorizing as “no ai” because the survey doesn’t allow them to distance themselves from the other population. Surveys are hard to write. Even ones where the users know it’s anonymous. Good students that use AI will see themselves as good students and the copy pasters as bad students. Human emotion plays a role and most people will not self categorize themselves further negatively than they feel they should be. They are more likely to select the imperfect category that is more positive.
Bad students don’t care though. They’ll admit to AI use. They already don’t care enough to learn the material. They have no reason to lie. So they’ll select “used AI”.
I disagree that it’s no different to before AI. These days you can prompt for the whole paper to be written for you.
The values aren’t a score. They normalized the average score to 100 for non-AI in both homework and exam. Then noted the deviation. People using AI to get very high scores (130% of the average) were tank the test (40% of the average) six months later.
I disagree. Bad students could fear being punished for using AI as it’s clearly cheating.