The Generative AI Learning Penalty: Evidence from Chinese Secondary Education

Using 30 months of panel data on 26,811 Chinese students in grades 7-12, we study how generative AI affects homework productivity and learning. The data combine monthly closed-book exams, high-school and college entrance exams, and homework scores and completion time across nine subjects. We exploit staggered AI adoption in a difference-in-differences design. AI adoption raises homework scores by 18% and reduces completion time by 30%, but lowers monthly exam scores by 20% within six months. High-stakes entrance-exam scores fall by 18 and 24%, with the full penalty emerging only after about two years. The losses are largest in social science subjects, followed by STEM and languages, and are especially large for junior students, high-achieving students, and boys. The learning losses are concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores. AI users who maintain similar homework completion time as non-AI users experience small learning losses.

Edit: moving my comment up here

Just in case as it’s formatted a bit weirdly

X-Axis: Homework scores

Y-Axis: Exam scores

    • assaultpotato@sh.itjust.works
      link
      fedilink
      arrow-up
      8
      arrow-down
      1
      ·
      1 month ago

      This is conjecture because I can’t access the full paper at this time, but based on:

      The learning losses are concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores.

      I’m guessing this is an artifact of “students under higher pressure to perform are more likely to use AI more thereby lowering their scores”. Given Chinese patriarchical cultural biases, I’d imagine boys are under higher pressure on average, thereby leading to greater usage.

      So perhaps it’s actually the second one, as the abatract is unclear if they’re studying performance relating to usage or not. Without checking methodology, it’s unclear if they’re controlling for usage.

    • THB@lemmy.world
      link
      fedilink
      arrow-up
      4
      ·
      1 month ago

      From other things I’ve read, men in general are more likely to use AI, not just in China. So I would assume the first.

    • assaultpotato@sh.itjust.works
      link
      fedilink
      arrow-up
      4
      ·
      1 month ago

      I suspect it’s the second but now that you mention it, it is a bit ambiguous. I’m assuming the study would control for usage to compare outcomes across similar usage metric cohorts.