A two-year randomized experiment at Vrije Universiteit Amsterdam produced a result its own author had not expected: the students barred from using ChatGPT came last both years. Thibault Schrepel split his AI Law course into three groups — no AI, AI with no instruction, and AI with hands-on training — gave all of them the same task and the same exams, ran it with 66 students in 2024 and repeated it with 164 in 2025, and came out of it publicly abandoning the position he started with.
The task was narrow enough to grade cleanly. In teams of four or five, students had 20 minutes to improve a single provision of the EU AI Act. Work was scored on substance, clarity, proportionality and novelty. The first group could not use ChatGPT at all. The second received ChatGPT-generated editing suggestions embedded in the text and could keep using the tool, but got no guidance on how. The third was trained in writing legal prompts and in checking the model's suggestions for consistency and accuracy. Everyone then sat the same multiple-choice test and the same take-home exam, which required reworking a different provision of the Act.
What happened inside the 20 minutes is the most useful part of the study. The no-AI teams mostly made small edits — tightening wording, removing repetition — and only a few subgroups attempted a substantive change to the legal norm. After 10 to 15 minutes, many of them regularly ran out of ideas, an effect Schrepel calls idea depletion. The absence of a tool did push them into deeper discussion with each other.
The untrained AI group did the opposite. They largely accepted the model's suggestions without checking, on the reasoning that the new wording simply sounded better. Some swapped out "shall" and "individual" for ChatGPT's alternatives without showing any grasp of what that does to a legal text. In every subgroup, at least one misleading or legally superfluous word from ChatGPT survived into the final version. Only the trained group actually argued with the model, testing formulations and working through the substantive questions underneath the task.
Then the training advantage evaporated. In 2024 the trained group was clearly ahead, especially on the harder take-home exam. A year later all three groups scored roughly the same. Schrepel attributes the collapse to familiarity: students now use chatbots in daily life, so formal instruction added less in 2025 than it had a year earlier. Ethical use and legal responsibility, he says, still have to be taught.
One thing held across both years. The no-AI group finished last each time, which is the basis for his conclusion that banning the tool produces worse average results than allowing it. Whether the advantage comes from structured training or simply from hands-on experience remains unresolved.
The second group surprised him most. He expected the uncritical errors made in class to reappear in the exams. They did not, and that group scored slightly above the students with no AI at all. His most likely explanation is that students learn to spot the model's weaknesses through their own use, particularly when accuracy determines the grade. That forced a reversal: Schrepel had been convinced AI was useful only under structured instruction and otherwise belonged outside the classroom, and now says he was wrong. Skipping AI training costs students a useful opportunity, but it does not produce the educational collapse some critics predicted.
His recommendations follow from that. Faculty leadership should drop blanket bans and let teachers experiment; universities should invest in training staff, many of whom feel they lack the skills. The master's thesis needs rethinking too, since its core value — sustained engagement with research and argument — can now be produced by a model in minutes. A literature review is no longer enough, he argues; programs should require practical or empirical work in which considered use of AI is itself part of the assessment.
Universities are moving the other way. Berkeley Law at the University of California has banned AI in almost all graded assignments, on the grounds that future lawyers need to build basic reasoning skills before the tool can help them.
That conflict is real, but I do not think this study settles it, and the reason is sitting in Schrepel's own list of limitations. He could not verify how much AI students used during the take-home exam — the half of the assessment that carries the most weight and the least supervision. The other evidence in the same record points straight at that seam. A University of California, Berkeley study covering more than 500,000 grades found that after ChatGPT launched, the share of A grades in writing- and programming-heavy courses rose by 13 percentage points, with no comparable movement on proctored exams. A longitudinal study of more than 26,000 K-12 students in central China found AI raised homework grades while exam results fell by as much as 24 percent, with the worst damage where the model replaced independent thinking rather than supporting it. Read together, they suggest the group that finished last may have been the only one graded on what it could do unaided.
The sample size is also small, and Schrepel concedes his participants — students enrolled in an AI law course — were probably more technically capable than average. Against that, a genuine randomized design run twice is rarer in this field than the confident policy pronouncements built on much less.
The question nobody in this debate is asking is what the no-AI group actually gained. The one thing the study observed about them that the others did not do was argue with each other more. Nothing in the rubric — substance, clarity, proportionality, novelty — measures that, because the rubric grades a document, not the thinking behind it.
Berkeley Law and Schrepel are not disagreeing about the evidence. They are grading different things: one the quality of the output, the other the formation of the person producing it. Neither side currently has an instrument that measures the other's.