The invisible crisis: when the grade is a false positive

The invisible crisis: when the grade is a false positive

There is a fact that should keep any rector awake at night. Outstanding grades on assignments have become a predictor of worse performance on exams. The use of artificial intelligence (AI) biases the ability to measure whether learning has occurred. Even so, from these biased readings, competencies continue to be certified and qualifications granted.

Read more Camello and Ratiu turn Butarque into Vallecas and give Rayo the first victory of the season

David Strömberg, from Stockholm University, along with Victor Lei and Yanhui Wu, from Hong Kong, followed 26,811 Chinese secondary students for 30 months. AI raised homework grades by 18% and reduced the time spent on completing it by 30%. After six months, closed-book exam results had dropped by 20%.

Robert and Elizabeth Bjork have spent decades describing “desirable difficulties,” crucial for learning. Learning happens when retrieving a fact with effort; when making mistakes and starting over; when refining an idea. These are slow and uncomfortable processes. They are what ensure learning. AI makes easy precisely the prerequisite — difficulty — for learning to occur.

Hamsa Bastani and her team demonstrated this with nearly a thousand Turkish students in PNAS. With the open interface, the subsequent exam result dropped 17% compared to those who never had it. With a version that only gave teacher hints and never the solution, practice improved by 127% and the negative impact disappeared.

Zara Contractor and Germán Reyes, from Middlebury, randomly assigned university students to study a new topic with or without a chatbot. AI raised the results of subsequent tests without help, and the advantage remained intact a week later. The improvement persisted among those who used it to have concepts explained. It was lost when it was used to generate text. The decisive question is not how much AI the student uses, but what they do when they use it.

What to do? Let’s consider the unfeasible alternatives. Prohibiting its use: already 92% of Latin American university students regularly use some tool, according to the 2026 survey by the Digital Education Council and Tec de Monterrey. In Spain, 59% of young people use it to study, the highest figure in the EU according to Eurostat. Detecting its use is a lost battle from the start.

An alternative can be derived from the results of the Chinese study. The impact on learning was concentrated in the 80% of users with abnormally short times and high grades. That shifts the debate from a moral discussion to a measurable problem and suggests an operational rule: the time the student dedicates to the task must not decrease. If they gained half an hour, they should use it to go further, not to finish earlier.

The replicable part is the design. The difference between Bastani’s two versions was not in the model but in the instructions governing it: asking instead of answering, the proven Socratic method. Today that pedagogical decision is made by the provider. It should be included in the terms of reference of each university’s contracting documents.

Read more Alma Guillermoprieto: “We are on the edge of a cliff, I don’t know if we will fall or not”

Evaluation remains. As long as the grade rewards the deliverable, no system will distinguish between the one who understood and the one who generated an output. Evaluation must shift to in-person: closed book, oral defense, new problems before the teacher, collaborative projects with a significant doing component. Princeton did this in one year: its faculty voted to supervise all exams, something prohibited since 1893. One piece is missing in almost all regulations: the debt rule. Work done with AI is acceptable if the student reconstructs their reasoning autonomously in three minutes. If they do not reconstruct it, it is not theirs. It takes time and oral fluency punishes the shy and non-native speakers. It would be a spot check verification.

Discernment remains. Sam Wineburg documented at Stanford that students and even historians valued the credibility of a source by the appearance of the site, while professional fact-checkers left the page to cross-check. With AI this worsens: the answer arrives without author or trace. Verification must be a graded exercise. Likewise, AI should be available from a certain age: UNESCO set 13 years in 2023 for autonomous use, while students consolidate reading and calculation.

In Latin America there is a long-standing learning debt. PISA 2022 already showed that three out of four 15-year-old students do not reach the basic level in mathematics, compared to 31% in the OECD. AI makes this reality invisible with impeccable assignments. The solution is costly: Princeton can afford supervision and small groups, while Mexico invests 3,650 dollars per student compared to the OECD average of 13,210. The tool is democratized; the antidote to prevent its inappropriate use is not. In Spain, the CYD Foundation records that 46% of university students say their faculty does not promote the use of AI — six points more than a year earlier — and that 61% have received no training on it. The lack of criteria is filled in by themselves. It is dangerous.

There is an experiment that offers the literal image of what is at stake. Nataliya Kosmyna and her team at the MIT Media Lab connected 54 Boston university students to an electroencephalograph while they wrote essays: some with ChatGPT, others with a search engine, others with no help. Those who wrote alone showed the widest and most distributed brain connectivity networks; those with the search engine, intermediate; those with AI, the weakest. 83% of the latter failed to cite a single sentence from their own essay. The authors called it cognitive debt.

The debt rule is not a metaphor. It is the recognition that this liability exists and that someone will have to honor it: the graduate at their first job, the patient before their doctor, the citizen before the bridge someone calculated.

The framework does not hold without teachers capable of recognizing when a student has understood and when they have delegated the work their own brain should do. It is the most effective investment that decides the outcome.

The inequality that PISA measures today in learning will tomorrow transfer to the credibility of the qualification certified by the diploma. A diploma that no one believes in ceases to be a problem of the student who received it: it becomes a problem of the institution that issued it.

Read more A 78-year-old man drowns while bathing at Lloret de Mar beach

Translated from

Leave a Reply

Your email address will not be published. Required fields are marked *