For decades, research in Artificial Intelligence focused on developing increasingly accurate models. While high predictive performance remains a necessary condition, it is no longer sufficient. As AI systems are increasingly used to support decisions in high-impact contexts, such as credit approval, fraud detection and hiring, it’s no longer enough to ask how accurately they make predictions. We must also ask who is most affected by their errors and under what circumstances those errors occur.
In January 2021, the Dutch government resigned following the childcare benefits affair. Between 2005 and 2019, tens of thousands of families were identified by an automated risk assessment system as being at high risk of committing fraud in relation to childcare benefits. This classification triggered a “zero tolerance” administrative policy, often resulting in the immediate suspension of benefits and demands for the repayment of funds that had already been received.
Subsequent investigations revealed that the system used dual nationality, among other factors, as an indicator of fraud risk. Because families with migrant backgrounds were investigated more frequently, the data and decision criteria reflected this institutional bias, leading the system to associate certain demographic characteristics with a higher likelihood of fraud.
The problem was not simply that the algorithm was “wrong”. Rather, it was a system that reproduced patterns already embedded in the data and administrative processes. Algorithms do not invent prejudice; they learn the patterns present in the data, including historical inequalities and discriminatory institutional practices. The case also demonstrated that a system could perform exceptionally well overall while still producing profoundly unfair outcomes for certain groups, particularly when transparency and auditing mechanisms are absent.
Cases such as this make it clear that building more accurate systems is no longer enough. The real challenge is also to develop robust methods for assessing properties such as fairness and explainability.
This is precisely where, in my view, one of the greatest challenges facing Artificial Intelligence research lies today. Unlike accuracy, which can be summarised by a single metric, fairness has multiple definitions that are often incompatible with one another. Measuring the fairness of an AI system is itself a scientific challenge.
This challenge inspired one of the main contributions of Sérgio Jesus’ PhD thesis, which received the Vencer o Adamastor 2026 award; the research addresses a fundamental methodological question in Responsible AI: how can we test and quantify the fairness of AI systems under conditions that closely resemble those encountered in the real world?
To address this question, this work introduces a fairness stress-testing methodology, built upon an experimental environment based on large-scale, real-world banking fraud data with privacy guarantees. This methodology enables the systematic evaluation of model fairness under different bias scenarios. The question therefore shifts from “Is this model fair?” to “Under what circumstances does it cease to be fair?”
The findings show that models delivering excellent performance under apparently normal conditions can become significantly less fair when the context changes. In certain scenarios, legitimate applicants from one age group were more likely to be incorrectly classified as fraudulent than applicants from another age group. This happened not because the algorithm itself had changed, but because the context in which it operated had changed.
More than a conclusion about a specific problem, this is a broader message for the entire field of Artificial Intelligence research. Trust in an AI system depends not only on the quality of the algorithm, but also on the quality of the data, the way the system is evaluated and our ability to understand limitations. When a model influences fundamental rights, it’s no longer enough to ask whether it works. We also need to know who it works for, who it fails, and under what circumstances it can no longer be considered trustworthy.
The Dutch case reminds us that the risks associated with Artificial Intelligence are particularly significant when fundamental rights are at stake. Yet it also offers a second lesson, perhaps a less obvious one. By automating decision-making, AI can amplify existing biases, making them more consistent and measurable than they would be in dispersed human decision-making. Paradoxically, this also creates an opportunity for research. If we can measure these biases, we can understand the conditions under which they emerge and develop more effective methods to limit them. Perhaps the real challenge of Responsible AI is not to eliminate every form of bias (which is probably impossible) but to build thorough methods for identifying, quantifying and mitigating bias before it affects people’s lives.
By Rita Ribeiro, AI researcher and lecturer at FCUP

News, current topics, curiosities and so much more about INESC TEC and its community!