The Model Was Never the Whole Story
MarketImpact researched and wrote two new reports for the Digital Convergence Initiative - DCI AI Hub for Social Protection. A global evidence review of 99 AI cases across at least 48 countries, and a taxonomy that gives the field a common language. Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ) GmbH has just published both. Four of the documented systems have been stopped and a fifth is under a court-stayed stop order, every halt driven by external accountability rather than technical failure, and the record of how they were stopped should worry anyone who trusts internal review to catch this.
The short version
We wrote the evidence base. Avril James and I authored AI Adoption in Social Protection: A Global Evidence Review (2020–2025) for the DCI AI Hub for Social Protection, and co-authored the companion A Taxonomy for AI in Social Protection with independent consultant Valentina Barca. Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ) GmbH has published both. We are also co-authors of the AI Hub's forthcoming risk assessment framework. Free, public, and rated for evidence quality case by case.
This is operational, not experimental. Of the 99 cases, 32 are in full production, 26 in limited rollout and 12 are scaled and institutionalised. These systems are already talking to claimants, verifying identities and flagging people for investigation.
The technology is plainer than the hype. Classical machine learning accounts for 60 of the 99 cases. Foundation models appear in 11.
Internal review has a poor record. In every documented case where a system was stopped, the decisive pressure came from outside the team running it. A court, a regulator, an audit office, a newsroom or a parliament. Where the warning came from an official watchdog but stayed out of public view, it was ignored for years.
Procurement is the fastest fix. If a system influences eligibility, investigation, verification or payment, make disaggregated outcome reporting, effective appeal routes and access to audit evidence contractual requirements.
By late 2020 Togo needed to get emergency cash to people the pandemic had pushed into crisis, and the government had no comprehensive social registry to find them with. What it did have was a recently updated voter register, satellite data, household surveys and mobile-phone records. The urban phase of the Novissi programme ran off the voter register. The rural expansion used machine learning in two stages. Satellite and phone metadata classified 10,119 grid cells and ranked all 397 cantons by poverty, selecting the poorest for coverage. A second model then used phone-based poverty prediction to prioritise roughly 57,000 individual beneficiaries inside those areas, without fresh household surveys. Together the two models helped deliver cash to about 140,000 people in the rural phase.
The approach was later evaluated in Nature. Against the geographic targeting the government could actually use during the emergency, machine learning cut exclusion errors by 4 to 21 per cent. Against methods that would have needed a comprehensive social registry, it increased exclusion errors by 9 to 35 per cent. Togo had no such registry. And people without a phone could not receive the digital payments at all, which is the identity and access gap no model solves.
So Novissi is an important case, and a warning against the fairy-tale version where AI finds objective truth in messy data. It was an emergency workaround. Its value depended entirely on the alternatives available, and its limits stayed real.
The Netherlands ends differently. SyRI linked data across government agencies to identify citizens considered at higher risk of fraud. On the State's own account it was closer to rule-based risk scoring than to modern machine learning, a simple decision tree rather than a self-learning system. The court couldn't check that, because the State never disclosed the model or its indicators, and the legislation left the door open to predictive analytics, deep learning and data mining anyway. In February 2020 the District Court of The Hague ruled that the legislation behind SyRI had no binding effect. The system was insufficiently transparent and verifiable, and the law failed to strike a fair balance between fraud prevention and the right to private life. It had not produced a single proven case of fraud in the deployments that triggered the lawsuit.
These two are not matched experiments and they prove nothing about sophisticated models being safe or simple ones being dangerous. What they show is that "which model are we using?" is too small a first question. What is the system pointed at, which decision does it influence, what evidence supports it, who can contest the result, and what legal, institutional and human safeguards surround it?
The model matters. It was never the whole story. After ten months of research we can now show the evidence behind that.
The evidence base is public now
Avril James and I authored AI Adoption in Social Protection: A Global Evidence Review (2020–2025) as consultants through MarketImpact Digital Solutions Ltd. It maps 99 documented cases across at least 48 countries, from biometric verification at humanitarian distribution points to fraud models inside European welfare agencies. Every case carries a source citation and an evidence rating. Twenty-five are rated A, resting on a peer-reviewed study or an independent audit. Seventy-three are rated B, drawing on implementer evaluations, government or multilateral reports, working papers, and legal or regulatory documents. One is rated C and rests on a media or single-source account.
We rated our own evidence base because the honest answer to "how much do we actually know?" matters more than the headline count. The answer is that three quarters of what this field knows about AI in social protection has never been peer-reviewed or independently audited.
The companion report, A Taxonomy for AI in Social Protection, the two of us wrote with independent consultant Valentina Barca. It exists because the field keeps talking past itself. "AI in social protection" currently covers everything from a chatbot answering questions about opening hours to a model deciding who gets investigated for fraud, and those are not the same governance problem. The taxonomy identifies eight AI capabilities in two families, analytical and generative, and separates them from ten use-case categories across the delivery chain. Then it connects them to decision criticality, the human oversight that level of criticality actually requires, and the wider governance risks. It came out of discussions with Digital Convergence Initiative - DCI partners including the The World Bank Group , ISSA, the International Labour Organization and the OECD - OCDE . The AI Hub's forthcoming tracker will put the same structure into a searchable resource, case by case.
One report gives the field a language. The other gives it a record.
Don't file this under ministries and welfare states. The review includes humanitarian applications wherever they meet government social-protection programmes. UNHCR, the UN Refugee Agency biometric authentication in Ethiopia and cash assistance for returnees in Afghanistan. World Food Programme 's SCOPE biometric verification in the Democratic Republic of the Congo. Displacement forecasting by the Danish Refugee Council / Dansk Flygtningehjælp in South Sudan, which is one of the few grade-A benefit findings in the whole review, at a reported EUR 6.60 saved for every EUR 1 spent. I've been writing about the ethics of identity matching and the case for AI in anticipatory action here since 2024. What's new is that the practice now has a documented record instead of a pile of anecdotes. If you work on humanitarian cash, registration or anticipatory action, this map already contains your work under a more precise vocabulary.
This is no longer a portfolio of pilots
Thirty-two cases are in full production, 26 in limited rollout, 12 scaled and institutionalised. Twenty remain at pilot stage, four are in design and five have been suspended or halted. That's a documented evidence base, not a census, so it can't tell you how common AI is across every social-protection system. It can tell you the future-tense conversation is over.
Four use-case categories dominate. User communication accounts for 22 cases, vulnerability and needs assessment for 16, compliance and integrity for 14, identification for 13. Those four hold two thirds of the evidence base. Systems are already talking to claimants, deciding who looks vulnerable, flagging people for investigation, and verifying who they are.
I wrote a year ago about watching AI move from the edges of operations into the day to day. Eligibility checks, grievance triage, fraud routines, anticipatory cash. That was an impression from practice. It's now a documented pattern.
The technology doing this work is also less exotic than the conference agenda suggests. Sixty of the 99 cases are classical machine learning, 19 deep learning, 11 foundation-model systems. Generative AI dominates the conversation. The systems actually making or informing consequential decisions mostly run on older, plainer methods.
That distinction matters for data sovereignty. For the 60 per cent running classical machine learning, sovereignty is technically achievable. These systems don't need large-scale compute or cloud API access, and as I argued in April, the excuse for not running things locally is disappearing fast. But technical possibility is not organisational control. Skills, procurement, maintenance, data pipelines and vendor dependencies still decide what can be run at home. And the documentation isn't there to check. Most cases in the review don't give enough information to determine their compute environment, data residency or cross-border transfer arrangements. The AI Hub's companion work on data sovereignty found two thirds of cases offer no public documentation of where data is stored or processed, and at least 8 per cent show confirmed indicators of vendor lock-in. Same blind spot I set out in Your AI Has a Passport Problem. The technology may permit more control than institutions are currently exercising. Or at least more than they are documenting.
Systems get stopped from outside, or they don't get stopped
The review's tables record five of the 99 as suspended or halted. On the live record it is four stopped and one still running, and the one still running is the most instructive, so I will take it separately. In every stop the decisive force was external accountability rather than technical failure. Look at how, and a pattern shows up that should trouble anyone who believes internal governance is doing this job.
SyRI ended because a court struck down the law behind it. Sweden's parental-benefit risk model came out of use in 2025 while the national data-protection authority had an open supervision running. Rotterdam suspended its welfare-fraud model in 2021, months after the municipal Court of Audit reported that its ethical choices were untraceable and its outputs risked bias, and by the city's own later account it had concluded the model could never be kept fully free of bias. Lighthouse Reports later reconstructed that model and documented discriminatory effects, and in July 2023, after that investigation and criticism from the Dutch data protection authority, the city paused work on a replacement. Denmark's Gladsaxe child-vulnerability pilot never got the legal basis it needed. The ministry refused the exemption, the press exposed the scheme, legal scholars attacked it publicly, a data breach exposed around 20,000 personal identity numbers, parliamentary support collapsed, and the municipality shelved it in January 2019.
Ireland shows the limit of the pattern. The Data Protection Commission concluded a four-year statutory inquiry in June 2025 and found multiple GDPR infringements in the biometric registration behind the Public Services Card. It issued a reprimand, fined the department EUR 550,000, and ordered it to stop processing biometric data within nine months unless it could find a lawful basis. The department appealed, the cease order was stayed pending the High Court, and the system is still running. External scrutiny found the problem. It hasn't stopped it.
That's four stops and one stayed order, and I'm not going to dress it up as a universal law. It doesn't prove every system still operating is well governed. It may partly show that scrutiny is unevenly spread. Some systems survive because they're sound. Others may survive because nobody with enough access and authority has challenged them yet.
What the five do show is that internal review has a poor record, and Sweden is the clearest case. In 2018 the Swedish Social Insurance Inspectorate, the government's independent watchdog for the agency, found its risk-based controls did not pass the test for equal treatment. The agency rejected the finding and questioned whether the inspectorate's fairness tests were an accepted way of testing equal treatment at all. Then it carried on. Investigative reporting exposed wider disparities in November 2024. The data-protection authority opened supervision in June 2025, and the agency withdrew the risk profile while that supervision was under way. The warning had been sitting on the table for six years. What changed was that the scrutiny became public, and then regulatory.
There's a structural reason for this and it's in our own data. Almost every one of these systems has a human involved. Fifty-four cases report a human in the loop, 42 a human on the loop, three no human at all. On paper that's reassuring. But a system described as human-in-the-loop is only as good as the human in it, and if that reviewer is under time pressure, untrained to judge the output, or routinely waving recommendations through, the oversight is weaker than the label. Our evidence base rarely reports override rates, reviewer training, or the conditions under which a human disagrees with the AI. Formal oversight is documented almost everywhere. Effective oversight is documented almost nowhere.
So external accountability is a backstop, not an evaluation system. Courts and regulators test legality, rights and institutional conduct. Journalists make hidden practices visible. Independent evaluation asks something different: does the system work, for whom, compared with what, and with what unintended effects. These overlap and none of them substitutes for the others.
Right now the evaluation layer is remarkably thin. The review puts it plainly. Independent evaluation of AI in social protection is rare, and that is the most important cross-cutting finding after governance. Most evidence comes from implementer reports and technical documentation. Peer-reviewed evaluation covers a small subset. Randomised and quasi-experimental evaluations are almost entirely absent. Vendor claims of accuracy and efficiency are frequently unverified.
This isn't confined to AI. I wrote in May about an open-source cash platform the sector still hasn't properly evaluated, which has been in production for years. We are not good at this. I expected the evidence base to be uneven when we started. Even I wasn't ready for how thin the independent layer turned out to be.
There are benefits worth taking seriously. Maidstone's predictive system generated 650 alerts for households at risk of homelessness, and the officer had capacity to attempt contact with only 260 of them. Of those, 0.4 per cent later presented as homeless. Of the alerted households nobody could reach, 40 per cent did, and the council estimated savings of more than GBP 225,000. That gap is striking, and it isn't a controlled comparison. The officer chose which households to contact, so if the easier ones were also the more tractable ones, the gap overstates what the intervention did. An analysis published by Naver with Yonsei University researchers estimated a 44.2 per cent lower rate of solitary deaths in areas using Korea's CLOVA CareCall than in areas without it, in a study the operator commissioned and published. Indian government reporting records average grievance-disposal time falling from 32 days in 2021 to 13 days across 2024, with AI-assisted triage forming one step of a ten-step reform package, so those figures don't isolate what AI contributed.
The numbers are promising. None of the three is an independent evaluation. And the distance between a documented outcome and a proven effect is exactly what the missing evaluation layer is supposed to close. Togo is the exception that proves it. Peer-reviewed, and it reports a finding against the method as well as for it.
Aggregate accuracy hides the distribution of harm
Sweden shows why disaggregated scrutiny matters. Its model selected temporary parental-benefit applications for fraud investigation, and independent analysis by Lighthouse Reports found women were more than 1.5 times more likely than men to be selected, and applicants with a foreign background close to 2.5 times more likely than those with a Swedish background. The investigators reported the model would still have passed the agency's own two-step fairness procedure without adjustment.
That's not a Swedish problem. Across the Dutch and Swedish fraud-detection cases the review found algorithmic selection over-flagging women by a factor of roughly 1.5 and people with foreign backgrounds by 2.5 to 3. Those patterns survived court and regulatory scrutiny and contributed to each system's eventual suspension.
Disaggregating outcomes by gender, disability and ethnicity is the single most common reporting gap in the whole evidence base. Fourteen cases engage directly with disability-related entitlements, mostly eligibility determination or case triage, and not one of them disaggregates outcomes by disability type. If nobody measures who is being flagged, excluded, delayed or wrongly classified, a respectable aggregate accuracy rate can hide the distribution of harm completely.
We were careful about what we didn't claim. Biometric identity systems, predictive targeting and compliance monitoring together build a data infrastructure that could, without adequate safeguards, enable disproportionate surveillance of vulnerable populations. Not one case in the evidence base reports surveillance as a realised harm. So we recorded it as a plausible systemic risk rather than a documented impact, and the AI Hub's forthcoming risk assessment framework, which we co-authored, takes it up properly. The distinction matters. The other harms in that chapter are evidenced. This one is inferred, and saying so is the difference between analysis and advocacy.
The review says the next wave of deployments should demand disaggregated evaluation design upfront. I'd go one step further, and this is the one thing I want you to take from all of it. If you are procuring a system that influences eligibility, investigation, verification or payment, make disaggregated outcome reporting, effective appeal routes and access to audit evidence contractual requirements. Not principles. Clauses. The most common gaps in this evidence base can close one procurement at a time.
One as a dictionary, one as case law
More than 100 pages is an unreasonable thing to hand a busy programme director without saying how to use it. So use the taxonomy as a dictionary. When a vendor, a consultant or an enthusiastic colleague proposes something "AI-powered", it gives you the questions that make the proposal specific. Which capability is involved, which use case does it serve, how consequential is the decision, and what oversight does that level of criticality require. A chatbot explaining opening hours and a model scoring eligibility are both called AI and they carry entirely different consequences. The taxonomy gives your team a shared language for refusing to treat them as one governance category.
It also settles a question I get asked constantly. Agentic AI is an architectural pattern that orchestrates several capabilities. It isn't a ninth capability of its own.
Use the Evidence Review as case law. Before you pilot anything, go to the case index. Find who has attempted the same use case, where, with what reported outcome, and on what quality of evidence. Read the safeguards as closely as the benefits. Read the failed and halted cases before you write the terms of reference. Nobody needs to be the first mover blind.
The map has edges
The review documents what we found, not everything that exists. It's a desk-research baseline built between June 2025 and March 2026 from peer-reviewed literature, institutional and legal documents and high-quality grey literature in English, French and Spanish, with particular emphasis on low- and lower-middle-income contexts. The high-income cases are not a systematic inventory.
Three languages beats one. It isn't the same as complete, and I learned that the hard way. On a recent Ebola desk review our first English-language search produced a confident and wrong picture of who was leading the Congolese response. Francophone sources corrected it. That wasn't a detail. It changed the finding. Any evidence base built in three languages has edges in the other seven thousand, and systems documented only in internal reports, local-language sources or channels we couldn't reach will be missing. So will the ones nobody is documenting at all, which is the same problem I described as Shadow AI inside humanitarian organisations, one level up.
The gaps have a shape. Upper-middle-income countries contribute only 11 cases. China, Turkey, Thailand and South Africa don't appear at all. Labour-market programmes contribute ten. AI-assisted casework, long-term outcomes and the lived experience of people assessed by these systems are all thin. Some of that may be under-deployment. Some is likely under-documentation. The review can't tell you which without more evidence.
So here's the ask. If you run, fund, study or know of a missing case, tell me, and everything credible goes to the AI Hub team for the next iteration. Please don't share personal data or confidential, protection-sensitive or otherwise restricted material. If a case can't safely be made public, share only what you're authorised to disclose and agree explicitly how it may be used.
Model choice matters. It just can't carry the governance burden by itself. A more sophisticated model doesn't repair a badly chosen purpose, weak evidence, absent legal safeguards or a system nobody can contest. And a simpler model can still do valuable work when its purpose is legitimate, its limits are visible and its outcomes get tested.
Download the Evidence Review and the Taxonomy. Use the taxonomy to name the capability and the use case. Use the review to see who has already tried it, what happened, and how strong the evidence is. Ten months of work like this is the research side of what we do at MarketImpact.
The review leaves one practical question for every organisation. How do you make your AI use visible before a court, an auditor, a journalist or a donor makes it visible for you?
That's the next piece.
One last thing
The evaluation layer is thin partly because checking is a habit almost nobody has been taught. Building that habit is most of what we do on AidGPT, the responsible AI training programme we run at MarketImpact. Six sessions over three weeks, capped at twenty, and every exercise runs on your own work rather than a case study.
The August cohort starts on 18 August and finishes on 3 September. The next one starts on 22 September. Each runs in two slots, 07:00 and 13:00 UTC, one that suits Asia-Pacific and Europe and one that suits Europe and the Americas. EUR 350, or EUR 280 for national NGO staff, local organisations and aid workers between roles. Details and registration at aidgpt.org/training.
If you're the person who has to ask a vendor which capability is involved and what evidence supports it, that's the room to be in.
Tom
Thomas Byrnes is CEO of MarketImpact Digital Solutions Ltd and runs the AidGPT responsible AI training programme. AidGPT is MarketImpact's training brand. Thomas Byrnes and Avril James of MarketImpact Digital Solutions Ltd authored the Global Evidence Review, and with independent consultant Valentina Barca co-authored the Taxonomy. Both were prepared for the DCI AI Hub for Social Protection and published by GIZ. MarketImpact is also a co-author of the AI Hub's forthcoming Risk Assessment Framework for Social Protection.
Tom's Aid and Dev Dispatches is my LinkedIn newsletter on humanitarian and development trends. AidGPT email updates are a separate service focused specifically on practical AI adoption, training and governance.
Previously in this series: Shadow AI in Humanitarian Work, November 2025. Use AI's Mind, Not Its Memory, January 2026. Most AI Training in the Humanitarian Sector is Teaching the Wrong Thing, April 2026. SAFE AI is Live, June 2026. Your AI Has a Passport Problem, July 2026.
AI disclosure. Claude helped draft and revise successive versions of this piece, including the summary, from the two published reports and my direction. OpenAI Codex audited the claims against the reports and official sources, tested the logic, and helped restructure and redraft the article. Claude then ran a full pre-publication audit against primary sources and against the published reports themselves. It found errors in earlier drafts, including two of its own, and I corrected them. Every figure retained here has been checked against its source. The argument, the selection of evidence and the final judgement are mine. AI tools were also used in producing the Evidence Review, with human review throughout, as its methodology chapter records. No beneficiary-level data, client-confidential material or unpublished operational datasets were used in preparing this article. If anything here is wrong, tell me and I will correct it publicly and say what changed. MarketImpact's full position on AI use is at marketimpact.org/how-we-use-ai.
Sources
DCI AI Hub for Social Protection (2026), AI Adoption in Social Protection: A Global Evidence Review (2020–2025), Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ), Germany.
DCI AI Hub for Social Protection (2026), A Taxonomy for AI in Social Protection: Defining Capabilities and Application Areas, Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ), Germany.
DCI AI Hub for Social Protection (2026), AI in Social Protection: Data Sovereignty Considerations, Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ), Germany.
DCI AI Hub for Social Protection (forthcoming), AI Hub Risk Assessment Framework for Social Protection, Deutsche Gesellschaft für Internationale Zusammenarbeit (GIZ), Germany.
Aiken, E., Bellue, S., Karlan, D., Udry, C. and Blumenstock, J. (2022), "Machine learning and phone data can improve targeting of humanitarian aid", Nature, 603, 864–870.
District Court of The Hague (2020), Judgment in case C-09-550982-HA ZA 18-388, ECLI:NL:RBDHA:2020:865, 5 February 2020; English court summary, 13 February 2020.
Rekenkamer Rotterdam (2021), Gekleurde technologie: Verkenning ethisch gebruik algoritmes, April 2021.
Lighthouse Reports (2023), "Suspicion Machines", 6 March 2023.
Lighthouse Reports (2024), "How we investigated Sweden's Suspicion Machine", 27 November 2024.
Inspektionen för socialförsäkringen (2018), Profilering som urvalsmetod för riktade kontroller, Rapport 2018:5.
Swedish Authority for Privacy Protection (2025), "Supervision concluded after Försäkringskassan stopped using AI system", 18 November 2025.
Data Protection Commission, Ireland (2025), "DPC announces conclusion of investigation into use of facial matching technology in connection with the Public Services Card", 12 June 2025.
Kristensen, K. (2022), "Hvorfor Gladsaxemodellen fejlede - om anvendelse af algoritmer på socialt udsatte børn", Samfundslederskab i Skandinavien, 37(1), 27-49. Available at: https://rauli.cbs.dk/index.php/SiS/article/view/6542/7067
Armed Conflict Location & Event Data (2024), "Helping the Danish Refugee Council anticipate and prevent forced displacement".
Crisis (2023), "Homelessness prevention by Maidstone Borough Council and Xantura", 15 February 2023.
Naver (2026), CLOVA CareCall social-value measurement, 19 March 2026.
Press Information Bureau, Government of India (2022), "Union Minister Dr Jitendra Singh releases the Annual Report of CPGRAMS for the year 2022", 20 December 2022.
Press Information Bureau, Government of India (2025), "The Department of Administrative Reforms and Public Grievances released the December 2024 CPGRAMS performance report", 31 January 2025.
Enjoyed this article?
This post is from Aid and Dev Dispatches, a LinkedIn newsletter with expert analysis on humanitarian reform, AI adoption, crisis economics, and the politics of aid. Join 9,000+ subscribers.