No AI Framework Sits at the Keyboard: What AidGPT's First Measured Open Cohort Showed

By Thomas Byrnes
• • 9

One participant in our June cohort is the case that keeps me up at night. Not a sceptic. Not a beginner. A capable user who, by their own account, had been trusting AI output more than they should. Using it without checking it.

That is the quiet failure mode in this sector. Not the person who refuses to touch AI, but the person who uses it well enough to trust it and not well enough to check it.

Article content

Last month, writing about SAFE AI , I argued that no framework can sit at the keyboard. Policy approval is not the real work. The real work is the moment a staff member gets a plausible answer under deadline and has to decide what to do with it. Is the source real? Is the summary quietly dropping the people least visible in the data? Is the tool helping with judgement, or replacing it?

No policy document reaches that moment. But I left the harder question open. If the framework cannot build the person at the keyboard, what does?

June gave me the first measured answer from an open cohort. I will come back to the person who keeps me up at night, because what happened to them is the answer in miniature. First, the room.

A different kind of room

Every set of numbers I have published before came from a single organisation. Action Against Hunger UK earlier this year, and several other humanitarian and refugee-response organisations before that. One client, one team, one culture, learning together.

June was not our first open cohort. It was the first open cohort where we had both baseline and endline survey data, and it was a different kind of room. We had already learned from earlier open delivery. What June gave us was a measured signal across a mixed room: participants from a UN agency, a major international NGO, a bilateral donor programme, a humanitarian research and innovation organisation, an academic centre, and independent consultants between assignments.

Fourteen enrolled and twelve completed the full course, with the other two moving to a later cohort. Starting points ran from someone who had barely opened an AI tool and was not sure they needed one, to daily users running several models at once.

That mix was the test. Moving a single team that shares a manager and a workflow is one thing. If the shift only happens inside one organisation's culture, it is a much weaker thing than we claim. If it happens across a room this varied, it is not a quirk of one culture.

Article content

The number I judge the course on

One participant ran a verification agent over a scoring tool they had built. It flagged two rows that came to 110 and 105 per cent instead of 100. An arithmetic error they had not caught by eye.

Five of the nine people who completed the endline survey self-reported a catch like that: something the course caught that they would otherwise have used. That is the figure that speaks to risk, and it is the one I would judge the course on.

The softer numbers moved too. Everyone who rated themselves below the top of the scale at baseline reported more confidence by endline. One moved a full three points on our five-point scale. Every endline respondent rated the course nine or ten. But high ratings are common and not very diagnostic, and feeling more confident is not the same as being more careful. Confidence is the easy thing to move. The catch rate is the one I trust.

How much weight can it bear? Nine of the twelve completed the endline survey. Nine self-selected respondents are a signal, not a study, and the catches are self-reported, not independently verified. Nothing here proves the shift holds across the sector, or that it lasts. It is a reason to keep measuring, not a finding to bank. Read every number in this piece in that spirit.

The pattern underneath is the argument I made in April with the matched Action Against Hunger UK numbers. There, the skill that moved least was writing clever prompts, and the skills that moved most were verification and data safety: checking whether the plausible answer in front of you is actually true.

June said the same thing in a different voice. Prompting is the easy part, and it gets easier every month as the models improve. Checking is the hard part, because a confident wrong answer looks exactly like a confident right one. What June added was the first measured open-cohort signal that the shift is not a quirk of one well-run in-house programme, and that it may hold across a room that shares no employer. Whether it lasts is the separate question I come to below.

The default that flipped

Back to the person who keeps me up at night. By the end of the course, by their own account, the default had flipped. Instead of asking whether an answer sounded right, they had started asking what they would need to check before using it. Whether that holds under a deadline three weeks later is the real test, and it is exactly what the thirty- and sixty-day check-ins exist to find out.

Put plainly, because this is the whole argument in one sentence: the framework I wrote about last month could never have reached that person. A policy on the shared drive would not have changed what they do at nine in the morning with a deadline and a plausible paragraph on the screen. Six sessions of doing the work, with other people watching and asking the same questions, did.

That is what building the person at the keyboard looks like. Not abstract. One specific habit, begun in one specific person who did not have it before.

Why the mixed room worked

The diversity I worried might dilute the room turned out to be the mechanism. When one participant described how an AI had flattened a needs assessment and dropped an affected group, the whole room recognised their own version of it. A donor-programme officer and an independent evaluator do not share a workflow. A consultant writing proposals, a UN officer writing situation reports, and a researcher pulling from a literature base hit the same wall from different sides, and in June they were in the same room saying so.

People learned as much from a stranger's question as from anything I said. The mix turned up failure modes no single team could have, and the case for responsible use landed harder because it was not one manager's rule. It was a room of peers arriving at the same conclusion. That is the argument for an open cohort over an in-house one, and I did not fully believe it until I watched it happen.

The part that does not work yet

Here is what the AI training market will not tell you, and what I would rather say plainly. The course builds the habit. It does not, on its own, embed it.

By the end, most participants had built or started something they intended to keep using, but had not yet folded it into how they actually work day to day. Two had. The rest were somewhere in between. Building happened in the room. Embedding is the work of the month that follows, and the month after that.

This is not a failure of the training. It is the honest shape of behaviour change, and pretending otherwise is how you get glowing evaluations and no lasting difference.

It is why we do not stop at the closing session. We run thirty- and sixty-day check-ins. We have started an alumni community of practice across the April and June cohorts, with the first call this month. And every respondent asked for an advanced follow-on. Nobody asked for more prompting tips. They asked for help making the habit stick under pressure. That tells you where the real work is.

If you are deciding

If you are choosing whether to commission AI training for your team, or whether to join a cohort yourself, here is the test I would apply. Ask what the training is organised around. If the answer is better prompts, cleaner outputs, and faster drafting, it is selling you the skill that was already going to improve on its own. If the answer is judgement, verification, and knowing when not to use the tool at all, it is selling you the skill that transfers.

Article content

Frameworks still matter. They decide which tools are approved, where the escalation routes run, and who is accountable. But no framework builds the habit of checking, and no model takes responsibility for whether its answer should be used. A room of peers, doing the work and checking each other, does.

The AidGPT Intro Lab runs on Monday 20 July: ninety minutes, hands-on, for anyone who wants to test the method before committing to the full cohort. It costs 75 euros, and that fee comes off the full cohort price if you continue.

The next full six-session open cohort begins on Tuesday 21 July, six ninety-minute sessions over three weeks, with further cohorts in August and September. A reduced rate of 280 euros on the 350-euro fee is available for national and local NGO staff in low- or middle-income or crisis-affected countries, and for aid workers between roles. Applications for all of them are open at aidgpt.org/training.

Tom

Thomas Byrnes is CEO and lead consultant at MarketImpact Digital Solutions, and leads the AidGPT Responsible AI in Practice programme for humanitarian and development professionals.

**Disclosure and AI use:**AidGPT is MarketImpact's own training programme, so I have a commercial interest in the conclusions here. The numbers above are from our own cohort and should be read as a first-party signal, not independent evaluation.

AI tools assisted with structure testing, drafting options, editing, and adversarial review of this piece. The cohort numbers, interpretation, final argument, and editorial judgement are mine. No beneficiary-level data, client-confidential material, identifiable participant data, or raw unpublished operational datasets were used in the AI-assisted drafting or review process.

In keeping with what the course teaches, I used AI as a drafting and review aid, then checked, edited, and stood behind the final piece myself. If I have got a material fact wrong, I will correct it publicly.

Enjoyed this article?

This post is from Aid and Dev Dispatches, a LinkedIn newsletter with expert analysis on humanitarian reform, AI adoption, crisis economics, and the politics of aid. Join 9,000+ subscribers.

Subscribe on LinkedIn

About the Author

Thomas Byrnes is a Humanitarian & Digital Social Protection Expert and CEO of MarketImpact.