We Let AI Draft a Full Market Assessment. This Is What It Took to Trust It.

By Thomas Byrnes
• • 13

This is my pre-read for "AI on the Ground: Realistic Use Cases for Market Systems in Fragile Contexts", the fireside chat this Wednesday 19 August, 15:00 to 16:00 CET, which is 14:00 in London and 16:00 in Nairobi. Co-hosted by the Markets in Crises (MiC) COP Community of Practice and the PRO-WASH and SCALE Award. Register here: https://crs-org.zoom.us/webinar/register/WN_rjc-dQdASxO8KQD58ellFg

The full protocol is free asAidGPT guidance paper MI-GUIDE-002,and theworking draft behind it is published alongside it.

AI disclosure: Claude helped draft and revise this post and the guidance paper it draws on, under my direction. The two cases were run by the MarketImpact team. The workflow, the judgements and the final text are mine, and I take responsibility for the content. Both pieces went through an adversarial audit before publication, with the figures checked against the case records and live sources, and I reviewed and stand behind that check. Our standard:marketimpact.org/how-we-use-ai

On Wednesday I'm joining the Markets in Crises fireside, AI on the Ground, to talk about the question that now sits under every assessment budget: can you trust analysis a machine helped produce? Rather than argue a position, I want to show you what happened when we put the question to real work, twice, and kept score.

The experiment

First, what this was and what it was not, because this audience knows the difference and the claim collapses without it.

A full Emergency Market Mapping and Analysis, done properly, is a two-week job at least. Field work, key informant interviews, traders, households, the market actors themselves. Nothing here replaces that, and I am not claiming it does.

What we needed was the thing that comes before it. A rapid desk review to build the initial analytical frame for the paper: what the transmission story looked like, which channels mattered, and where the field work should be pointed when it happened. On my own, from scratch, that is a working day.

So the question was narrow and answerable. Could AI produce that desk product in a single session, to a standard we would put our name on?

The answer was yes. The interesting part is what yes cost, and I want to be precise about that, because the tidy version of this story is a lie.

The setting matters here. These sessions were preparatory work for a real commission, the analysis we carried out for Mercy Corps on the Iran war's secondary economic impacts, published in May as From Hormuz to the Frontlines of Hunger, with Somalia one of its six contexts: https://www.mercycorps.org/research-resources/hormuz-to-hunger. That series is the finished, human-verified product. What I'm sharing this week is the internal method that sat underneath it.

What came out was a ten-section desk-based EMMA for Somalia, structured to the toolkit. Fifty-one cited sources. More than twenty-eight validated data points. Fifteen explicitly flagged gaps, each with a named follow-up. And when we ran an adversarial audit over the finished draft, hunting specifically for fabricated claims, it found zero hallucinations.

Where the zero came from

Zero is a suspicious number. Everyone knows these tools make things up. I expected the audit to come back with at least one invented citation, and I would have bet money on it. So it's worth being precise about where the zero came from.

The zero did not come from the model. It came from the second hour.

Here's what I mean. AI has collapsed the cost of producing analysis and left the cost of trusting it untouched. The draft is now the cheap half, call it the first hour. Everything that makes the draft defensible, the setup, the evidence rules, the audits, the judgement, is the second hour. That is where the zero was manufactured. And because the second hour is a sequence rather than a talent, you can copy it, which is why I'm writing it down.

The arithmetic nobody puts in the sales deck

Here are the real numbers.

That desk review would have taken me a working day from scratch. The AI produced a complete first draft in about an hour.

I did not get the day back.

What I got was an hour of drafting, then four more hours of reading, cross-checking, chasing sources back to their origin, arguing with the output and rewriting it. Five hours, near enough, from blank page to something I would put my name on.

So call it the second hour if you like. In this case the second hour took four.

Now the part that matters. That five-hour product was better than the one-day product would have been. Not marginally. Substantially. More sources, an explicit register of what we did not know, an evidence grade on every single claim, and two structured attacks on the draft that I would never have run on my own writing at five o'clock on a Friday.

That is the real trade, and it is worth naming precisely, because the sector is being sold a different one. The time the AI saves in drafting does not go into your pocket. It moves. The level of effort stays roughly where it was. What changes is where it goes: out of producing, into verifying. Bank the saving and skip the second part and you have not bought speed. You have bought an unchecked draft with your name on it.

Some days that trade is worth it and some days it is not. A short piece you know cold, you are better off writing yourself. This one was worth it, and I can say why: not because it was faster, but because the finished thing was stronger.

Where the five hours went

Article content

Eight steps, five hours. The first four produced the draft. The last four decided whether anyone should believe it.

It starts before the drafting does

The session never ran in a blank chat window. We loaded the full EMMA toolkit and guidance into the project as its standing reference library, and gave the AI a clear brief: the role it was playing, the methodology that governed the work, and the rules it could not break.

These tools perform best under clear instructions, and market analysis is unusually lucky here, because EMMA already specifies in detail what a good assessment contains. The page was blank. The workspace was not. The framework gave the document its shape, which left the real question, whether each claim inside it was true, to the rest of the sequence.

Then the sourcing rule: local and regional media before the international wires. The wires give you the direction of change. The local press gives you the numbers, the baselines and the texture, in one case the actual pump prices, and the arrest at a fuel protest that opened the whole domestic political dimension. English-language sources lag by weeks, and in a fast-moving crisis those weeks are the analysis.

No tag, no claim

Then the evidence rule, enforced from the first sentence: the AI was not permitted to write a claim without a tag, and a tag forces a source.

Article content

Four grades, applied inline. A model that has to attach a source to every claim has nowhere to put an invention.

The effect on the reader is as useful as the effect on the model. You can see the strength of every sentence at a glance, and so can whoever reviews it. Nothing in the draft is allowed to look more confident than its evidence. A reviewer who wants to attack the weakest claims can find them in about thirty seconds, which is exactly what you want.

Two checks, two different questions

Then the audit, as a hard gate, run as a separate pass so the drafting session never marks its own homework. Fixed categories, severity tiers: unsourced claims, overstatement, arithmetic, circular citations. It found eighteen issues, three of them critical. A confident overstatement that the project's own data contradicted. A shipping route framed as dependent on a chokepoint it doesn't actually transit. A draft citing the project's own earlier output as evidence, which is how errors launder themselves into fact.

Then a robustness review, asking something the audit structurally cannot ask.

The audit interrogates the sentences that exist: is this sourced, is it overstated, does the arithmetic hold. The robustness review interrogates the document: if this is a market assessment, where is the healthcare-access transmission chain? What happens when the religious calendar collides with peak livestock exports? Both were real finds in the Somalia case, two of seven whole dimensions missing from a draft that read as complete.

A draft cannot flag what it never mentions, and neither can an audit of that draft. One pass of either is never enough, and if you only have budget for one, know which blindness you are buying.

What it caught in a finished draft

We then pointed the same audit at an older AI-drafted assessment from the same programme, one with over a hundred validated data points. It caught the two failures that worry me most.

A fact that had gone stale: a plant reported as shut down had quietly resumed operations between drafting and review. And a missing piece of structural context, a pre-crisis gas surplus, that reframed the entire supply story in a draft where every individual fact was true.

Facts can all be right while the story is wrong. The model's editorial decisions are more dangerous than its lies, and two questions catch them: what does this answer assume, and what did it leave out?

The part that does not automate

Which brings me to the bit that took me longest to say plainly. We still had to know what questions to ask, and how to read what came back.

The transmission hypothesis that gave every search a target came from understanding market systems. Deciding that two sources were independent, rather than two outlets carrying the same wire copy, was a judgement. When the audit returned eighteen issues and the review proposed seven missing dimensions, somebody had to tell the real findings from the noise, which takes the same expertise as finding them. The shipping-route error was caught because a human could see the map in their head.

The workflow multiplies that expertise and runs it without fatigue. Run by someone who cannot read its outputs, it would produce well-formatted confidence instead of verified analysis. You cannot properly verify work in a domain you do not know. No workflow changes that.

So when is it good enough

Anyone selling you "the assessment in an hour" is selling you the first hour alone, the unaudited draft, which is the dangerous product. The draft in an hour is real. The warranty is everything after it, and that is what decides whether the output can be defended in front of a donor, an auditor, or a programme team about to act on it.

Which raises the question every trainee asks first: if verification is the expensive half, when do you stop? It is not endless. Severity tiers say what must be fixed. A stop condition says when you are done. And the human assigns the final confidence based on what the answer actually rests on. Without that, the second hour really does become four, then eight, and the tool stops paying for itself.

If you commission analysis rather than produce it

The second hour is still yours to buy, and you can buy it in writing.

We've published the whole protocol as a guidance paper: the sequence above, the nine failure modes we hit with the control that caught each one, and ten requirements you can write into any terms of reference where AI touches the production chain. Three of them, to give you the flavour:

Evidence grading on every claim, inline, from the first draft. An untagged claim is an unaccepted claim.

An adversarial audit log as a deliverable, with its categories, severity tiers and the fixes applied. The audit is the product's warranty, and a supplier who cannot show you one has not run one.

A recency sweep in the sign-off week on every claim about current operational status, with the sweep date recorded. Staleness is the silent killer, and in a fast crisis a claim can be validated, multi-sourced and wrong inside ten days.

None of these slows a good supplier down in any way that matters. Each one removes a place where speed currently hides risk. They cost your supplier nothing except the ability to sell you the unaudited draft.

Free, CC-licensed, deliberately short: https://www.aidgpt.org/AidGPT-MI-GUIDE-002-Verification-First.pdf

And because a methods paper should show its receipts, the Somalia working draft itself is published beside it, every tag and every flagged gap intact: https://www.aidgpt.org/AidGPT-MI-GUIDE-002-Somalia-Case-Exhibit.pdf

One thing I'd flag about our own evidence, applying our scale to ourselves. Two cases, one team, and we are the people who ran them. On the four grades above, that is not validated. Grade it accordingly, and test it on your own work rather than taking mine.

What we teach

None of this stays theoretical once people take it back to their desks. On this month's AidGPT alumni call, which runs under the Chatham House Rule so I'll keep it general, the discussion kept returning to two of the protocol's pressure points: the expertise limit above, which is as true of a human consultant as it is of a model, and that verification depends on disclosure, because you cannot check what you cannot see. Watching a group of alumni argue those points against their own live work is the course doing its job. The community where that argument happens, a WhatsApp group and a monthly call, comes with it.

If you want to go deeper than a paper, this is precisely what we teach. The tagging discipline, the adversarial audit, the robustness review, and the judgement calls that stay human are the spine of the AidGPT course, taught by doing, on real crises, with your own failed drafts as the course material. The rule underneath all of it: every AI output is a draft until you decide it is reviewed. https://www.aidgpt.org/programme

Wednesday is the place to argue with any of this

"AI on the Ground: Realistic Use Cases for Market Systems in Fragile Contexts", co-hosted by the Markets in Crises Community of Practice and the PRO-WASH and SCALE Award.

Wednesday 19 August, 15:00 to 16:00 CET (14:00 London, 16:00 Nairobi) Register: https://crs-org.zoom.us/webinar/register/WN_rjc-dQdASxO8KQD58ellFg

Karri Goeldner Byrne of the MiC Advisory Committee is moderating, I'm speaking alongside Emil Heidkamp p of Sonata Intelligence, and a good chunk of the hour is open to the floor. You can also submit a question when you register, so if you want to put something hard to me, that is the place to do it.

Come with your own second hour, and tell me what yours contains that mine misses.

Thomas Byrnes is CEO of MarketImpact Digital Solutions Ltd and runs AidGPT, a responsible AI training programme for humanitarian and development professionals.

Enjoyed this article?

This post is from Aid and Dev Dispatches, a LinkedIn newsletter with expert analysis on humanitarian reform, AI adoption, crisis economics, and the politics of aid. Join 9,000+ subscribers.

Subscribe on LinkedIn

About the Author

Thomas Byrnes is a Humanitarian & Digital Social Protection Expert and CEO of MarketImpact.