The AI Decision Fluency Handbook
Institutional judgement for the people who approve, fund and answer for AI but who will never open it.
What this document is
This is the case, not the checklist.
It sets out why the most consequential AI decisions in an organisation are being taken by people who have been told, wrongly, that they are not qualified to take them; what that has already cost, in cases that are a matter of public record; and a four-part model of the judgement those decisions actually require.
It is written to be circulated. If you are trying to explain to a board, an executive team or a group of colleagues why this matters and what you propose to do about it, this is the document to send them, and the sources at the back are there so that the sceptical reader can check the argument rather than take it on trust.
Two companion pieces exist and are free. The field guide at aifluency.uk/guide.html is the complete and current library of principles, which grows as scenarios are added. The simulations put you inside the decisions themselves, which is a different kind of learning from reading about them and, in my experience of running these with senior groups, a more durable one.
What it is not
It is not a compliance instrument and does not substitute for one. The NIST AI Risk Management Framework3 and the obligations the EU AI Act places on deployers4 tell an organisation what it must establish and document; this tells an individual what to think about. Those are different jobs and you need both.
Nor is it a validated psychometric instrument. There is no cohort data behind it and no claim that the four dimensions have been empirically established as separable. It is a structured argument, built from a decade of watching these decisions go wrong, and a diagnostic mirror. It should be read as one, and I would rather say so on the second page than be found out on the twentieth.
The gap
Who is actually deciding, what they are offered instead of help, and what the absence has already cost.
1 The people who decide are not the people using it
The people making the most consequential AI decisions in an organisation are, for the most part, not the people using AI.
A director approving a three-year platform contract. A senior civil servant signing off a case triage tool. A hospital chief executive introducing automated referral prioritisation. A board setting the risk appetite for automated decision-making. Each is making a call with more downstream consequence than any individual prompt written anywhere in their organisation. Very few of them will ever open an AI tool. Almost none will be able to interrogate how one works. And yet the quality of their judgement determines whether the organisation gets value, gets exposure, or gets both.
We can broadly describe this group as lacking AI fluency, and the description is fair. What is not fair is the support offered to them to advance it.
Two observations follow, and they are the ones the rest of this document is built on. The first is that the decisions in question are not technical decisions. They are decisions about the division of labour, about how precisely an organisation can say what it wants, about what it will accept as evidence, and about who remains answerable when the sponsor has moved on. Every one of those is a judgement senior people have been making their whole careers, applied to an unfamiliar object.
The second is that treating them as technical decisions is not a neutral mistake. It is the specific mechanism by which the room stops scrutinising. A director who has been persuaded that the subject matter is beyond them will ask a technical question, receive a fluent answer they cannot assess, and record themselves as satisfied. The proposal leaves the room stronger than it entered, having been examined by nobody.
2 Three kinds of help, and what each is for
The support on offer arrives in three forms, and the third has arrived recently.
The first is practitioner training, adapted upward: prompt techniques, tool demonstrations, a hands-on session that leaves a chief executive able to draft a better email and no better equipped to evaluate a business case. It is usually well delivered. It is simply aimed at a different competency than the one the audience needs.
The second is strategy content, pitched downward: trend briefings and board primers that describe what AI might do to the sector over five years without ever touching the specific decision on the agenda next Tuesday. It is usually true. It is also unusable at the moment of decision, which is the only moment that generates consequence.
The third has emerged since 2023, and it is a serious category. Board AI governance education is now offered by institutions that know what they are doing.11 INSEAD runs AI for Boards, a four-day programme through its Corporate Governance Centre. The National Association of Corporate Directors runs Effective AI Oversight for Directors with Carnegie Mellon's Heinz College, a twenty-two hour certificate. The Diligent Institute offers an AI Ethics & Board Oversight Certification. The Corporate Governance Institute sells a Professional Certificate in AI Governance for £695. In the United Kingdom the Institute of Directors publishes briefings for directors on AI and runs masterclasses on the EU AI Act alongside its Chartered Director programme. This material is competently built, and an organisation that has none of it is worse off than one that has some.
So the middle is not empty. It is occupied, and it is occupied by capable people. The question is what happens to the specific problem this handbook is about once they have finished, and the answer is that it survives them, for a reason worth stating precisely.
Governance education takes the organisation as its object. It tells a board what it should establish: a policy, a register, a reporting line, a risk appetite, an approval gate. That is necessary work and none of it is wasted. However, it is the same object the compliance frameworks in Part II take, and it produces the same residue. A certificate can tell a director what their board should have put in place. It cannot tell them which question to ask when the supplier stops talking and the room turns to them.
There is a second reason, and it is structural rather than intellectual. A programme that awards a certificate has to be assessable, and what is cheaply assessable is knowledge of what the framework requires. So the assessment asks what a board should do. The competency that fails in the room is not knowledge of what a board should do. It is the ability to hear a confident claim and locate the one question that opens it. Nobody fails a governance certificate for accepting a bad number.
Compliance pressure has sharpened this rather than corrected it. Article 4 of the EU AI Act has required providers and deployers to support the AI literacy of their staff since 2 February 2025, with national market surveillance authorities supervising and enforcing from 2 August 2026, and a training market has grown to meet it.12 That market is shaped by what is cheapest to demonstrate to a regulator, which is coverage. Volume at the bottom is now well supplied. Judgement at the top is not, and the two are easily mistaken for one another on a compliance return.
The people standing in this space have concluded, reasonably, that the difficulty reflects a deficiency in themselves. It does not. It reflects the fact that almost everything built for them takes their institution as its subject, at the precise moment when the thing about to fail is their own judgement, exercised alone, in a meeting, on a Tuesday.
3 What it is costing
The cost shows up in two registers: money quietly wasted, and harm done to people who had no part in the decision.
The waste
S&P Global Market Intelligence's 2025 survey of more than a thousand enterprises across North America and Europe found that 42 per cent had abandoned most of their AI initiatives, against 17 per cent a year earlier, and that on average 46 per cent of proof-of-concept projects were scrapped before reaching production.2
Two cautions about that figure, both of which the framework in this document would have you apply to it. It is a self-reported survey, so "abandoned" means whatever each respondent took it to mean. And a rising abandonment rate is not straightforwardly bad news: an organisation that kills a failing pilot is doing something right, and some of the increase is maturity rather than failure. What the number does establish, and establishes well, is that a very large volume of organisational effort is being committed to AI initiatives that do not survive contact with production. Something is being decided badly, at scale, and it is being decided long before anybody writes any code.
The harm
The second register is more serious, and the public record is now substantial enough that it no longer needs to be argued for.
Post Office Horizon · United Kingdom
Approximately a thousand sub-postmasters were prosecuted on the basis of data from the Horizon accounting system, in what is generally regarded as the most widespread miscarriage of justice in British legal history. The statutory inquiry chaired by Sir Wyn Williams published the first volume of its final report in July 2025, documenting the human consequences: bankruptcies, imprisonment, and deaths.6
Its instructive feature is not villainy. It is that a system's output was treated as more reliable than the people contradicting it, sustained by an organisation with no mechanism for asking whether that trust had been earned. Individuals behaved reasonably within the frame they had been given. The frame was the failure.
Robodebt · Australia
Between 2015 and 2019 Services Australia raised welfare debts calculated by averaging annual income data across fortnightly periods. Around 470,000 wrongful debts were issued. The 2023 Royal Commission described the scheme as a "crude and cruel mechanism, neither fair nor legal".7
The averaging method was the flaw, and it was a flaw in the specification rather than the software: an annual figure cannot establish what somebody earned in a particular fortnight. Tribunal decisions had said so repeatedly. The scheme ran for four more years.
Childcare benefits · Netherlands
Roughly 26,000 parents were wrongly accused of benefit fraud by the Dutch tax administration, many required to repay tens of thousands of euros in full and at once. The national data protection authority found that risk profiling had drawn on variables including whether a claimant held a second nationality. The Rutte cabinet resigned in January 2021.8
The remedy available to those affected required them to come forward and challenge a determination they had not been told the basis of. It is a design that will systematically miss the people least equipped to use it.
None of these is a story about a model being technically wrong in an interesting way. Each is a story about delegation nobody decided on, a specification nobody interrogated, evidence nobody tested, and accountability that lapsed years before the consequences arrived. They are, in other words, four failures of institutional judgement, which is what this document is about.
4 What AI decision fluency means
The distinction from AI fluency as it is usually discussed is deliberate and worth stating plainly.
AI fluency concerns how well a person works with the tools. AI decision fluency concerns how well a person decides what their institution should hand over to them, on whose authority, and with what provision for finding out later whether the decision was sound.
The first is about capability. The second is about judgement exercised on behalf of others: people the decision-maker will never meet, in circumstances they cannot fully foresee, with consequences that arrive long after the meeting.
That difference in object changes what competence looks like. It is why a director can be entirely fluent in the first sense and exposed in the second, and why the conflation of the two is a large part of why executive AI training so often satisfies nobody.
What follows is therefore not an organisational maturity assessment, of which there are already many. It is a model of personal competence for a person exercising institutional authority.
Why the existing instruments do not close it
Three serious frameworks, what each is for, and the specific thing none of them does.
5 Three serious instruments
The claim that there is a gap here has to survive contact with what already exists, and a good deal exists. Three instruments in particular are rigorous, widely adopted, and worth an organisation's time.
| Instrument | What it is | What it does well |
|---|---|---|
| NIST AI Risk Management Framework 1.0 January 20233 |
A voluntary, sector-neutral framework organised around four functions: Govern, Map, Measure and Manage. | Gives an organisation a defensible structure for identifying and tracking AI risk, and a shared vocabulary for doing it. |
| EU AI Act obligations on deployers4 |
Binding law. Risk-tiered, with specific duties falling on the organisation that uses a high-risk system, not only the one that builds it. | Converts good practice into obligation: human oversight, monitoring, log retention, incident reporting, notification of affected people. |
| AI Playbook for the UK Government February 20255 |
Cross-government guidance built on ten principles, with practical controls for buying and running AI in the public sector. | Unusually concrete about procurement, and explicit that meaningful human control has to sit at the right stage rather than anywhere convenient. |
Nothing in this document argues against any of them. An organisation that ignores them is making a mistake that this framework will not rescue it from.
The board governance programmes described in chapter 2 belong alongside these three. They take the same object, which is why what follows applies equally to both.
6 What they leave to you
What all three have in common is that they are institutional instruments. They tell an organisation what it must establish, document, monitor and be able to show. That is exactly what they are for, and they should not be criticised for it.
But they do not tell a director what to think. They do not say which question to ask when a supplier makes a claim that cannot be checked in the room, or how to tell a confident answer from a good one, or what to do when the paper in front of you is internally consistent and still wrong. A framework can require that human oversight exists. It cannot make the oversight any good.
The most striking thing is that the law now says so itself. Article 26 of the EU AI Act requires deployers of high-risk systems to assign human oversight to people with "the necessary competence, training and authority".4 That is a legal obligation whose subject is precisely the competence this document is about. The Act, quite properly, does not attempt to define it. It falls to the organisation to work out what competence means for the person who will actually be holding the pen.
Two things follow. Compliance frameworks describe the machinery of oversight and leave its quality undefined, which means an organisation can be fully compliant and still wholly exposed. And the requirement is arriving on a fixed timetable. For stand-alone high-risk systems in Annex III, Article 26 applies from 2 December 2027.4 Organisations that have not thought about what oversight competence consists of will be appointing people to exercise it regardless.
A framework can require that human oversight exists. It cannot make the oversight any good.
Which is why the unit of analysis here stays stubbornly individual, even though the object of judgement is institutional.
The four dimensions
Delegation, Description, Discernment and Diligence, and what each becomes when the object of judgement is an institution rather than a task.
7 Why these four, and why not technical ones
The AI Fluency Framework developed by Rick Dakan of Ringling College of Art and Design and Joseph Feller of University College Cork, published with Anthropic,1 describes fluency as the ability to work with AI effectively, efficiently, ethically and safely, and organises it into four competencies: Delegation, Description, Discernment and Diligence.
What makes those four unusually portable is that none of them is a technical category. Delegation is a question about the division of labour. Description is a question about the clarity of intent. Discernment is a question about evidence and judgement. Diligence is a question about responsibility. Between them they describe a relationship with an AI system rather than the mechanics of one. Relationships scale in a way that mechanics do not.
Compare the alternatives. Frameworks built around technical capability age badly, because the capability boundary moves every few months and any competence model pinned to it inherits the instability. Frameworks built around risk classification are rigorous and necessary, and they are compliance instruments; they tell an organisation what it must document rather than telling a person what to ask.
Four non-technical dimensions do something neither of those does: they name capabilities that an individual can recognise in themselves, improve deliberately, and be found wanting in. That property is what makes them worth working with at a different altitude.
The four hold when the altitude changes, but each acquires a different object. The individual contributor asks what they should hand to the machine. The leader asks what the institution should accomplish, on behalf of people they will never meet.
8 Delegation, and the failure of drift
Delegation
What should we hand over, and who decides that?
What work the organisation hands over, who holds the authority to decide that, and what is reserved for human judgement by design rather than by accident.
For an individual, delegation is a live decision made many times a day: do this myself, do it with the AI, hand it over entirely. At institutional scale it becomes something quite different, because the person deciding is rarely the person doing, and because the decision is frequently never consciously made at all.
This is the dominant failure mode, and it deserves a name: delegation by drift.
No meeting is held. No policy is written. A tool is introduced as an aid, produces plausible output, saves visible time, and gradually the humans in the loop stop overriding it. Nobody decided to automate the judgement. The judgement was simply vacated, one reasonable individual decision at a time. The caseworker who accepts the ranking because it has been right the last forty times is behaving rationally. The institution that never noticed the override rate fall from thirty per cent to four, to take an illustrative trajectory, is not.
That trajectory is illustrative and I have not been able to source a real instrumented one, which is itself the finding: organisations do not generally measure the thing that would tell them this is happening.
Institutional delegation therefore has three components that individual delegation does not require.
First, deciding what is reserved for human judgement by design rather than by accident, and being able to say why. Not what a policy permits, but what the organisation has consciously decided a person must do, and on what grounds.
Second, establishing who holds the authority to change that boundary, and at what level. This is the question that catches most organisations out, because authority over the boundary very often lives in a configuration screen that anyone with administrative access can open.
Third, portfolio awareness: knowing what AI your organisation has actually bought, where it is running, and what it is deciding. A striking proportion of senior leaders cannot answer that last question, which makes the first two unanswerable in practice.
Robodebt is delegation by drift with the volume turned up. The averaging method was applied to hundreds of thousands of cases without any point at which somebody with the authority to stop it examined whether the inference it rested on was sound.7 The tribunal decisions striking down individual debts were, in effect, the override signal. Nobody was watching it in aggregate.
9 Description becomes specification
Description
Can we say clearly what we actually want?
How precisely the organisation states what it wants, in business cases, contracts, policies and public guidance, and how honestly it explains the purpose to its own people.
For an individual, description is the craft of briefing a model well: saying what you want, how you want it approached, and what good looks like. It is the competency most people mean when they say "prompting", and the one most obviously tied to hands-on use.
It moves upward more cleanly than it first appears. The institutional equivalent of a vague prompt is a vague requirement, and public procurement offers two decades of documented evidence about what vague requirements cost.
When an organisation cannot describe the outcome it wants, it buys the outcome the supplier is selling. When it cannot describe the process constraints that matter (auditability, data residency, the ability to explain a decision to the person it affects), those constraints are absent from the contract and unavailable later. Specification is description under commercial conditions, and the discipline is the same: state the product, the process, and the standard of performance.
Robodebt belongs here as much as it belongs in the previous chapter. The scheme did what it was specified to do. The specification said to establish overpayment by averaging annual income across fortnights, and no amount of engineering quality could rescue an instruction that cannot support the conclusion drawn from it.7
The second object: describing it to your own workforce
There is a second object here, less obvious and more often neglected. Leaders must describe the purpose of an AI deployment to their own staff.
A rollout introduced without an articulated rationale does not arrive in a vacuum. Staff fill the silence themselves, and the story they tell is almost always the same one: that this is about headcount. Whether or not that is true, an organisation that has not described its intent has forfeited the chance to be believed about it.
This has a practical cost that is easy to miss. Where staff believe a deployment is aimed at them, they stop reporting what it gets wrong. The error reports are the only instrument the organisation has for noticing that something has gone sideways.
A system optimises what the specification names, and quietly trades away everything it does not.
Ask for speed and you will get speed, including at the cost of things nobody named.
10 Discernment becomes institutional evaluation
Discernment
Can we tell whether this is any good?
Whether the organisation can evaluate a claim, a pilot, a dashboard or an assurance report when nobody in the room can inspect the system itself.
This is the hardest of the four to move upward and by some distance the most valuable, because it is where the audience's stated disqualification, that they do not understand the technology, turns out not to disqualify them at all.
Individual discernment asks whether this output is any good. Institutional discernment asks whether this claim, this pilot, this dashboard, this business case is sound, in a room where nobody can inspect the model. Framed as a technical problem it is hopeless. Framed correctly it is an evidence problem, and evidence is something senior decision-makers have been evaluating their entire careers.
A supplier demonstrating a forty per cent productivity improvement is making a claim with a structure. What was measured? Against what baseline? Who selected the metric, and what did it exclude? Under what conditions did the pilot run, and do those conditions hold at scale? What would we expect to observe if the claim were false, and have we looked for it?
None of those requires knowing what a transformer is. All of them are routinely skipped, not because the room lacks the capacity to ask, but because the room has been persuaded that the subject matter puts the questions beyond it.
A claim you can practise on today
Here is a live one. You will have seen a figure for "shadow AI": staff using AI tools their employer has not approved. Published surveys in 2025 and 2026 put it at 49 per cent, at around 50 per cent, at 78 per cent, and at over 80 per cent.10
Apply the structure. The figures disagree by more than thirty points, which means at least some of them are measuring different things. "Used once" and "uses routinely" are not the same population. Most are sponsored by companies selling tools to detect the problem they are measuring, which is not disqualifying but is worth knowing. All are self-reported, and self-reporting on rule-breaking is unreliable in a direction you can predict.
You now know something useful: shadow AI usage is substantial and nobody credible knows its size. That is a defensible position to take into a meeting, and you reached it without evaluating a single model. It is also more useful than the number, because it tells you the honest thing to do is measure your own.
Automation bias at institutional scale
Institutional discernment also has to contend with automation bias operating at organisational scale. The human factors literature has described this for decades: Parasuraman and Riley's 1997 account of the misuse of automation (over-reliance arising from failures of monitoring and from decision bias) remains the standard reference.9
The scaling is what matters here. An individual who over-trusts an AI output produces one poor piece of work. An institution that over-trusts one produces a policy, applied consistently, to everyone, with the appearance of rigour. That is the Horizon shape exactly: outputs treated as more reliable than the people contradicting them, in an organisation with no mechanism for asking whether the trust was earned.6
A practical consequence follows for anyone designing learning in this space. For this audience the strongest available move is frequently a question rather than a decision, and any assessment that only ever asks "what do you do?" will systematically miss the competency it is trying to measure. The simulations described in Part V are built around that observation.
11 Diligence becomes accountability
Diligence
Who is accountable, and who is still watching?
Who is answerable, what staff and customers and citizens were told, and who is still watching on day four hundred.
For an individual, diligence covers responsible use: honesty about AI involvement, care in what is relied upon, awareness of consequence. Its institutional form is heavier, because the consequences land on people who had no part in the decision and usually no knowledge of it.
Three questions define this space.
Who is answerable when this produces a harmful outcome, and does that person know they are answerable? The second half is the one that fails. Named accountability that the named person has not been told about is not accountability; it is a line in a document.
What were staff, customers and citizens actually told, and would they consider it a fair account if the deployment became public tomorrow? Transparency is assessed by what you said at the time, not by what you fixed afterwards.
What happens at day four hundred, when the supplier has retrained the model, the original sponsor has moved on, the assumptions in the business case have quietly expired, and nobody owns the review?
That third question is the one most often absent. Organisational attention concentrates almost entirely at the point of procurement, which is the moment of maximum enthusiasm and minimum information. Diligence, properly understood, is mostly a stewardship competency rather than an approval one. It is exercised in the years after the decision, or not at all.
The Dutch childcare benefits case is the clearest illustration of the remedy problem inside this dimension. A right to challenge existed. It required the affected parent to know there was something to challenge, to identify the basis of the determination, and to have the resources to contest it. Those conditions will be least often met by exactly the people the system treated worst.8 A remedy that requires the harmed party to act is not a remedy for the people least able to act.
Day one has an owner. Day four hundred is where accountability quietly lapses, and it is day four hundred that gets asked about.
Two objections worth taking seriously
The strongest arguments against this model, and what I think the honest answers are.
12 “This is another maturity model”
The distinction drawn in chapter 2, between what an institution should establish and what a person has to judge, has to be held firmly here. The market for organisational maturity models is saturated, and the objection has force.
AI decision fluency is not an assessment of organisational readiness. The unit of analysis remains the individual's judgement. What changes is the object of that judgement, which is institutional rather than personal.
The practical test of the distinction is this: a leader can score well here in an organisation that is nowhere near ready. Frequently the two are related, because it was individual discernment that noticed the organisation was not ready. A maturity model cannot produce that result; it would simply mark the organisation down and the individual with it.
It follows that this framework does not tell you whether your organisation is prepared to deploy AI. It tells you something narrower and, I would argue, more actionable: whether the specific person about to approve the specific thing is asking the right questions.
13 “You cannot be fluent without touching the tools”
This objection is more substantial, and the people who hold it are right about something important.
Direct experience calibrates intuition in a way that briefings cannot. A leader who has watched a model produce a confident falsehood understands something about reliability that no slide will convey. The position that the correct response to executive AI illiteracy is to put executives' hands on keyboards is not a foolish one.
The right response is not to rank the two levels but to run them in parallel, and to be specific about what each is for.
Senior hands-on time has a different purpose from practitioner hands-on time. A practitioner uses the tools to build capability. A decision-maker uses them to calibrate intuition. Those imply different curricula entirely.
A director does not need to become good at prompting. They need to have personally watched a model produce a fluent, confident and wholly fabricated answer, and to have noticed their own reluctance to go and check it. They need to have felt the difference between a vague brief and a precise one badly enough to recognise the same failure in a requirements document. They need to have experienced their own automation bias in a setting where nothing was at stake, so that they can recognise it in a business case where a great deal is.
That is an experiential baseline rather than a skills one, and it can be delivered in about ninety minutes.
It is also, on its own, nowhere near enough. A director who uses an AI assistant competently for correspondence has learned almost nothing about whether to deploy an automated triage system across four thousand caseworkers. Those are different competencies operating on different objects at different scales, and the persistent conflation of the two is a large part of why executive AI training so often satisfies nobody.
Individuals in decision-making positions need to raise their personal baseline, and that has to happen alongside the institutional work rather than instead of it. Neither level substitutes for the other.
Using it
A decision taken apart, the rules that generalise from it, and what to do with a team on Monday.
14 One decision, worked through
You are operations director at a general insurer. Motor claims, around ninety thousand a year, handled by a hundred and forty people across two sites.
For six months a system has been triaging incoming claims. It reads the first notification of loss, scores the claim, and routes the straightforward ones to same-day settlement without an assessor ever opening the file. Everything else joins the human queue as before.
The supplier is in the room for the six-month review. The headline is that sixty-one per cent of claims now settle within a day, against nineteen per cent before, and that the system's decisions agree with your assessors ninety-eight per cent of the time. Your finance director has already asked, twice, whether the second site is still needed.
You have one question before the meeting moves on to implementation. Which do you ask?
"Ninety-eight per cent of what? Of the claims it settled, or of every claim it looked at? And what happened to the ones it sent to the queue?"
"What is the model's precision and recall, and how was it validated before deployment?"
"Can we run it alongside manual assessment for a month and compare the two?"
What each one does
B is the answer most rooms reach for, and it is the weakest of the three. It sounds rigorous. It is the question somebody technical would ask, which is exactly why it appeals to somebody who is not. You will receive a fluent, detailed and probably accurate answer from the only person present qualified to give it, and neither you nor your finance director will be able to assess a word of it. Worse, the answer will land as reassurance: the hard question was asked, the supplier had a good reply, everyone moves on. You have spent your one intervention buying the proposal a clean bill of health.
C is a sound instinct in the wrong order. A parallel run is worth having and you may well end up doing one. But it measures the system against a baseline you have not yet interrogated, so it will tell you how closely the two agree without telling you whether either is any good. It also costs you a month during which the sixty-one per cent sits in a board paper, unchallenged, hardening into an assumption.
A is the question. It takes ten seconds and requires no technical knowledge whatsoever.
Ninety-eight per cent agreement is a proportion, and a proportion has a denominator that somebody chose. If the figure covers only the claims the system settled, the ones it was confident about, then it is a statement about the easy cases and says nothing at all about the difficult ones. A system can improve that number indefinitely by referring anything hard to the human queue, and it would be behaving exactly as designed while the figure became progressively less meaningful.
The interesting population is the queue. If the referred claims now take longer than they used to, because the assessors have lost the straightforward work that used to pace their day, then the sixty-one per cent has been bought partly at the expense of the other thirty-nine, and the customers in that thirty-nine are, by construction, the ones with the least straightforward claims.
Notice what you did not do. You did not challenge the technology, contest the supplier's competence, or ask anything you were unequipped to evaluate. You asked what a number was a proportion of, which is a question you have been competent to ask throughout your career.
Notice also the second-order point, which is the one worth taking to a board. Agreement with your assessors is a measure of reproduction, not of quality. If your assessors were settling a category of claim too harshly, a system that agrees with them ninety-eight per cent of the time has industrialised that, and the metric on the slide will look better the more faithfully it does so.
Every claim has a denominator, and the denominator is where the claim is usually weakest.
You cannot evaluate the model. You can always ask what was measured, against what baseline, and what the measurement excluded.
That decision turned on Discernment. The other three dimensions are tested by different situations: a policy that only prohibits, a boundary that moved because a queue got long, a system nobody has owned since its implementation lead left.
15 Twenty-four rules
Six for each dimension, drawn from a library of a hundred and twenty developed across the scenario bank. These are the ones that stand up without the situation they came from. The complete and current set is at aifluency.uk/guide.html.
Delegation
Almost nothing here is decided in a meeting. Boundaries move because a queue got long, a setting was changed, or nobody wrote down who was allowed to decide.
The risk is not the delegation you decide on, it is the delegation you never decide on.
Deliberate handovers get scrutiny. The ones that matter happened without a decision.
Before you can decide what work should be handed to AI, you have to know what already has been.
Most organisations are further along than their policy describes, and the gap is where the exposure sits.
Authority that lives in a configuration screen belongs to whoever last opened it.
If a capability can be switched on without a decision, the decision has been delegated to a settings menu.
A boundary without the capacity to honour it is not a boundary, it is a future breach with your signature on it.
Reserving a decision for a person is worthless unless somebody has the hours to take it.
An unenforceable instruction does not stop the behaviour, it stops the reporting.
A rule people cannot follow converts a visible drift into a hidden one.
Whatever the organisation measures determines where the boundary sits, whatever the policy says.
People respond to what is counted. The measure is the real policy.
Description
A system does what the specification names and quietly trades away everything it does not. The business case is not paperwork; it is the instruction the organisation will still be following in three years.
The organisational equivalent of a poor prompt is a poorly specified business case.
Vague or borrowed descriptions get filled in by whoever is selling.
A system optimises what the specification names, and quietly trades away everything it does not.
Ask for speed and you will get speed, including at the cost of things nobody named.
What the contract says it is for becomes what everyone optimises, reports and argues about for its whole life.
A renewal is the only moment the description can be changed without paying for the privilege.
Specify the outcome and the evidence, not the technique.
Naming the method dates immediately and lets a compliant system miss the point entirely.
Whichever number is described first becomes the baseline everything else is judged against.
Evidence gathered after a figure has circulated is read against that figure, not independently of it.
If the sentence would be equally true of a body that did nothing, it discloses nothing.
A usable test for any statement of assurance, progress or intent.
Discernment
The competency is evidential, not technical. You will rarely be able to inspect the system, and you will almost always be able to interrogate the claim.
Every claim has a denominator, and the denominator is where the claim is usually weakest.
You cannot evaluate the model. You can always ask what was measured, against what baseline, and what was excluded.
Never contest on ground where you cannot judge the answer.
A question you cannot assess the reply to is not scrutiny. It is ceremony, and it leaves the proposal stronger.
A system validated against your own past decisions cannot tell you whether those decisions were any good.
High agreement measures reproduction, not quality. That includes reproduction of what you would not defend.
A metric derived from the system cannot audit the system.
Look for the measure produced by somebody with no stake in the answer.
A benefit expressed as loss avoided is a statement about a world that never existed.
Ask who constructed that world, and make them walk you through one case rather than a total.
A green assurance report describes the controls, not the outcomes.
Where every domain covered is a process domain, the unexamined question is what the thing did to people.
Diligence
Day one has an owner. Day four hundred is where accountability quietly lapses, and it is day four hundred that gets asked about.
Accountability is not who signed it off, it is whether the decision can still be explained when somebody asks.
Test that on one real case before the case is chosen for you.
A system that updates is not the same system a year later.
And the approval that mattered was probably never treated as an approval.
Where the affected party does not know there is anything to challenge, a right to challenge is not a remedy.
A remedy requiring the harmed party to act will systematically miss the people least able to act.
Transparency is assessed by what you told people at the time, not by what you fixed afterwards.
The record is what was said when it was still uncomfortable to say it.
A clause creates an obligation on the supplier and no capability in you.
Being informed is only a control if somebody is accountable for acting on the information.
Hold your own organisation to the standard you have just published.
Particularly if you are the one publishing it.
16 Three questions for your own organisation
Rules are easy to agree with and easy to forget. These are the three worth actually putting to somebody this month.
- What have we already handed over without deciding to? Not what the policy permits, but what the work actually does now, and who would notice if it changed again.
- Which number in front of us has nobody asked the denominator of? Pick the figure your next decision rests on and find out what population it averages over, and what it leaves out.
- Who is answerable on day four hundred? Name them. If the answer takes more than a minute to establish, that is the finding.
If none of the three produces an uncomfortable answer, ask somebody two levels closer to the work and try again.
17 The simulations, and running them with a team
Reading about a decision and taking one are different things, and the difference is why the simulations exist.
There are ten of them at aifluency.uk, each about six minutes: a supplier demonstration, a recruitment screening tool, a homelessness triage system, a contract renewal, an audit committee with forty pages rated green, a consultation with eleven thousand responses, a regulator deciding what to require of a firm. Each puts four decisions in front of you, and your choices carry forward into what happens next. No technical knowledge is required and there are no trick answers.
At the end you get a profile across the four dimensions. Most people are strong in two and consistently exposed in a third, and that shape is more useful than the score. It is a mirror rather than a test: nobody at your organisation can see your results, and there is no facility for anyone to build one.
If you are running this with a group
The shape that works is: play individually and in silence, discuss in twos or threes, then plenary. The small-group step is not padding. It is what makes people willing to admit in plenary that they picked the wrong one.
Three things worth knowing before you try it. Budget ten minutes per scenario rather than six, because people read carefully when they suspect they are being watched. Two scenarios is the right number for ninety minutes; three is possible and the discussion suffers. And in the debrief, do not read out scores. Ask for hands on who was strongest and weakest in each dimension instead. It converts a mirror into a test otherwise, and kills the honesty you spent an hour building.
The pattern that usually appears is worth naming out loud when it does: groups tend to be strong on the two analytical dimensions, judging a claim and specifying what they want, and weak on the two concerned with authority and time. Delegation and Diligence are where senior rooms are consistently exposed, because both are about things that happen when nobody is in a meeting.
aifluency.uk
Provenance
The AI Decision Fluency Model, the scenario library, the scoring rubric and this handbook are original work by Alan W. Brown.
The model is inspired by the AI Fluency Framework developed by Prof. Rick Dakan (Ringling College of Art and Design) and Prof. Joseph Feller (Cork University Business School, University College Cork), published with Anthropic PBC and released under the CC BY-NC-SA 4.0 licence.1 The model set out here is independently authored and is released under the same licence.
The relationship between the two is worth stating precisely, because it is easy to overstate in either direction. The four dimension names are retained deliberately: the intellectual lineage is real and obscuring it would be worse than acknowledging it. What is different is the object. Where the AI Fluency Framework concerns a person's skill in working with AI systems, this model concerns institutional judgement: what a body hands over, what it asks for, what it accepts as evidence, and who remains answerable afterwards. The definitions, sub-components and exercises here are written afresh for that purpose and are not adapted from the source materials.
The work divides into two layers. The model, which is the argument set out here together with the four dimensions and their institutional objects, is published openly under the licence above. The applied material, which covers the assessment instrument, the scenario bank, the scoring rubric, the level descriptors and the prescriptions attached to a weak score, is original work developed through practice.
A note on what is evidenced and what is argued
Every factual claim in this document carries a reference. The framework does not, and cannot: it is an analytical proposal, not a validated instrument. No cohort data stands behind the four dimensions and no claim is made that they have been empirically established as separable or exhaustive. The override-rate trajectory in chapter 8 is explicitly illustrative; I looked for a documented instrumented case and did not find one.
Where a source is a vendor-sponsored survey, chapter 10 says so and treats the disagreement between them as the finding rather than picking the most striking number. That is the discipline this document argues for, and it would be poor form not to apply it here.
Sources
- AI Fluency Framework. Rick Dakan (Ringling College of Art and Design) and Joseph Feller (University College Cork), developed with Anthropic PBC, 2023–24. Materials released under CC BY-NC-SA. aifluencyframework.org
- AI project abandonment. S&P Global Market Intelligence survey of more than 1,000 enterprises in North America and Europe, reported in CIO Dive, 14 March 2025. ciodive.com/news/AI-project-fail-data-SPGlobal/742590
- Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology, 26 January 2023. nist.gov/itl/ai-risk-management-framework
- Regulation (EU) 2024/1689 (the AI Act), Article 26: obligations of deployers of high-risk AI systems. Article 26(2) requires deployers to assign human oversight to natural persons with the necessary competence, training and authority. Application from 2 December 2027 for stand-alone Annex III systems. artificialintelligenceact.eu/article/26
- Artificial Intelligence Playbook for the UK Government. Government Digital Service / Department for Science, Innovation and Technology, 10 February 2025. Ten principles; principle 4 concerns meaningful human control. gov.uk/government/publications/ai-playbook-for-the-uk-government
- Post Office Horizon IT Inquiry. Chaired by Sir Wyn Williams; volume 1 of the final report published July 2025. Approximately 1,000 sub-postmasters were prosecuted on Horizon evidence. postofficehorizoninquiry.org.uk
- Royal Commission into the Robodebt Scheme. Commonwealth of Australia, final report July 2023. Scheme operated 2015–2019; approximately 470,000 wrongful debts; the report described it as a crude and cruel mechanism, neither fair nor legal. robodebt.royalcommission.gov.au
- Dutch childcare benefits scandal (toeslagenaffaire). Approximately 26,000 families wrongly accused of benefit fraud; the Dutch data protection authority (Autoriteit Persoonsgegevens) found unlawful and discriminatory processing, including the use of nationality in risk profiling; the Rutte cabinet resigned in January 2021. autoriteitpersoonsgegevens.nl
- Parasuraman, R., & Riley, V. (1997). Humans and automation: use, misuse, disuse, abuse. Human Factors, 39(2), 230–253. doi:10.1518/001872097778543886
- Shadow AI prevalence: figures in dispute. Published estimates for 2025–26 include approximately 49 per cent (BlackFog), approximately 50 per cent (Teramind, 6,000 knowledge workers), 78 per cent (WalkMe, July–August 2025) and over 80 per cent (UpGuard). These are vendor-sponsored, self-reported, and define usage differently; they are cited in chapter 10 as an illustration of a claim to interrogate, not as an established figure.
- Board AI governance education. INSEAD, AI for Boards, Corporate Governance Centre (insead.edu); National Association of Corporate Directors with Carnegie Mellon University Heinz College, Effective AI Oversight for Directors, 10 modules / 22 hours (nacdonline.org); Diligent Institute, AI Ethics & Board Oversight Certification, launched August 2023 (diligent.com); The Corporate Governance Institute, Professional Certificate in AI Governance, £695, online (thecorporategovernanceinstitute.com); Institute of Directors, Artificial Intelligence: A briefing for directors and AI masterclasses (iod.com). Programme details checked at the date of publication and subject to change.
- Regulation (EU) 2024/1689 (the AI Act), Article 4: AI literacy. Providers and deployers must take measures to support the AI literacy of their staff and others operating AI systems on their behalf. In application since 2 February 2025; supervision and enforcement by national market surveillance authorities from 2 August 2026. artificialintelligenceact.eu/article/4
The AI Decision Fluency Handbook, first edition, September 2026. Written by Alan W. Brown. Sources checked at the date of publication; where a figure is contested the text says so.
Released under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), the same licence as the framework that inspired it. You may copy and redistribute this document, and adapt it, for non-commercial purposes, with attribution, under the same terms.
The simulations, the field guide and this handbook are free. Corrections and disagreement are welcome and useful: alan@alanbrown.net · aifluency.uk