Measuring AI fluency
Most companies I talk to are a year into an AI push. They have bought licences, run training, appointed someone to own it. Ask them whether fluency actually went up and the room goes quiet, because nobody measured where it started.
That is a strange position for HR to be in. Measuring whether a capability improved is the one thing HR is supposed to be good at.
The 4D framework
Anthropic built a framework for this with Rick Dakan and Joseph Feller. They call it the 4D. It treats working with AI as a learnable skill set with four parts, and it comes with an assessment, so you can find out where people actually sit.
Read the four out loud to a room and watch what happens. Almost everybody assumes AI skill means prompting. Only one of the four is prompting.
Delegation
Deciding what to hand to AI, and what to keep.
This is the one people skip, and it is the one that separates someone useful from someone who has simply been given a tool. It asks two questions at once. Is this task a good fit for a model at all, and am I the right person to be deciding that?
In HR the cost of getting this wrong is high in a specific way. Handing a model a drafting job is fine. Handing it a decision about an individual, or a piece of work where the input is personal data that should not be leaving your building, is a different matter, and no amount of skill at the other three dimensions rescues you.
Good delegation also means knowing when to do the work yourself because the thinking is the point. Some tasks exist so that the person doing them understands the problem afterwards.
Description
Telling it what you want, and how you want it worked.
This is prompting, and it is the part everyone means when they say AI skills. It matters. Being able to say what good looks like, supply the right context, set constraints, give an example, and ask for the working rather than only the answer makes a large difference to what you get back.
It is also the most teachable of the four, which is why training goes here. A day of practice moves people a long way. That is worth doing, and it is worth being honest that you have then improved one quarter of the thing.
Discernment
Judging what came back.
This is where most of the risk lives. A model produces something fluent every time, whether or not it is right, and fluency is persuasive. Discernment is the habit of asking what would have to be true for this to be correct, spotting the number that is confidently wrong, noticing the citation that does not exist, and recognising the answer that is plausible because it is average rather than because it fits your situation.
People with deep domain knowledge are usually good at this immediately, in their own domain, and poor at it everywhere else. That is worth remembering when you decide who to put in front of a model and on what.
I find this the hardest of the four to raise through training, because it is mostly expertise wearing a different hat. What does help is making people show their checking, out loud, in front of each other.
Diligence
Owning the result: responsibility, transparency, accountability.
Whoever put the work out is answerable for it. That means being open about what was AI assisted, keeping enough of a trail that a decision can be explained later, and not treating a model’s involvement as a reason the answer was somebody else’s fault.
For HR this is the dimension with a compliance edge. If a model touched a process that affects people’s pay, progression or employment, you will at some point be asked to explain how. Diligence is having an answer ready before you are asked.
What to do with the score
An assessment gives you a number per person across four dimensions. That number on its own is mildly interesting and easy to over-read.
It gets interesting the moment you connect it to everything else you already know about your workforce, which is exactly the kind of joining HR should be able to do. Four questions I would ask first:
| Question | Why it matters | |
|---|---|---|
| 1 | Are we losing people at fluency level one, or level five? | Tells you whether your AI push is retaining your best people or driving them out |
| 2 | Do we reward higher fluency? | If pay and promotion ignore it, the training will not stick |
| 3 | Do our most fluent people sit in the most critical positions? | Capability in the wrong place is not capability |
| 4 | Are we still hiring and promoting people who are not fluent at all? | The intake decides where you are in three years |
Every one of those needs the fluency score joined to turnover, reward, performance, position criticality and hiring. None of them are answerable from the assessment alone.
How I would run it
Baseline this month, before any more training. Take the assessment across a population you care about, and record where people start, because a year from now the only honest way to say whether it worked is to compare.
Then measure again in six months. Split the result by the four dimensions rather than reporting one blended number, because the shape tells you what to do next. A team that scores well on description and badly on discernment needs something completely different from a team with the opposite shape.
And expect the leadership scores to be the awkward part. Fluency is cultural, and culture follows what leaders actually do rather than what they announce. If the executive team has signed the budget and delegated the whole thing, that shows up in the numbers, and it is the most useful thing the assessment will tell you.