Skip to main content
Framework

How we measure, and what we refuse to claim

Published before there is anything to measure. Every rule below is read straight from the code that enforces it, so the version you are reading cannot quietly differ from the version we apply.

There are no results on this page

That is deliberate. A public counter needs a definition, a period, a source and a date it was last updated, and we have no deliveries behind one yet. Publishing the method first is the part we can do honestly now — and it means the standard was set before anybody had a number they wanted it to accommodate.

First

What kind of thing a number is

Six levels. Every one of them is worth knowing, none of them is the one below it, and the collapse always runs the same direction — attendance gets reported as transformation, because attendance is the number you have on the day.

activity

01

Something happened. People came, a session ran, a tool was opened.

Not reach, not engagement, and never impact. Forty people in a room is forty people in a room.

output

02

Something was produced. A workflow designed, a policy drafted, a scan completed.

Not adoption. A workflow designed and never run is a document.

behaviour

03

Somebody did it again without being asked. The first level that is not about us.

Not capability. Repeating something is not the same as being able to explain or adapt it.

capability

04

Somebody can now do a thing they could not, and showed it to a person who checked.

Not outcome. Being able to do it is not evidence that anything changed because of it.

outcome

05

Something measurable changed in the work — time, error rate, response time, adoption.

Not impact. Four hours reclaimed is a fact about a week, not about a community.

impact

06

The change mattered to somebody beyond the person who made it. The hardest to evidence and the easiest to assert.

Not a total. It is the level most often claimed and least often measured.

The registry

9 metrics, not five hundred

Each one exists because a decision depends on it. Anything added later has to answer what decision it changes. The definition matters more than the number: a design counted as a deployment inflates a figure without anybody lying.
The metric registry: what each metric means, which level it sits at, and whether it may be combined across cases.
MetricLevelDefinitionCombined across cases
AttendancepeopleactivityPeople present for at least half a session.Yes, where the cases share a period and a method.
Workflows designedworkflowsoutputCompleted workflow blueprints. A design, not a deployment — the distinction is the whole point of publishing the definition.Yes, where the cases share a period and a method.
Workflow still in use at 30 daysbehaviourA workflow the person reports still using a month after designing it.Yes, where the cases share a period and a method.
Competency demonstratedcapabilityA competency confirmed against its criteria by somebody with standing to confirm it.Yes, where the cases share a period and a method.
Hours reclaimed per weekhours/weekoutcomeTime no longer spent on the named task, measured the same way before and after. Never reconstructed afterwards.Median and range, never an average. A mean of it describes none of them.
Correction burdenoutcomeHow much rework AI output needs before it is usable: none, light, moderate, substantial.Median only. A rank is not a quantity.
ConfidencepointsbehaviourSelf-rated confidence using AI safely in their own work, 1 to 5.Median only. A rank is not a quantity.
Organisations with an active AI policycapabilityAn organisation whose policy has an owner, a date and a status of active.Yes, where the cases share a period and a method.
Who owns what was builtimpactThe ownership model recorded on a delivered project. Reported as a distribution rather than a total.No. A kind has no middle value — only how many fell into each.

Why the third column exists

Whether several cases belong in one figure and what may be done with the numbers once they are there are different questions. Ownership models belong together and still have no average. A kind. There is no middle value of a kind — only how many fell into each.

Second

What else could explain it

The question underneath every impact claim, and the one a before-and-after cannot answer on its own. Seven alternatives, each with what it would take to set it aside — and the honest answer is almost always that we contributed.

The people who came

Whoever signs up for an AI workshop is already the person most likely to try one. Their improvement is partly who they were before they arrived.

Set aside by: A comparison group of similar people who did not attend.

It was happening anyway

AI use rose across every sector during this period. Some of the change would have arrived without us.

Set aside by: A comparison group, or a long enough baseline to see the slope before the work started.

The organisation was already moving

Organisations that seek out training are usually already changing. The training joins a direction of travel it did not start.

Set aside by: A baseline taken before contact, measured the same way.

Something else changed too

A hire, a grant, a new system, a departure. Any of these moves the same numbers.

Set aside by: Asking what else changed, and recording the answer — including when the answer is nothing.

They started paying attention

Measuring a thing changes it. Some of the improvement is the counting, not the work.

Set aside by: A measure that exists independently of the programme — system timestamps rather than a log somebody started keeping for us.

It was unusually bad when they called

People seek help at the worst point. Unusually bad periods are followed by ordinary ones whatever anybody does.

Set aside by: More than two points in time, so an ordinary fluctuation is distinguishable from a change.

They are telling us

We taught them, we asked them, and they like us. Self-report to the people who ran the session is warm by construction.

Set aside by: A measure that does not pass through us, or somebody else doing the asking.

The first one is the one we cannot fix

Whoever signs up for an AI workshop is already the person in that organisation most likely to try one. A before-and-after on that group measures the session and the self-selection together, and no amount of care afterwards pulls them apart. It is not a criticism of the workshop. It is a fact about what a voluntary group can demonstrate about itself, and it caps what we will say.

Third

The sentence the evidence permits

The failure this prevents is not lying. It is the ordinary slide from one participant reported reclaiming about five hours a week to Heirloom Praxis saves customers five hours a week — which happens over three drafts, by people who each made a small edit, and nobody involved decided to overstate anything.
anecdotalOne participant reported… / A tool modelled…
case specificIn one documented case…
observed patternAcross three implementations we measured…
aggregate resultAcross N cases, the median was…
validated program resultParticipants across N deliveries showed…
externally verifiedAn independent review found…
What we cannot say

Two levels nothing here reaches

Both are kept in the framework rather than removed. A ladder with its top rung missing looks complete; a ladder with an empty top rung shows you where you are.

Sole cause

Nothing else explains it. Unreachable without knowing what would have happened otherwise, which means a comparison group. Kept here so the top of the ladder is visible and empty.

No path in the assessment returns it. Establishing that nothing else explains a change means knowing what would have happened otherwise, and that means a comparison group we do not have.

Independently verified

Checked by somebody with no stake in the answer. Nothing here currently meets this, and saying so is the point of having the level.

Nothing here has been reviewed by anybody outside this organisation. The level exists so that the gap between what we can say and what would be strongest stays visible.

When somebody’s story appears here, it appears the way they agreed to it — one of three ways, chosen by them:

named
The organisation or person is named.
role only
Described by role or sector, without a name. "A community land trust in Connecticut."
unnamed
No name, no role, nothing that points back. The hardest to actually achieve.

Removing a name is the easy half: a farm co-op in New Haven with eleven members is one organisation to anybody local. Consent is per channel, and withdrawing it means the piece comes down rather than stops being used next time.

How we handle data more generally