In 1954 Paul Meehl compared trained clinicians making judgements against simple statistical rules doing the same job, and found the rules won.1 The finding has been replicated for seventy years. It has also been resisted for seventy years, and the resistance is the interesting part: it is not that the professionals doubt the evidence. It is that being outperformed by a short list feels like an insult, and the list is free, which somehow makes it worse.
The same result turns up in surgery, where a checklist that any nurse can read outperforms the memory of people who have done the operation four hundred times.2 Nobody involved got worse at surgery. They just stopped relying on remembering.
Expect to feel some of that resistance reading what follows. The questions are simple and you know all the answers already, which is exactly what makes them worth writing down.
Eight questions
Answer each honestly. There is no scoring page, no email box, and nothing to submit. If you want to score it, give yourself a point for each clear yes and read the section after.
- Can you say, in one sentence, what finishes?Not “improve operations.” A thing that either happened or did not: the invoice reached the customer, the reminder went, the stock figure updated. If you cannot name the finished thing, nobody can build it and nobody can tell you later whether it worked.
- Does the information already exist as data, or only on paper and in heads? Software cannot read a notebook or a memory. If the answer is paper, the first project is not automation, it is writing things down, and that is a smaller and cheaper project that somebody should tell you about.
- Is there one place that is right when two disagree? If the sheet and the billing system disagree about yesterday, is there an agreed answer to which one wins? Where there is no agreed answer, automation does not resolve the disagreement. It industrialises it.
- Will you switch the old way off? On a date. In writing. If not, you are running a demonstration, and it will quietly end in week three.
- Is there one person whose week gets worse if it stops? Not you, and not the vendor. If nobody feels it break, nobody will notice it break.
- Do you know what one outcome is worth in rupees? One recovered booking, one avoided error, one hour returned. Every decision after this is a comparison against that number, and without it you will be comparing against a feeling.
- Does it still work if the price of the tool doubles? Rate cards move, and the one under most of these builds moved twice in eighteen months. A flow that only works at today’s price belongs to somebody else’s pricing committee.
- Have you run the four numbers? Fully loaded monthly cost of the person, the fraction of the role it really takes, monthly running cost, build cost. If the payback is past about a year, the honest answer is often a person with a better tool.
Notice that seven of the eight are about your business and one is about the software. That ratio is the finding, and it is why buying a different tool so rarely changes the outcome.
Reading your own answers
This part is judgement, not a scoring key, and I am not going to dress it up as one.
- Question 2 answered “paper” means stop. Nothing further on the list matters yet. The project is to get the information into a form that exists, and it costs a fraction of what you were about to spend.
- Question 4 answered “no” means stop. Everything else can be perfect and it will still end in week three.
- Questions 6 and 8 unanswered means you are not ready to hear a quote, because you have nothing to compare it against.
- Mostly clear yeses means the risk in your project is ordinary technical risk, which is the kind worth paying somebody to take.
What would make this an instrument, and why it isn’t one yet
There is a version of this paper that says “businesses scoring below five never shipped.” It would be a better paper to read and a much better thing to put in a sales deck. We are not writing it, because we have not done the work that would make it true.
Doing that work means, specifically:
- Scoring a set of finished projects on these eight questions as they stood at the start, from records made at the time, not from memory after the fact. Scoring from memory guarantees the answer.
- Recording each project’s real outcome on a definition fixed before scoring — still running at twelve months, say.
- Having the scoring done by somebody who does not know the outcomes.
- Reporting the correlation whatever it is, including the case where these eight questions turn out to predict nothing, and reporting how many projects were in the set so you can judge the weight of it.
Until that is published, treat the list as a considered checklist from people who have watched this go wrong, which is a useful thing but is not the same as a tested one. When we do publish it, hold us to the four steps above — they are on this page so that you can.
How this paper was made
The eight questions are derived from mechanism, not from data. Each one corresponds to a failure we can explain: the handoff arithmetic in Paper 1, the discovery split in Paper 2, the cost-per-outcome sum in Paper 3, the threshold rule in Paper 4, the workflow-not-tool argument in Paper 5, the comparability problem in Paper 6, and the local wage maths in Paper 7.
What we have NOT done is score a set of finished projects against their real outcomes and report the correlation. That is the study that would turn this checklist into an instrument, and we have not run it. The card above says so, and the last section of this paper says exactly what running it would require, so that when we publish a scored version you can check whether we actually did it.
Nothing here describes a client. No project, named or anonymised, is characterised in this paper.
On the date at the top of this page. This paper is dated 10 August 2026 because that is its slot in the series. The writing and the working were done on 26 August 2026, when the series was compiled and released together. We would rather say that here than have you find it in the page history.
References
- Meehl, P. E. (1954). Clinical versus Statistical Prediction: A Theoretical Analysis and a Review of the Evidence. University of Minnesota Press. The original demonstration that a simple formula outperforms trained judgement, and that the professions resist the finding.↩
- Gawande, A. (2009). The Checklist Manifesto: How to Get Things Right. Metropolitan Books. The same finding in an operating theatre. Expertise is not the problem; remembering under pressure is.↩