A field guide for operators: where it pays back inside a services business, how to run a pilot that survives contact with clients, and what to leave alone for now.
Most AI programs inside operating companies begin with a tool and go looking for a job. The ones that pay back begin with a line on the income statement and go looking for the hours behind it.
In a services business the candidates are narrow and familiar: research and first drafts, quality review, project administration, sales qualification, reporting, and the internal questions that consume a senior person's afternoon. Each is measurable in hours per month before anyone buys anything.
If you cannot name the hours you expect to recover, you are not running a pilot. You are running an experiment on your clients.
Adoption is close to universal. Earnings impact is not. The distance between these bars is the entire subject of this paper.
First four bars: McKinsey, The State of AI: Global Survey 2025, 1,993 respondents across 105 nations; nearly two-thirds of organisations have not begun scaling AI across the enterprise.4 Last bar: MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, which found roughly 95% of enterprise generative AI pilots produced no measurable profit-and-loss return against an estimated $30–40 billion of spend.1
“Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact.”
MIT Project NANDA · The GenAI Divide · July 20251This paper is written for the person who has to answer for the result — a managing director, a head of delivery, an owner. It assumes no enthusiasm and no skepticism, only a budget and a quarter.
Before any pilot, three questions. A no to any of them is a no.
The third question is the one most programs skip, and it is the one that decides whether the saving reaches the income statement. An agency that recovers four hundred hours and has no demand to fill them has bought a morale problem. The same four hundred hours in a business with a waiting list is margin. Ask the commercial question before the technical one.
Three questions, each answerable in a week, each fatal on its own.
| Question | What a no means | How to test it in a week | Evidence to keep |
|---|---|---|---|
| Is the output checkable? | You are shipping unverifiable work and calling it delegation | Have a reviewer grade twenty existing outputs and time each check | Minutes per check, and the disagreement rate between two reviewers |
| Is the input already ours? | This is a data and contracts project wearing a pilot's costume | Name the systems the work draws on and who holds the rights | A one-page data map with a legal position on each source |
| Does the hour saved go somewhere? | You are buying idle capacity, not margin | Ask what the person would do with a recovered day, and whether it is billable or sold | Pipeline or backlog that the recovered hours will absorb |
Our own filter. It exists because the failure MIT identifies is not model quality but the gap between a tool and the workflow it was bought to change.1
Ranked by how reliably we have seen the return arrive, most reliable first.
Start where the hours are concentrated and the output can be checked in minutes. Everything else can wait a year.
Our own placement, from pilots across the group. Open points are workflows where we have either failed to bank the saving (internal question answering, where the recovered time is spread across many people) or decided not to try (final client judgment). MIT found the same asymmetry from the other direction: more than half of enterprise budgets went to sales and marketing, while the highest measured returns came from back-office work.1
Research and first drafts. The work between a brief and something a senior person can react to. High volume, checkable in minutes, and the reviewer was already reviewing. This is where nearly every services business should start.
Internal question answering. The afternoon a senior person loses to questions whose answers already exist somewhere in the company. The return is real but it is spread thinly across many people, which makes it easy to claim and hard to bank.
Project and account administration. Status notes, timelines, meeting records, the reporting nobody bills for. Unglamorous, safely checkable, and consistently underestimated.
Sales qualification. Sorting and preparing for inbound rather than deciding on it. Pays back where volume is high; a waste of effort where the pipeline is twenty relationships.
Quality review. As a first pass that catches the mechanical, never as the last pass. Useful, and the easiest place to talk yourself into removing the human signature.
Reporting. Assembly is a good candidate; interpretation is not. Draw that line explicitly, because it will move on its own if you do not.
Ranked by how reliably we have seen the return arrive, with our own confidence stated.
| Workflow | Where the hours sit | Checkable | Our evidence | Verdict |
|---|---|---|---|---|
| Research and first drafts | Delivery, junior and mid | In minutes | Measured more than once | Start here |
| Project and account admin | Delivery and PM | In minutes | Measured more than once | Start here |
| Internal question answering | Spread across seniors | Sometimes | Real but unbanked | Run it, expect no line item |
| Sales qualification | Sales, volume-dependent | Sometimes | One pilot, high volume only | Only above real volume |
| Quality review, first pass | Senior review | In minutes | Useful, never final | First pass only |
| Reporting | Finance and accounts | Assembly yes, reading no | Assembly measured | Draw the line explicitly |
| Final client judgment | Partner and principal | No | Not attempted | Leave alone |
Our own ranking across six operating companies, measured against a common baseline. “Our evidence” is deliberately unflattering where it should be: two of these rows rest on a single pilot.
Ninety days, one team, one workflow. Measure the baseline for two weeks before anything changes. Name the person accountable for the result, not for the tool.
Keep a human signature on client work. Someone with a name and a job title reviews and owns every deliverable that leaves the building. This is not a compliance formality; it is the reason the client keeps paying.
Write the disclosure before the first client sees the work, not after they ask. Clients tolerate a great deal and forgive very little.
Kill it on schedule. A pilot with no end date becomes a permanent cost with no owner.
One team, one workflow, a measured before, and a date on which someone decides.
Our own pilot shape. The two weeks of baseline before anything changes are the cheapest insurance in the sequence, and the step most often skipped.
The last line deserves emphasis. A workflow that halves drafting time and doubles review time has not saved anything; it has moved the cost from a junior person to a senior one, which is usually a loss. Measure the reviewer or do not bother measuring.
Two pilots that both look like a success on the drafting line. Only one of them is.
Illustrative, on the pattern we have seen twice. Both pilots cut drafting from 100 hours to 40. In Pilot B senior review went from 20 hours to 70, so total effort barely moved and the cost migrated to the most expensive person in the room. Reported as a drafting metric, Pilot B is a triumph for about two months.
We have run enough of these across the group to see the same four failures repeat.
No baseline. The team never measured the before, so the after is a matter of opinion. This is the most common failure and the only one that is free to prevent.
The reviewer absorbs the cost. Output volume goes up, quality holds, and one senior person is now working evenings. The pilot reports success for two months and then the reviewer resigns.
Ownership sits with the enthusiast. The person who championed the tool is accountable for the outcome, so the evaluation is not independent. Separate the two roles even when it feels bureaucratic.
The pilot has no end date. It becomes infrastructure by inertia, appears in no budget, and is discovered a year later by finance.
Four failures, the signal that shows up first, and the month you can catch each one.
| Failure | Early signal | Visible by | The fix, before you start |
|---|---|---|---|
| No baseline | The debrief argues about whether it worked | week 12 | Two weeks of measurement before anything changes |
| Reviewer absorbs the cost | Output volume rises, one senior person starts working evenings | week 5–8 | Log reviewer hours per deliverable from day one |
| Owned by the enthusiast | Every reported number is favourable | week 6 | Separate the champion from the person who owns the verdict |
| No end date | The tool appears in no budget and has no owner | a year late | Write the kill date into the pilot brief |
Our own record across the group. The wider evidence points the same way: MIT attributes the 95% to a learning and workflow gap rather than model quality,1 McKinsey finds fundamental workflow redesign to be the strongest correlate of EBIT impact,4 and Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027 on cost, unclear value and weak controls.5
“The single strongest correlation with EBIT impact is fundamental workflow redesign.”
McKinsey · The State of AI: Global Survey 20254“For now” is doing real work in that heading. This list is a position on today's tools and today's contracts, not a principle. Review it annually and be honest about which items are still true and which have become habits.
Disclosure is a commercial decision disguised as an ethical one, and getting it wrong is expensive in both directions. Say nothing and you are one question away from a conversation about trust. Over-announce and you invite a discount request on work that took the same craft it always did.
What has worked for us is narrow and factual: name where the tools are used, name who reviews and owns the output, and confirm what happens to the client's data. Put it in the engagement terms rather than in a press release. Then have the answer ready in one sentence when a client asks in a meeting, because they will.
Two things not to do. Do not present efficiency as a reason to lower the price unless you have decided to compete on price. And do not let a client's enthusiasm move the human signature — the client who asks you to skip review is not the one who will accept the consequences.
Most of the risk in this work is not model behavior. It is data leaving a boundary it was not supposed to leave, under a contract that did not permit it.
The minimum is unglamorous: a named list of approved tools and a stated position on everything else; a check that the vendor does not train on your inputs; identity and access managed centrally so leavers actually leave; and a review of client contracts for subprocessor and confidentiality clauses that predate all of this. Most standard agreements written before 2023 do not contemplate any of it.
None of this requires a security function. It requires one person spending a week and writing down the answers.
One week of work, seven answers, written down where someone else can read them.
| Item | The question to answer | Evidence it is done |
|---|---|---|
| Approved tool list | Which tools are permitted, and what is the position on everything else? | A named list with a date and an owner |
| Training on inputs | Does the vendor train on our data or our clients' data? | The clause, quoted, per vendor |
| Identity and access | Are accounts issued centrally, so leavers actually leave? | A single sign-on record and an offboarding test |
| Client contracts | Do subprocessor and confidentiality clauses permit this use? | A reviewed list of exceptions to renegotiate |
| Disclosure language | What do we tell clients, and where does it live? | Two sentences in the engagement terms |
| Retention | Where does the output live, and for how long? | A retention setting, not an intention |
| Named human signature | Who reviews and owns each client deliverable? | A role in the workflow, by title |
Our own checklist, run once for the group rather than six times. Most standard client agreements drafted before 2023 do not contemplate any of this.
A single company runs one pilot and learns one thing. A group of companies runs six and can compare them — same measurement, same review standard, different contexts. It can also buy once, negotiate once, and set one security posture rather than six.
The collective's job is not to choose the tools. It is to make sure that when one company learns something expensive, the other five get it free.
The second-order advantage is more useful. Six companies produce six baselines, which means a claim about hours recovered can be checked against a comparable rather than against a vendor's case study. A group that measures the same way in every company knows, within a quarter, which claims are real.
The published record of enterprise spending is a map of where not to start.
MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, based on more than 300 public deployments, 52 structured executive interviews and 153 leader surveys.1 The 95% headline counts P&L impact within roughly six months of deployment; the fair criticism of it attacks that precision rather than the direction.3
We are three years into this and the honest position is that our confidence is uneven. We are confident about first drafts and administration, where we have measured the return more than once. We are not confident about internal question answering, where the saving is real and almost impossible to bank. We have no evidence at all on whether any of this changes what a client is willing to pay, in either direction.
We also do not know how much of today's guidance survives the next capability step. The three-question filter should; the list of things to leave alone almost certainly will not.
The operators who do well with this are not the enthusiasts. They are the ones who treat it as ordinary capital allocation: a defined cost, a measured return, a date by which the answer is known, and a willingness to stop.
If you are running a services business and want to compare notes on what has and has not worked, we are glad to. We will tell you about the pilots we stopped as well as the ones we kept.
Third-party research in this paper describes surveyed and published deployments, not our companies. Everything labelled as ours — the ranking in Exhibit 04, the pilot shape, the checklist, the failures — comes from pilots run inside the group against a common measurement standard, and we have marked the places where it rests on a single data point.
Figures are quoted as published, with their sample sizes where they are small. If a number here is wrong, write to us and we will correct it in the next revision.