The most common way an AI project fails in a small business is not technical. It is that somebody bought licences for the whole team, sent an email announcing them and nothing measurable happened for six months.
What follows is the opposite approach. One task, a handful of people, a number at the start and the same number at the end. It takes about ninety days and it works whether you are five people or fifty.
Weeks 1 and 2: choose the task
Do not start with the most valuable task. Start with the most annoying one, because the people doing it will forgive an imperfect first attempt.
Ask every team the same question: what do you do every week that you resent? Then filter the answers against four tests. A good first task is one that recurs at least weekly, takes more than an hour each time, produces text or pulls data out of documents, and carries no legal or financial consequence if the first draft is wrong.
Tasks that consistently qualify: writing up meeting notes and actions, producing first-draft quotes from a spec, triaging a shared inbox, summarising tender documents, turning supplier invoices into spreadsheet rows, drafting job adverts and interview questions.
Tasks that do not: anything touching payroll, anything that forms part of a contract, anything where being wrong once damages a client relationship permanently.
Measure the baseline before you touch anything
This is the step everyone skips and it is the reason so many AI projects end in an argument about whether they worked. For two weeks, have the people doing the task record how long it takes. Not an estimate. An actual figure.
You will be surprised twice: once by how long it really takes, and once at the end when you have something concrete to compare against.
Weeks 3 to 6: run it small
Three or four people. Proper business licences, never personal accounts, because the data terms are different and because you cannot audit what you cannot see. Whether that is Microsoft 365 Copilot inside the tools they already use, or a general assistant on a business plan, depends on where the task lives.
Before anyone starts, agree three things in writing:
- What may be pasted in. The green, amber and red categories from your AI policy.
- Who checks the output. A named person owns anything that leaves the business.
- How long the old process runs alongside. Two weeks is usually right. Long enough to catch problems, short enough that people commit.
Then leave them alone and meet once a week for fifteen minutes. The single most useful question in that meeting is not is it working. It is what did you have to fix this week, because the pattern in the corrections tells you whether the task was right.
If the same correction comes up three weeks running, that is not a prompting problem. That is the boundary of what the tool can do for this task, and the process needs to be redrawn around it.
Weeks 7 to 10: measure honestly
Put the baseline next to the new figure and be ruthless about what it tells you.
- Under twenty per cent saving. The task was wrong, not the tool. Do not conclude that AI does not work for your business. Pick another task and run it again.
- Twenty to fifty per cent. Real but fragile. Worth keeping if the people doing it want to. Not yet worth rolling out.
- Over fifty per cent. You have found something. Move to week eleven and build it properly.
Ask the third question too, the one that is not a number: do the people in the pilot want to carry on? If the answer is no despite a good time saving, something about the output quality or the checking burden is not showing up in your figures, and it will kill the rollout later.
Weeks 11 and 12: make it permanent
A pilot that works and is never written down becomes a story about the time you tried AI. Three things turn it into a process.
Write it down
One page. The tool, the steps, the prompt or template if there is one, what to check before sending and what to do when the output is obviously wrong. Store it where the rest of your process documentation lives, not in the head of the person who ran the pilot.
Train the people who were not in it
Half an hour, hands on, doing the real task with their own real work. Not a demonstration. People learn this by doing it badly once with somebody sitting next to them.
Decide what happens when it breaks
Who notices if the tool changes and the output quality drops? What is the fallback if the service is down on a day the task must happen? These are dull questions and they are the difference between a process and a dependency.
What to do in month four
Pick the next task. Not five next tasks. One. The compounding comes from a business that has genuinely absorbed six processes over eighteen months, not from one that announced twenty and completed none.
Somewhere around the third or fourth task you will hit the real ceiling, which is almost never the AI. It is that the data lives in four systems that do not talk to each other, and the person is the integration. That is the point at which the conversation moves from AI tools to business automation and system integration, and where the larger savings actually are.
If you want somebody to walk your processes with you and point at the right first task, that is what our IT consultancy service is for. Tell us what your team resents most and we will tell you honestly whether it is a good candidate.








