The common pattern is: an impressive demonstration, enthusiasm, a pilot, and then nothing. Six months later the tool is used occasionally by two people and nobody quite likes to say it did not work.

Avoiding that is mostly about sequence and honesty.

Stage one: find the tasks, not the technology

Start from where time goes rather than from what a product can do.

Spend a week noting the tasks in your business that are: high volume, repetitive, involve reading or writing text, and currently done by capable people who could be doing something better.

Then filter for consequence. The right first candidates are ones where being wrong is visible and cheap — a draft that gets edited, a summary somebody reads, a classification that gets checked.

The wrong first candidates are anything where an error reaches a customer unreviewed, anything with a regulatory dimension, and anything where you could not tell whether the output was wrong.

Typical good starting points: drafting routine correspondence, summarising long documents, categorising incoming enquiries, extracting structured information from unstructured documents, and first-draft content that a person then edits.

Stage two: establish the baseline

The step that separates projects that succeed from pilots that evaporate.

Before introducing anything, measure the current process: how long the task takes, how often it is done, how often it goes wrong, and what a good output looks like.

Without this you cannot tell whether the tool helped. And in the absence of measurement, judgement defaults to impressiveness, which is not correlated with usefulness.

Also collect twenty or thirty real examples — including the awkward ones. These become your test set.

Stage three: evaluate against the baseline

Run the tool against your real examples, not against a demonstration.

Measure the same things you measured before. Time taken including review and correction. Proportion of outputs usable as-is. Proportion needing minor correction. Proportion unusable. Errors that would have reached a customer if nobody had checked.

That last figure is the important one. A tool producing plausible, confident, wrong output is more dangerous than one that fails obviously, because the failures are the ones people stop checking for.

Be prepared for the answer to be no. A tool that saves ten minutes and requires fifteen minutes of checking has not helped, and recognising that early is a success rather than a failure.

Evaluate on your own awkward examples, not on the demonstration. Impressive output on curated cases predicts almost nothing about performance on your Tuesday afternoon backlog.

Stage four: design the process, including failure

A tool is not a process. Before deployment, decide:

Who reviews the output, and against what standard? "Somebody checks it" is not a process. What are they checking for, and what do they do when it is wrong?

What is never sent without review? Anything reaching a customer, anything with a number in it, anything with legal or financial consequence.

Who is accountable? The person who sends it, always. "The system produced it" is not a position anybody can defend to a customer or a regulator.

What data may be used? Customer information, personal data and commercially sensitive material each need a decision, and that decision depends on the service being used and where it processes data.

How do you know if it degrades? Models and services change. A periodic re-run of your test examples catches it.

Stage five: deploy narrowly, then widen

One team, one task, for a month. Fix what they report before anybody else sees it.

Then widen with those people advocating rather than with management announcing. The adoption dynamics are the same as any system — see getting staff to use the system.

Resist the temptation to apply it to five more tasks before the first is genuinely working. Breadth before depth is the main cause of pilot graveyards.

Governance, proportionate

A short written policy, actually applied, covering: what tools are approved, what data may be entered into them, where human review is mandatory, who is accountable for output, and how new tools get approved.

Keep a register of what is in use. Businesses routinely discover that staff have adopted three tools nobody sanctioned, because the tools are useful and the approval process was unclear.

Under UK GDPR, entering personal data into a service means you need a lawful basis and to know where processing happens. That is a real obligation and it is straightforward to satisfy if considered in advance — covered in AI and UK GDPR.

The wider policy question is covered in an AI policy for a small business.

What to expect realistically

The businesses getting genuine value are mostly doing unglamorous things: drafting, summarising, extraction, classification, and first drafts that people finish.

The returns are real and modest — hours a week rather than transformation. Accumulated across several tasks, that is worthwhile.

What consistently disappoints is anything requiring the output to be right without review, anything depending on facts the tool does not have, and anything where the business could not previously articulate what "good" looked like.

The order, briefly

  1. Find the repetitive, low-consequence tasks.
  2. Measure the current process and collect real examples.
  3. Evaluate against those examples, honestly.
  4. Design the process including review and failure.
  5. Deploy to one team, fix, then widen.
  6. Write the policy and keep the register.
  7. Re-test periodically.

Our consultancy team works with UK businesses on this sequence, including telling you when the answer is that a tool did not help. Start a conversation.