Why does the same AI give different answers to the same task?
Because you never actually gave it the same task. Like any technology: inconsistent input, inconsistent output - regular readers will have seen me put it less politely. Two colleagues ask AI to draft the same client letter; one includes the role, the source documents and the rules, one types a single sentence. Same tool, different universes of output. Then the letters go out, a client notices the tone has changed between Tuesday and Thursday, and AI gets the blame for being "inconsistent".
The research says this runs deeper than effort. A study from the University of Washington and the Allen Institute found that trivial changes to prompt formatting - spacing, punctuation, layout - can swing a model's task accuracy by up to 76 percentage points. Not changes to the question; changes to the formatting of the question. If professional researchers get performance differences that large from details nobody would think matter, what chance does an unguided team have of producing consistent output from freehand prompting? That's the wild west of AI: everyone using it, some well, some badly, and the organisation seeing no consistent return.
Isn't better prompt training the answer?
It's a part of the answer - a smaller part than the training industry wants it to be. I've sat in plenty of AI workshops where a specialist blows the room's socks off with brilliant prompting, and everyone scribbles the prompts down. Good prompting matters and will lift your outputs. But it's one piece of the puzzle, and it has a ceiling: you can never guarantee that everyone will prompt correctly for their role, every time, under deadline pressure, forever. A method that depends on every individual remembering the rules isn't a method - it's a hope.
Notes get lost. People leave. New starts copy whoever sits nearest. Prompt training on its own produces a few individual power users and no organisational consistency - which is exactly the pattern behind the 77% of AI-adopting businesses reporting no change in revenue.
What happens when inconsistency meets agentic AI?
The stakes change completely, and this is the part I'd want every MD to hear. With traditional AI, a weak prompt has a small blast radius: Copilot mangles the spreadsheet tidy-up, you bin the output, you still have the original. With agentic AI - AI that completes work rather than suggesting it - a weak prompt acts on the real world. Point an agent at a folder of files to cleanse without a backup, and a bad instruction doesn't ruin a draft; it loses the files.
Take a VAT return. Traditional AI: one person's prompt includes the rule to pull tax calculations directly from HMRC, another's doesn't - and the second picks up incorrect guidance from an EU source and calculates it wrong. Annoying, and caught in review - hopefully. Agentic AI with the same freehand approach: an incorrect VAT return doesn't sit in a draft - it gets submitted. And the liability for that sits with a person, not the tool. It always does.
Roll out agentic AI without governance, training and proper adoption, and you haven't automated your work. You've automated your errors.
How do you make AI reliable?
Stop relying on prompting and start locking in method. Co-design the workflow with the team who own the work: map the task end to end, bake the rules into the workflow itself - the VAT calculation pulls from HMRC's authoritative source, the letter uses the firm's approved structure, the backup happens before any file is touched - then save it as a trained, repeatable skill and roll it out across the team. The rules stop living in people's heads and start living in the workflow. Right output, every time, whoever runs it.
Approval gates go where the risk is: nothing gets submitted, sent or deleted without a human sign-off at the points that matter. That's how you trust the output without checking everything twice - the checking is designed into the workflow once, not performed by everyone forever.
The evidence sits firmly on this side. McKinsey's State of AI research finds the organisations getting real bottom-line impact from AI are the ones that fundamentally redesigned workflows - not the ones that trained harder on prompts. And it's what the data shows in our own programmes: against a public benchmark of about 2.2 hours saved per week, organisations we work with get 6 to 10 hours back per person per week - because the method is in the workflow, not the individual. One approach speeds a few individuals up. The other compounds efficiency across a team.
Frequently asked questions
Why does AI recreate documents that need an exact format?
Because freehand prompting invites it to improvise. When a document must match an exact format - a statutory return, a client template - the format belongs in a trained skill the AI runs, not in a prompt the human remembers. Locked in once, the output is deterministic.
How do I trust AI output without checking everything myself?
Design the review in, rather than doing it all yourself: approval gates at the points where work leaves the business, spot-checks where risk is low, full human sign-off where it's high. If every hour AI saves is spent re-checking, that's not a trust problem - it's a workflow design problem.
Who is accountable when AI gets it wrong?
A person or a company - always. UK law gives AI no legal personality; responsibility for what it produces sits with the human who deployed it. Which is precisely why reliability has to be engineered, not hoped for. I've written more on data safety and accountability here.
If your team's AI output changes depending on who pressed the button, that's fixable - type out the workflow, and you'll see exactly what locked-in looks like for your team - mapped, free.