
AI Models Broke Into Four Real Companies During Safety Testing. Nobody Audits the Test Lab.
2 August 2026
Saying ‘Lean’ Is Not a Skill
26 August 2026In June, a customer operations director at a major top tier retailer told me she was eighteen months into her AI rollout and still couldn’t say if it was actually working.
It wasn’t a lack of data. Her dashboard was full of promising containment rates, declining average handling times, and a complaints chart quietly creeping in the wrong direction. What she lacked was a single defensible metric she could present to her CEO that would still hold true in six months.
She followed that admission with a hesitant question: is everyone else further ahead?
Short answer: no.
On 7 July, Octopus Energy published trial results for Arlo, its AI email assistant. Arlo achieved a 76% customer satisfaction score, nudging past the 72% scored by human advisors on comparable responses. The media coverage was warm, and the numbers looked great.
But the metric that matters most was tucked into the third bullet of the press release. Arlo handled roughly 8,000 emails a week. That represents 4% of Octopus’s total UK email volume.
Octopus built its own tech stack in Kraken, spun it out as a separate company, and is widely regarded as the sharpest consumer technology operator in British energy. Their production AI boundary sits squarely at four per cent.
Over in retail banking, NatWest operates across 19 million customers and reported well over 200 AI projects in flight as of early 2025. By the end of Q1 2026, its agentic financial assistant within Cora was accessible to 25,000 users. That is about one in every 760 customers.
Neither project is a failure. But treating them as evidence of an imminent full-scale automation takeover misses what is actually happening on the ground.
Consumer resistance isn’t ideological, it’s pragmatic
The strongest argument for AI customer operations isn’t found in vendor slide decks or in Gartner’s March 2025 projection that agentic AI will resolve 80% of common issues by 2029. It’s found in how consumers behave when a problem is solved cleanly.
In controlled conditions, Octopus’s system outperformed the human baseline by four points on comparable queries. That isn’t vendor marketing. That’s a major utility publishing comparative trial data it wasn’t required to share, at a company whose service reputation is its main marketing asset.
A March 2026 global study by Ada and NewtonX, covering 2,000 consumers and 500 CX decision-makers, found that 59% of consumers prefer instant, round-the-clock AI support over waiting for a human agent. Ada sell AI agents, so weigh the framing accordingly, but the caveat attached to that number is the whole finding: the preference holds only when the issue actually gets resolved. In the same study, just 24% of respondents said their last AI service interaction was fully resolved without a human.
Customers aren’t defending the sanctity of human interaction. They’re defending their afternoon. Solve the issue instantly and the objection largely evaporates. Fail to resolve it, and the AI becomes an irritating barrier between the customer and a solution.
Why aggregate dashboards lie
On the other side, Qualtrics found that only 5% of UK consumers actively prefer dealing with automated assistants, a third don’t trust the information those assistants give them, and 58% fear being trapped in an automated loop with no human escape hatch. That last figure is a worry about architecture, not about intelligence.
AnswerConnect, which sells human answering services and therefore has a dog in this fight, tracked UK, US and Canadian consumers between October 2025 and April 2026. On their numbers, frustration with AI agents rose from 54% to 59%, preference for a real person from 83% to 85%, and the proportion who would hang up on reaching a bot from 29% to 31%. Discount for the incentive and the direction of travel still doesn’t favour anyone selling full automation.
Then there is Klarna, which I include reluctantly because it has been chewed over so thoroughly it has stopped working as evidence and started working as a fable. The useful detail isn’t that its OpenAI-built assistant handled 2.3 million chats in its first month, equivalent to the workload of roughly 700 agents. It’s that in May 2025 Sebastian Siemiatkowski told Bloomberg the company had cut human support too aggressively, and reopened hiring for complex and premium cases. Whole-population CSAT held up. The long tail didn’t. If your reporting only tracks the average, you’ll find out about the tail from the Ombudsman.
Which is why Gartner followed its optimistic 2029 projection with a rather different note in June 2025, predicting that over 40% of agentic AI initiatives will be abandoned by the end of 2027 on grounds of spiralling costs, unclear business value and inadequate risk controls. The same analysis coined “agent washing” for rebadged chatbots and RPA, and reckoned that of the thousands of vendors claiming agentic capability, roughly 130 were real.
Meanwhile the national baseline sits still. The July 2026 UK Customer Satisfaction Index came in at 78.3 out of 100, up a single point on a year ago and 0.1 points on January. Three years of unprecedented capital expenditure in support technology, and the barometer has moved by a tenth of a point in six months.
Software compiles. Policy doesn’t.
Why does AI progress feel so effortless in software engineering and so treacherous in customer service?
In software development, a coding assistant works alongside a developer. They share the same objective, the developer stays in the loop, and broken code announces itself. If the agent gets it wrong, there is a reset button.
A customer service agent isn’t an assistant. It is a representative. It speaks on the company’s behalf, makes commitments and sets expectations, and there is no undo on what it has already told someone. The party on the other side of the screen isn’t a collaborator testing code. They want a full refund where policy permits a partial credit, an exception the agent can’t grant, an answer the agent isn’t sure of. None of that is hostile. It’s just Tuesday.
Code either compiles or it fails. Policy documents are written in human language. They cover typical scenarios and leave the edges to judgment, and the edges are where real customer service lives.
A late £12 impulse purchase and a delayed £400 birthday present look identical to a tracking API and represent entirely different human conversations. A first-time buyer with a missing parcel needs a different touch from a ten-year account holder suffering their third delivery failure this month. No knowledge base in Britain captures that distinction on its own.
The bottleneck isn’t the underlying model. Everyone has access to the same models. It’s the governance layer: the explicit rules defining what the agent handles, what it must never promise, when it stops, and what it hands over with.
The CX vendor Quiq put this well in a recent ebook, and the diagnosis is sharper than the pitch surrounding it. Stop programming conversations, they argue, and start governing them. Most organisations have done neither. They have a large prompt, some glue code, and one engineer who is quietly a single point of failure.
Four per cent isn’t a ceiling on what Arlo could handle. It’s the size of the space Octopus was prepared to declare in writing.
The regulatory clock started running on 2 August
Article 50 of the EU AI Act became enforceable on 2 August 2026.
This has been widely missed, and the reason is instructive. The Digital Omnibus extended compliance deadlines for high-risk systems under Annex III to December 2027 and Annex I to August 2028. Headlines said the AI Act had been delayed. Boards heard “delayed” and moved the item down the agenda. Article 50 was deliberately excluded from that deferral.
If you deploy customer-facing AI in the EU, regardless of where your head office sits, you are now required to tell the person they are dealing with a machine at or before the first interaction, in the interface rather than in a terms-of-service link nobody opens. A separate obligation to embed machine-readable markers in AI-generated content bites on 2 December for systems that were already live in August. Penalties run to €15 million or 3% of global annual turnover.
Once you explicitly inform a customer they are speaking with a machine, two questions follow immediately. What authority does this machine actually have? And where is the human escalation button?
You cannot answer either from inside a sprawling 4,000-word prompt. They require documented, reviewable bounds that a compliance officer can read.
UK operators face similar pressure by a slower road under the FCA’s Consumer Duty and its vulnerable customer expectations, without a date attached to concentrate the mind. Look closely at how Octopus structured Arlo. Vulnerable customers, sensitive cases and complex complaints were routed around the AI entirely, and every generated email was labelled as AI-written. That wasn’t a last-minute compliance patch. It was a deliberate design boundary, drawn months before Brussels required one.
Three actions for Chief AI Officers this week
Check that every customer-facing AI deployment identifies itself at the start of an interaction, in the interface itself. If you serve EU customers, that stopped being a design preference on 2 August.
Then write down what your agent is forbidden to say. Every organisation I walk into has a can-do list, because the can-do list is what got the business case approved. Almost nobody has the can’t-say list: the refund threshold it may not exceed, the discounts it has no authority to offer, the account changes it may not make, the sentiment triggers that force an immediate handover regardless of how confident the model sounds. That document is unglamorous, it will cost you a fortnight of arguments with legal and operations, and it is the only artefact in your entire AI programme that will survive contact with an angry customer, a regulator, or the Financial Ombudsman. Start it with vulnerable customers, because that is where the exposure is worst and the internal guidance thinnest.
And stop treating containment as resolution. Containment measures whether the customer went away, which is not the same as whether they were helped. Segment your resolution reporting and put the complex tail on its own chart, because that is the chart your board has never been shown.
Octopus’s 4% isn’t a sign of stalled progress. It is the exact proportion of customer interactions they were willing to guarantee, in writing, under governance.
The companies that look slow right now are generally the ones that defined their agent’s boundaries before exposing them to live customers. The companies that look fast are generally the ones nobody has audited yet. That asymmetry has been comfortable for about two years.
Octopus put 4% in writing. Most of your competitors have never been asked what their number is.
That question arrives on 2 December.
Andy McGurk is a Fractional Chief AI Officer and founder of AMVEN, advising UK and European organisations on AI governance, operational design and continuous improvement. He is based in Peterborough and recently needed three separate chatbots to change one direct debit.
Sources
- Octopus Energy, “Octopus Energy’s AI trial wins customer approval”, 7 July 2026 — https://octopus.energy/press/more-news-press-releases/octopus-energy-s-ai-trial-wins-customer-approval/
- Institute of Customer Service, UK Customer Satisfaction Index, July 2026 — https://www.instituteofcustomerservice.com/research-insight/ukcsi/
- Gartner, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027”, 25 June 2025 — https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- Ada and NewtonX consumer and enterprise research, March 2026 — https://www.businesswire.com/news/home/20260324961586/en
- AnswerConnect, 2026 AI Customer Service Attitudes Report, May 2026 — https://www.answerconnect.com/blog/news/consumers-turning-away-from-ai-customer-service/
- Qualtrics UK Consumer Experience Trends research, via Customer Experience Magazine — https://cxm.world/customer-experience/qualtrics-uk-consumer-research-highlights-low-trust-in-ai/
- Gibson Dunn, “EU AI Act Omnibus Agreement: Postponed High-Risk Deadlines and Other Key Changes” — https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/
- EU Artificial Intelligence Act, “The EU AI Act’s Transparency Rules: A Practical Guide to Article 50” — https://artificialintelligenceact.eu/transparency-rules-article-50/
- NatWest Group, technology, data and AI — https://www.natwestgroup.com/who-we-are/about-natwest-group/tech-data-and-AI.html


