Ask a bank where its biggest AI risk sits, and the answer is almost always governance. Few industries have approached AI as carefully as banking. According to Personetics, 80% of global banking executives see gen AI as an opportunity, but only 18% report it's fully integrated into day-to-day operations, with the accuracy, reliability and compliance of the output the biggest barrier to adoption.
With regulators watching and customer trust on the line, banks are spending real money to get AI customer communications right. In Sinch's AI Production Paradox study, a survey of more than 2,500 enterprises, including over 500 financial leaders, “trust, security and compliance” ranks as the number one AI investment priority across financial services, ahead of AI development itself. And it’s not only money: Most AI engineering teams now spend at least half their time building safety controls rather than shipping new features.
Still, the agents keep getting rolled back, and in big numbers. The bigger worry, though, is the bank that's never found a reason to.
Every bank’s AI fails, but not every bank sees it
Roughly two in three (69%) banking AI customer-facing agents have already been rolled back after a governance failure, and that rate doesn't get better as programs mature. In fact, the most governed financial organizations in the study report the highest rollback rate of all: 80%.
That seems counterintuitive, but it isn't, once you separate two things: how often an agent fails, and how often a bank catches it. The most mature programs are rolling back more because they're seeing more. Failures are inevitable in a non-deterministic system, and no amount of governance spending drives them to zero. So, a rollback is a sign of detection: It means the failure got caught, which is the system working. That’s why the banks that should worry most are the ones reporting none. A zero rollback rate is an indication of weak observability, not real success.
So, why do so many failures go unseen? Because the tools watching for them were built for a different job. Standard monitoring (latency, error rates, uptime) tells you whether a system is running, not whether its answers are right. An agent that surfaces the wrong account balance would log a clean interaction, and only the customer would know it failed. Sinch’s study found 16% of rollbacks can’t be fully diagnosed because there’s no audit trail. That means for many banks, the cause of failure is a total mystery.
In banking, every failure is a trust failure
Globally, the biggest consequence of AI failures is an impact on the support queue. But in banking, the impact costs more. Customers hand a bank their money because they trust it to get things right. A wrong balance, an account detail shown to the wrong person or a fraud alert that never fires or lands too late is a service failure and a breach of the relationship at the same time. And often, it becomes a compliance event on top.
“For critical communications like OTP and high-risk notifications, we must send messages in a timely manner, often within seconds. If we fail to meet these timelines, we are required to report it to regulators, which comes with a risk of penalty. Most importantly to HSBC, any delay can impact customer experiences, and their day-to-day lives.” - Phoenix HY Li, Head of Messaging, HSBC.
When an agent floods the support queue, service bounces back once it's online again. Reputational damage doesn't, and for a bank it rarely stops at reputation. One visible failure can invite regulatory scrutiny and hand customers who are already a tap away from switching a reason to leave. No surprise, then, that banks name reputational damage as the consequence they fear most (39%), one of the few sectors where brand risk outranks the hit to the support queue. Whatever breaks behind the scenes, the customer sees one thing: their bank getting it wrong.
Not every rollback is a model problem
So, what's setting off these rollbacks? Sinch’s study found customer data exposure is behind more than a quarter of them (27%), the leading cause for financial institutions. Hallucinations and off-brand responses rank second at 21%.
Telling the two failure modes apart is key to understanding the fix. When a customer's account details surface where they shouldn't, that's an infrastructure failure, not a prompting one. An infrastructure failure you can design out: Mask the data before it ever reaches the model, and it can't leak, because it was never there. A hallucination or an off-brand answer is a model failure. You can't eliminate that, only contain it, by narrowing what the agent is allowed to do and checking its work. The fix needs to start with an accurate diagnosis of the root cause.
The biggest cause of rollbacks, data exposure, is an infrastructure failure. So are others the study surfaced, like lost context between channels or reliability that fails at scale. But the communications infrastructure decides more than individual failures. Sinch found it's the strongest predictor of AI success, ahead of governance maturity, model choice or budget.
Most banks agree in principle, but few act on it: 83% call high-performance communications infrastructure essential to deploying AI safely, yet only 56% fund it as a priority.
Where guardrails belong
None of this is an argument against guardrails. Safety controls are critical in AI customer communications, even more so in a regulated industry like financial services. So is strong monitoring. Guardrails stop the failures you can anticipate, while observability catches the ones you can't.
But with 84% of engineering teams spending at least half their time building safety controls, banks must evaluate where each guardrail belongs: built on top of the platform as part of a defense-in-depth strategy, or provided by the communications infrastructure the agents run on.
“Every team needs to decide what controls belong at the platform layer and what their engineers should build on top, because the cost of building custom guardrails compounds over time, especially as the team moves through the product lifecycle. Each new agent, each new channel, each new deployment adds to the pile. And eventually you lose that momentum when it comes to outperforming on the market.” - Anton Efimenko, SVP of Software Engineering at Sinch.
Banks are spending more on trust than ever, and it still isn't keeping their agents live. That won't change until the money reaches the platform the agents run on, where many of these failures start. For now, most of it is going into more guardrails, not better infrastructure.
Sinch's AI Production Paradox report digs into what keeps enterprise AI agents in production and what pulls them back. Explore the findings.
Alejandro Murcia is the Global Director of Financial Services at Sinch, the communications infrastructure provider banks and financial institutions trust to power intelligent customer communications at scale.