Every fintech conference now has the same room: a vendor on stage, a slick demo running, a crowd nodding along at how effortlessly the AI just did something that used to take a team of people. Then everyone goes back to the office and tries to actually build it and that’s where most of it quietly falls apart.
Last year, 42% of companies abandoned most of their AI initiatives. Not stalled, not delayed, walked away from. That’s according to S&P Global Market Intelligence’s 2025 Voice of the Enterprise survey, which polled more than 1,000 IT and business leaders across North America and Europe and found the figure had jumped from 17% in 2024, with the average organisation scrapping close to half of its AI proofs of concept before they ever reached production (S&P Global Market Intelligence, “AI experiences rapid adoption, but with mixed outcomes”). It’s an enterprise-wide figure rather than one specific to fintech, but it tracks closely with what regulated industries like payments experience on the ground.
For a technology that’s supposedly transforming the industry, that’s a striking amount of wreckage. And it’s not really an indictment of AI – it’s closer to a sign of an industry growing up. Gunnar Már Gunnarsson, CTO and co-founder of PAYSTRAX, and Petras Abaravičius, PAYSTRAX’s Head of IT, have watched that failure rate up close, running their own AI vendor evaluations inside a highly regulated business. Their experience is a useful anatomy of exactly how these pilots die — and it looks nothing like a technology problem.
The vendor shortlist that shrank to a handful
Nineteen. That’s how many AI vendors PAYSTRAX shortlisted to explore. By the end of the process, only a handful of strong proofs of concept remained. In an unregulated environment, Petras points out, you can do whatever you please, but in payments, that freedom collapses fast once real requirements show up.
That’s not really a failure of the vendors. It’s a mismatch between what a demo can show and what a regulated production environment actually demands. Gunnar’s read on it is blunt: what’s said to be possible in a sales pitch often turns out to contradict the requirements, or simply produces contradictory results once you dig in.
Demos are human theatre, not integration tests
Maybe the sharpest insight into why so many pilots die is also the simplest: demos are built to be understood by humans, not to be plugged into other systems. Petras frames it directly: demos are done by humans, for humans, a vendor performing for a client in a room. They’re calibrated to convince a person, not to prove that one system can actually talk to another.
The failure mode follows naturally from that: buyers expect the technology to give them the answer they think is correct, but when it comes to integration between systems, that readiness just isn’t there yet. The early wave of AI tooling was built to communicate well with people. It wasn’t necessarily built to integrate cleanly with the tangle of legacy processors, compliance systems, and data pipelines that make up a real payments stack. A demo answers “can this look intelligent?” It doesn’t answer “can this plug into what we already have?”
Regulation isn’t the obstacle – it’s the filter
It’s tempting to hear “regulated industry” and assume that’s simply where AI goes to die. Gunnar pushes back on that framing directly: regulators tend to want people to make the decisions, not the technology, even in places where the technology could clearly help. That’s a real constraint on what a pilot can become. But it works less as a brake and more as a quality filter. It’s exactly why 19 vendors turned into a handful of serious contenders, and exactly why PAYSTRAX invested heavily in frameworks like DORA compliance before, not after, trying to move fast.
Petras calls this “compliance by architecture” and “compliance by design” – building the audit trail and governance model into the infrastructure itself, so any new proof of concept inherits compliance automatically instead of needing it bolted on afterwards. That one design choice, more than any specific model or vendor, is what lets a regulated company move with confidence instead of caution.
The skill nobody priced in: moderating the agent
A lot of the disappointment around AI pilots was never really about the technology, it came from the wrong usage, from missing moderation. Knowing how to direct an AI system, set its context, and catch where it’s blind to dependencies it can’t see is, as Petras puts it, still very much a skill. Treat AI like something you simply switch on and you get noise. Treat it like a system that requires real expertise to operate and you get results.
That’s a big part of why payments companies are now hiring dedicated AI engineers and building internal AI expertise rather than assuming off-the-shelf tools will run themselves. There’s no silver bullet, as Gunnar puts it, no single vendor or model solves the whole problem. Development, legal, and QA may all end up needing different tools entirely, and the job is figuring out what’s genuinely efficient for each rather than betting everything on one platform and hoping it generalises.
Fraud is the exception that proves the rule
One area where AI in payments is genuinely proven, not hype, is fraud prevention. This isn’t new — Visa and Mastercard have published research on machine learning in fraud detection for years, well before generative AI entered the conversation. Visa’s own AI models helped it block around $40 billion in fraudulent transactions in a single year, from October 2022 to September 2023 – roughly 80 million transactions, and nearly double the year before. While Mastercard’s 2025 payment fraud prevention report, produced with Financial Times Longitude, found that a significant share of issuers and acquirers have saved millions in fraud losses thanks to AI. It works precisely because it’s forced to: fraud decisions have to happen in subsecond time, on real transaction volume, with real financial consequences for getting it wrong. There’s no room for a proof of concept that only works in a sandbox. That pressure (not looser oversight) is exactly what produced results solid enough to publish.
The actual lesson
The takeaway here isn’t “AI doesn’t work in regulated industries.” It’s closer to the opposite: regulation is what separates the pilots that were always going to fail from the ones built to last. A 42% failure rate isn’t a verdict on the technology. It’s the cost of an industry figuring out, vendor by vendor and POC by POC, which parts of the hype were real.
Want to hear more on AI in payments, agentic systems, and where the real risks are? Watch the full episode of PAYSTRAX TALKS with Gunnar Már Gunnarsson and Petras Abaravičius below, or listen on Spotify 👇
