CMOtech US - Technology news for CMOs & marketing decision-makers
United States
Backbase banking AI beats GPT-4.1 in live deployment

Backbase banking AI beats GPT-4.1 in live deployment

Thu, 30th Jul 2026 (Today)
Karen Joy Bacudo
KAREN JOY BACUDO Finance Editor

Backbase has published a study showing that a 12-billion-parameter banking AI model outperformed GPT-4.1 in a live deployment at a large US financial institution. The company's AI research team developed the model.

Over seven months, it resolved 7.1 percentage points more customer queries across a sample of 3,297 interactions, according to the study. The system was trained to refuse answers when the available evidence did not support a response, rather than produce a confident reply regardless.

The findings add to a wider debate over how generative AI systems should behave in regulated sectors such as banking, where inaccurate answers on fees, rates or policies can create compliance and customer service risks. According to the study, a model that was more willing to say it did not know still solved more customer problems than alternatives that either guessed too often or declined too frequently.

Backbase's researchers said 22% of the training examples used for the model had no correct answer. That design was intended to teach the system to identify the limits of the underlying information and refuse to respond when evidence was missing.

As a result, the model recorded a 12% refusal rate. That was higher than the 4.3% rate for its untuned base model, but lower than the 20.2% rate reported for GPT-4.1 in the same research.

According to the study, the higher refusal rate did not reduce usefulness. The model still achieved better customer query resolution in the live environment, which Backbase described as statistically significant.

Independent evaluation also placed the model ahead of GPT-4.1 on answer quality. It scored 6.21 against 5.72 for GPT-4.1 on a 10-point scale, while citation grounding improved by 2.3 points through answers tied directly to source documents.

The model also produced stronger results on FinanceBench, a public benchmark based on questions about US Securities and Exchange Commission filings. According to Backbase, the work is among the first peer-reviewed accounts of a banking-focused language model assessed in live production.

Cost and speed

The economics outlined in the study are likely to draw attention from banks weighing whether to build or buy AI systems. Backbase said the model ran at roughly USD $0.001 per query on a single graphics processing unit, making it 20 to 50 times cheaper than GPT-4.1 and three to five times faster.

Training costs were put at about USD $1,800. That combination of lower operating cost and improved performance is notable at a time when many financial institutions are still struggling to show clear returns from generative AI projects.

The research team was led by Denys Katerenchuk, Head of AI Research at Backbase, who previously worked at Google and IBM. One key lesson from the study was that data order mattered more than data volume.

On the same underlying data, training the model first on general financial language and then on calibrated refusal produced the best results. Combining all data at once reduced answer quality by more than 40% and pushed refusals to nearly half of all queries.

Sector debate

The work arrives as AI developers and users wrestle with the trade-off between answer rates and reliability. Recent industry research has highlighted how standard training and evaluation methods can encourage systems to guess confidently instead of admitting uncertainty.

Backbase argues that its production results show an alternative approach can work in practice. Rather than treating refusal as a weakness, the study presents it as a mechanism for improving accuracy and trust in customer-facing banking systems.

Some of the production analysis drew on anonymised customer interactions at a large US retail credit union. Backbase also said the broader research was validated across more than 40 financial institutions.

This is the first published work from Backbase AI Research, the team that joined the company through its acquisition of Kasisto. The group focuses on banking-specific AI issues and on moving research into deployed products.

"Off-the-shelf models tend toward hallucination and sycophancy - confident, agreeable answers even without evidence. That's especially risky in banking, where information is complex, technical, and scattered across dozens of documents," said Denys Katerenchuk, Head of AI Research at Backbase.

Backbase's Chief Executive Officer also framed the findings as a challenge to prevailing assumptions in both AI and banking.

"For three years, the AI industry has rewarded models for speed and confidence. Banking has rewarded itself for the same thing for three decades. Saying 'I don't know' got treated as a weakness, not a feature," said Jouk Pleiter, Founder and Chief Executive Officer at Backbase.

"Our research shows the opposite: a model that knows the limits of its own evidence earns more trust, not less."