Latest from Cambridge Spark

Why Your AI Prototype Is Lying to You About Being Finished

Written by Cambridge Spark | August 26 2026

Ask most teams building with large language models what stage they're at, and "we have a working prototype" comes up almost immediately. Ask whether that prototype is ready for customers, and the honest answer is usually no. The gap between those two things is, according to Alberto Romero, the single most common failure mode in enterprise AI.

Alberto has spent close to a decade building machine learning systems that get used, first in mobility, then in banking at Citi, then as a founder building AI agents for regulated industries, and now as director of AI engineering at Aviva. In the latest episode of Inside the Algorithm, he makes the case that the industry's excitement problem is just as dangerous as its risk aversion problem.

The Prototype Trap

It's never been easier to build something that looks finished. A good-looking demo can be assembled quickly, and that's exactly what makes it dangerous. Alberto is direct about this: a prototype is built within a specific context to cater to specific use cases, but production readiness is everything else. All the edge cases. All the unforeseen scenarios. Adversarial attacks. The parts nobody thought to test because the demo never surfaced them.

It's a mistake he sees repeated across teams, not just his own, and it sits opposite an equally real problem. Many organisations remain overly cautious about deploying large language models at all, unsettled by their non-deterministic nature in a way that traditional machine learning, with its well-understood risk profiles, never quite triggered. Alberto's view is that both instincts, excessive caution and excessive confidence, come from the same place: not fully understanding what LLMs are actually capable of and what controls make them safe to rely on.

Explainability Has to Be Earned, Not Assumed

One of the sharper moments in the conversation concerns explainability. Large language models are very good at producing a plausible-sounding justification for a decision. The risk, particularly in a regulated business like insurance, is mistaking that justification for the actual reasoning behind the output. Alberto compares it to a child inventing an excuse on the spot. It sounds coherent. It isn't necessarily true.

His answer is a multi-step, agentic approach, closer to a ReAct-style pattern, where a model reasons over a sequence of actions and available context rather than jumping straight to an answer. That produces something genuinely traceable rather than a single plausible sentence generated after the fact. It's a meaningful distinction for any organisation where being able to show your working isn't optional.

Standards Before Tooling

Alberto's move from founding startups to operating inside one of the UK's largest insurers reframed what "the hard part" actually is. At a startup, the constraint is resources. Inside Aviva, running AI across more than seventy use cases, the constraint becomes organisational. His advice is to resist the temptation to go looking for a piece of technology that will solve everything before your standards are settled. Get the standards right first, because the tooling will keep changing underneath you regardless.

That same discipline shows up in how Aviva approaches fine-tuning. Alberto draws a clear line between instruction fine-tuning, used to lock in tone and brand consistency for customer-facing agents handling something as sensitive as a first notification of loss, and knowledge-based fine-tuning, which requires ongoing retraining as information changes and is often better served by retrieval instead. Fine-tuning isn't a default. It's a decision made once the use case and its constraints are actually understood.

Watching Both Ends of the Pipeline

On fraud detection, Alberto frames the problem in adversarial terms, closer to the generator-discriminator dynamic in GANs than a static classification task. His broader point extends beyond fraud: teams tend to obsess over monitoring the distribution of a model's outputs, watching for drift, while paying far less attention to the distribution of its inputs. The world keeps changing after deployment, and a model that was accurate on historical data can quietly stop generalising well without anyone noticing until it's too late.

As AI use cases scale and models become cheaper and faster to run, Alberto's closing point is less about any single technique and more about foundations. The organisations that will handle that scale well are the ones building solid governance and standards now, not the ones waiting for the next model release to sort it out for them.

______________________________________________

Chapter Markers: 

  • (00:00) - Cold open: why prototypes get mistaken for production

  • (02:53) - Avoiding common AI adoption pitfalls in regulated sectors

  • (05:44) - Real explainability versus post-hoc justification in LLMs

  • (09:18) - From startup founder to enterprise: the mindset shift

  • (11:38) - Managing AI across 70+ use cases at Aviva

  • (13:31) - Standards first, technology second

  • (17:19) - Where fine-tuning earns its place

  • (20:33) - Building Aviva's own governed AI platform

  • (23:51) - Fraud detection as an adversarial ML problem

  • (28:34) - Quick fire: the most common AI failure mode

  • (29:37) - What deserves more attention as AI scales

Useful Links:

If you're a data and AI leader looking to build deep technical capability across your organisation, Cambridge Spark is here to help  - explore our AI solutions