Navigating the Future of Agentic AI in Banking, Financial Services, and Insurance

Financial institutions globally are encountering an acute operational paradox: while modern software and artificial intelligence models have never been easier to deploy in sandbox environments, the integration of autonomous systems into core production workflows remains constrained by severe structural hurdles. Core banking and insurance processes rely heavily on fragmented, unstructured data, demand nuanced human judgment under pressure, and must operate within intensely regulated legal frameworks. According to a comprehensive operational report published by the Bank for International Settlements (BIS), the accelerating integration of artificial intelligence across financial services introduces profound model, data, and operational governance risks whenever systems are deployed in the absence of standardized, rigorous controls.
To dissect these systemic challenges, Emerj’s Yolandi de Weerdt recently hosted an in-depth conversation series featuring Yoav Naveh, Co-Founder and Co-CEO of Reindeer, and Ajay Swamy, Senior Executive Product Director of GenAI Products, AI/ML Platform Management and Governance at JPMorganChase. The discussions focused on how leading global financial institutions are successfully migrating from isolated, superficial AI experiments to deeply governed, operationally sound agentic systems capable of handling high-stakes workflows.
The Anatomy of Data Fragmentation in Modern BFSI Environments
The friction experienced by financial institutions attempting to automate complex workflows largely stems from the sheer heterogeneity of enterprise data. Legacy architectures ingest disparate document formats across core transactional databases, complex risk engines, and fragmented document repositories. Industry research published by the CFA Institute indicates that an astonishing 90 percent of enterprise data exists in an unstructured format.
When data heterogeneity is left ungoverned, it inevitably drives up processing error rates in document-intensive, mission-critical operations such as Anti-Money Laundering (AML) checks and Know Your Customer (KYC) onboarding. Furthermore, autonomous decision-making engines that operate as "black boxes" without clear explainability conflict directly with core financial compliance mandates. These mandates legally require a traceable, transparent lineage for every financial decision executed on behalf of a client or institution.
Research highlighted by the ProSight Financial Association illustrates the severe financial gravity of these compliance requirements, demonstrating that regulatory non-compliance costs institutions an average of $14.82 million annually—a figure approximately 2.71 times higher than the cost of maintaining robust compliance and governance infrastructure in the first place.
From Testing to Production: The Operational Deployment Gap
Despite high enthusiasm among corporate leadership, financial automation initiatives frequently encounter immense friction when transitioning from controlled testing environments into live, high-volume production. Analysis published by the IEEE Computer Society reveals a striking disparity: while 83 percent of global technology leaders actively initiate artificial intelligence projects, only 9 percent successfully operationalize them at scale. This massive drop-off results in stalled deployments, wasted capital, and technical obsolescence as institutional policies and underlying IT systems continue to evolve.
Moreover, true end-to-end autonomy remains severely limited when handling non-standard, multi-system workflows. Recent research published by MR Online indicates that current artificial intelligence task execution achieves an average success rate of only 30 percent for complex, multi-step workplace processes. Guidelines published by the National Institute of Standards and Technology (NIST) emphasize that reliable enterprise deployment requires continuous human oversight, explicit fail-safes, and clear fallback mechanisms whenever model confidence degrades or encounters anomalies.
Workflow Redesign as an Operational Engineering Imperative
Addressing these limitations requires treating workflow redesign as an engineering discipline rather than a superficial software upgrade. Yoav Naveh of Reindeer emphasizes that financial institutions do not struggle because any single operational step is inherently complex, but rather because the moment ten simple steps span five disparate systems, three distinct teams, and inconsistent document formats, the nature of the work fundamentally changes.
"It’s no longer just a question and answer. You actually have to take a task across multiple systems, multiple teams, and it’s a lot more difficult to just close a task. Exceptions and edge cases become much more frequent," Naveh explains.
According to Naveh, financial leaders must map the real path that work takes across the enterprise, pinpoint exactly where human judgment is applied, and stabilize exception handling before introducing autonomous execution. Ajay Swamy of JPMorganChase reinforces this perspective by noting that institutional workflows rarely fail at scale; instead, they fail at the seams—where fragmented data, inconsistent user entitlements, and shifting policy interpretations collide. Automation-ready workflows demand that institutions expose every system touched by the process, every data format consumed, and every decision point modified by human oversight. Without this granular visibility, automation only accelerates existing structural fragmentation.
Governance-First Architecture for Explainable Decisions
In the realm of regulated finance, governance cannot be treated merely as a protective boundary around an AI model; it must serve as the primary foundational condition that determines whether the technology can be deployed at all. Ajay Swamy argues that financial institutions routinely misjudge their true risk surface by focusing excessively on raw model capabilities rather than rigorous supervisory controls.
"A number or a decision that you cannot trace is a number or a decision that you cannot defend. You’re not really evaluating the technology; the technology is the easy part because the demos all look great. What you’re evaluating is whether you can actually supervise it and explain every step end-to-end," Swamy asserts.
This philosophy establishes explainability as a core design constraint rather than a retrospective reporting requirement. If an autonomous agent acts without immediate traceability, the host institution inherits immediate, unmitigated regulatory exposure. Naveh complements this by highlighting that while organizations obsess over model accuracy, the true operational exposure emerges when an agent encounters uncertainty or when business processes inevitably shift. Consequently, governance systems must be engineered to detect unknown variables early and route them smoothly into structured human intervention, ensuring that silent failures are entirely eliminated.
Exception-Aware Escalation and the Two-Loop Learning Model
When an autonomous agent reaches the boundaries of its programmed competence, the system’s reaction dictates its enterprise viability. Traditional automation systems typically stall, fail silently, or generate erroneous, hallucinated outputs when faced with ambiguous inputs.
"Everybody’s measuring accuracy, but they’re missing the part of what happens when the agent doesn’t know. You don’t want it to stop, and you definitely don’t want it to hallucinate. You want the agent to raise its hand, reach out to a subject-matter expert, and ask a specific series of questions to get itself out of the hole," Naveh explains.
To solve this, Naveh advocates for a sophisticated two-loop operational model. In the first loop, agents encountering uncertainty actively engage human subject-matter experts, articulating precisely what information is missing to resolve the case. In the second loop, the system aggregates these human interactions and feedback over time, continuously incorporating newly acquired institutional knowledge into future executions. This iterative mechanism ensures that the agent expands its operational coverage, improves its performance metrics, and actively reduces recurring points of failure across the enterprise.
Building a Sustainable Strategic Operating Model
As financial institutions look toward the future of agentic systems, the ultimate challenge lies in long-term ownership and maintenance. The accessibility of modern AI development frameworks has enabled internal teams to rapidly prototype impressive applications in mere days. However, these prototypes frequently transform into enterprise liabilities when no clear organizational accountability exists for maintaining, governing, and evolving them.
"It’s so easy to build today, everybody wants to build it, and nobody wants to maintain. People get swept up in how easy it is to build and they don’t think about how it’s going to work in the long run. If you don’t put someone in charge of building a strategy, you end up dependent on tools you can’t govern," warns Naveh.
Mitigating this risk—often described as accumulating "agent debt"—requires financial institutions to apply the same disciplined release management, continuous auditing, and rigorous model-risk oversight to AI agents that they traditionally apply to mission-critical core banking infrastructure. By coupling robust governance frameworks with exception-aware escalation and strategic workflow redesign, financial services firms can successfully transition agentic AI from an experimental novelty into a secure, scalable, and enduring operational asset.






