Navigating the Frontier of Agentic AI in Regulated Financial Services: Operational Insights from Industry Leaders

Financial institutions globally are confronting a complex operational paradox: while customer expectations demand instant, frictionless service, core workflows continue to rely on fragmented, unstructured data, complex human judgment, and stringent regulatory frameworks. A recent report published by the Bank for International Settlements (BIS) underscores this challenge, warning that expanding artificial intelligence integration across financial services introduces major operational, model, and data governance risks when systems are deployed without standardized controls.
To explore how leading institutions are safely transitioning from isolated AI experiments to operationally governed agentic systems, Emerj’s Yolandi de Weerdt recently hosted an insightful conversation series. She spoke with Yoav Naveh, Co-Founder and Co-CEO of Reindeer, and Ajay Swamy, Senior Executive Product Director for GenAI Products, AIML Platform Management and Governance at JPMorganChase. Their discussions illuminate the structural requirements necessary for deploying agentic AI safely inside complex, heavily regulated banking, financial services, and insurance (BFSI) environments.
The Structural Reality of Unstructured Data and Core Workflows
BFSI environments ingest disparate document formats across core databases, risk engines, and document repositories. According to a research report published by the CFA Institute, a staggering 90% of enterprise data is unstructured. When institutions attempt to manage this data heterogeneity without robust governance, processing error rates skyrocket in document-intensive workflows such as Anti-Money Laundering (AML) and Know Your Customer (KYC).
Compounding this data challenge is the requirement for explainability. Autonomous decision engines that lack transparency inevitably conflict with financial compliance mandates, which require a traceable decision lineage for every action taken. Research highlighted by the ProSight Financial Association demonstrates that non-compliance costs institutions an average of $14.82 million annually—a figure approximately 2.71 times higher than the cost of maintaining robust compliance infrastructure in the first place.
Furthermore, moving financial automation initiatives from testing environments to live production is notoriously difficult. Analysis published by the IEEE Computer Society indicates that while 83% of technology leaders initiate AI projects, a mere 9% successfully operationalize them. This massive attrition rate results in stalled deployments as institutional policies and underlying systems evolve faster than the models can adapt.
Chronology of Automation: From Basic Chatbots to Complex Agentic Systems
The evolution of enterprise automation has shifted dramatically over the past several years. Initial deployments focused primarily on rigid, rules-based robotic process automation (RPA) that struggled with edge cases. As generative AI emerged, organizations rushed to deploy conversational interfaces and basic question-and-answer tools, only to discover that modern financial tasks require multi-step execution across disparate legacy systems.
Recent research published by MR Online reveals that current AI task execution achieves an average success rate of just 30% for complex, end-to-end workplace processes. Recognizing these limitations, standards bodies such as the National Institute of Standards and Technology (NIST) have published guidelines emphasizing that reliable deployment requires continuous human oversight and explicit fallback mechanisms whenever model confidence declines.
Workflow Redesign for Automation-Ready Operations
Addressing the operational friction of multi-system workflows requires treating process redesign as an engineering challenge rather than a simple software upgrade.
"It’s no longer just a question and answer," explains Yoav Naveh, Co-Founder and Co-CEO at Reindeer. "You actually have to take a task across multiple systems, multiple teams, and it’s a lot more difficult to just close a task. Exceptions and edge cases become much more frequent."
Naveh points out that BFSI institutions do not struggle because any individual step is inherently complex, but because ten simple steps spanning five systems, three teams, and inconsistent document formats create an entirely new operational paradigm. Leaders must map the real path work takes, surface where human judgment is actually required, and stabilize exception handling before introducing autonomous execution.
Ajay Swamy of JPMorganChase adds the governance dimension to this redesign, arguing that workflows fail not at scale, but at the seams—where fragmented data, inconsistent entitlements, and shifting policy interpretations collide. Automation-ready workflows require institutions to expose every system the process touches, every data shape it consumes, and every point where human judgment modifies the outcome. Without this visibility, automation only accelerates underlying fragmentation.
Governance-First Control for Explainable Decisions
Governance cannot remain a peripheral safeguard; it must serve as the primary condition determining whether agentic AI can be utilized at all. Institutions routinely misjudge their risk surface by focusing excessively on model capabilities rather than supervisory controls.
"A number or a decision that you cannot trace is a number or a decision that you cannot defend," notes Ajay Swamy. "You’re not really evaluating the technology; the technology is the easy part because the demos all look great. What you’re evaluating is whether you can actually supervise it and explain every step end-to-end."
Swamy emphasizes that explainability is the operational backbone of regulated AI. Every autonomous action must carry lineage, policy context, and a defensible chain of reasoning that can be surfaced instantly. Meanwhile, Naveh adds that institutions often assume accuracy is the primary risk, whereas true exposure emerges when an agent encounters uncertainty. Agents must be engineered to escalate ambiguity rather than mask it, ensuring that silent failures are avoided entirely.
Exception-Aware Escalation and the Two-Loop Model
When an agent reaches the edge of its knowledge base, traditional automation systems typically stall, fail silently, or produce fabricated outputs. Regulated financial workflows—such as complex AML reviews, onboarding checks, treasury operations, and vendor validations—contain edge cases that require institutional context and policy interpretation.
"Everybody’s measuring accuracy, but they’re missing the part of what happens when the agent doesn’t know," Naveh highlights. "You don’t want it to stop, and you definitely don’t want it to hallucinate. You want the agent to raise its hand, reach out to a subject-matter expert, and ask a specific series of questions to get itself out of the hole."
To achieve this, Naveh champions a two-loop governance model. The first loop enables agents to engage subject-matter experts whenever they encounter uncertainty. The second learning loop aggregates those interactions and incorporates the resulting knowledge into future execution. The ultimate objective is to continuously expand agent coverage, improve operational performance, and systematically reduce recurring points of failure over time.
Strategic Operating Models for Sustainable Deployment
As financial institutions look toward the future of enterprise automation, the challenge has shifted decisively from building prototypes to sustaining them over the long term. The ease of developing AI tools has created a blind spot where prototypes multiply rapidly, yet few teams remain accountable for their ongoing maintenance and governance.
"It’s so easy to build today, everybody wants to build it, and nobody wants to maintain," Naveh warns. "People get swept up in how easy it is to build and they don’t think about how it’s going to work in the long run. If you don’t put someone in charge of building a strategy, you end up dependent on tools you can’t govern."
To prevent this accumulation of operational debt, financial leaders must establish disciplined operating models that couple rigorous technical oversight with continuous compliance monitoring. By treating agent evolution with the same seriousness applied to mission-critical financial software, institutions can harness the immense efficiency of agentic AI while maintaining absolute adherence to regulatory mandates. Ultimately, the successful deployment of agentic AI in BFSI will not be defined by how well systems perform under normal operating conditions, but by how intelligently they manage uncertainty when established protocols break down.






