Artificial Intelligence in Finance

Bridging the Data Divide: How Pharma and Biotech Leaders Are Building AI-Ready Foundations for Drug Discovery

Drug discovery remains one of the most resource-intensive and high-risk endeavors in modern enterprise research and development. According to benchmark workshop proceedings published by the National Academies of Sciences, Engineering, and Medicine, guiding a single novel therapy from initial laboratory concept to full Food and Drug Administration (FDA) approval typically spans 10 to 15 years. Compounding this prolonged timeline is an alarmingly high attrition rate: comprehensive analyses published in JAMA indicate that roughly 90 percent of drug candidates entering clinical development ultimately fail before reaching patients. Many of these failures occur only after organizations have poured years of financial investment and scientific labor into a single, isolated research direction.

A significant proportion of these prohibitive costs and operational delays trace directly back to foundational inefficiencies in how research data is generated, stored, and managed. Across the biotechnology and pharmaceutical sectors, discovery data is frequently fragmented across multiple proprietary platforms, incompatible formats, and distinct organizational silos. Researchers from the University of Pennsylvania have highlighted these systemic architectural divisions as persistent barriers to efficient biomedical research, impeding data sharing, cross-lab reproducibility, and the long-term scientific reuse of empirical findings. In practice, this fragmentation manifests as high-throughput assay data trapped in legacy relational databases, severely limited interoperability between specialized research software platforms, and cross-functional teams operating from inconsistent, unstandardized data sources.

Recognizing the urgency of these structural roadblocks, industry and academic working groups have begun mapping out pathways toward data modernization. A data science roadmap authored by a collaborative working group from the Structural Genomics Consortium and published in Nature Communications emphasizes that centralized data architectures, standardized scientific vocabularies, and tightly integrated research workflows are absolute prerequisites for generating truly AI-ready datasets. As enterprise organizations aggressively expand the integration of artificial intelligence and machine learning (AI/ML) across their R&D pipelines, the intrinsic quality, structural governance, historical traceability, and baseline accessibility of scientific data have emerged as critical determinants of success.

Further reinforcing this perspective, researchers from the University of Maryland, Baltimore County, and the University of Illinois Chicago analyzed FDA workshop perspectives on the role of AI in drug development. Their findings underscore that successful, scalable AI adoption is entirely dependent upon robust data management practices and fit-for-purpose datasets capable of supporting reliable predictive model development while instilling necessary regulatory confidence.

To examine what it truly takes to transition AI from isolated proof-of-concept pilot projects to organization-wide operational integration, Emerj recently convened a series of deep-dive discussions with leading voices in preclinical data science and pharmaceutical informatics. In an in-depth internal podcast interview, Dr. Barry Bunin, CEO and President of CDD Vault, drew upon two decades of data-unification successes and failures observed across the global pharma and biotech landscapes. In a companion industry webinar titled Building AI-Ready Foundations for Drug Discovery, Dr. Xiong (Sean) Liu, Director of Data Science and AI at Novartis, joined Mitchell Buckley, Application Scientist and Head of Partnerships at CDD Vault, to dissect these exact friction points from inside a major enterprise R&D environment.

The insights shared by these leaders offer a strategic blueprint for organizations striving to modernize their preclinical informatics infrastructure, eliminate costly operational silos, and establish reliable foundations for scalable computational drug discovery.

The Chronology of Industry Informatics: From Isolated Silos to Enterprise Integration

The structural challenges currently plaguing pharmaceutical R&D informatics are rooted in the historical evolution of laboratory software. For decades, individual scientific departments—ranging from medicinal chemistry and structural biology to high-throughput screening and in vivo pharmacology—procured specialized software tools designed to optimize their localized workflows. While these point solutions successfully accelerated localized data collection, they inadvertently created isolated data repositories that lacked cross-platform communication protocols.

By the early 2000s, as high-throughput screening campaigns began generating unprecedented volumes of chemical and biological assay data, the limitations of these fragmented systems became acute. Early attempts at data integration typically involved manual data extraction, transformation, and loading (ETL) processes, supplemented by cumbersome ad-hoc spreadsheets. These manual interventions introduced transcription errors, degraded data integrity, and made cross-programmatic meta-analyses nearly impossible.

The advent of modern machine learning techniques in the 2010s placed entirely new demands on these legacy architectures. While computational models required vast, clean, and interconnected training datasets to achieve predictive accuracy, enterprise R&D databases remained stubbornly partitioned. Recognizing that traditional software retrofitting was insufficient, forward-thinking organizations began migrating toward unified, cloud-native informatics platforms during the late 2010s and early 2020s. Today, the industry stands at a critical juncture: while advanced generative AI and deep learning algorithms are readily available, their ultimate enterprise utility is bounded almost entirely by the maturity of the underlying data foundations.

Unifying Discovery Data for Scalable Machine Learning Models

Addressing the root causes of data fragmentation requires confronting the cultural and technical divides that traditionally separate experimental scientists from computational modelers. Mitchell Buckley of CDD Vault identifies fragmented data infrastructure as the primary obstacle any drug discovery organization must clear before artificial intelligence can deliver tangible business value. When databases across distinct project teams, research programs, or international geographic sites operate in isolation, organizations lack the unified data pool required to train and validate robust machine learning models.

Dr. Barry Bunin traces this friction further back to the historical communication gap between experimentalists generating physical lab results and computational scientists attempting to model them. Without a unified informatics platform designed around the daily operational realities of both groups, organizations forfeit what Bunin defines as the "economics of specialization"—the collaborative efficiency gained when biologists, chemists, and data scientists build seamlessly upon each other’s empirical discoveries.

Historically, this disciplinary divide fostered mutual skepticism. Computational modelers occasionally overslept the predictive capabilities of their early algorithms, while experimentalists were left shouldering multi-year synthesis and assay projects when those predictions failed to materialize in the wet lab. As Bunin observes, overcoming this historical baggage requires addressing organizational alignment just as rigorously as technical architecture.

Mitchell Buckley elaborates on the core operational hurdles that organizations must resolve:

"The primary issue is different data silos. If databases aren’t talking to each other, there isn’t an opportunity to pool that data together and feed it into these models. A secondary issue: even a well-organized database creates problems if the data isn’t properly annotated, both for reproducibility and because that metadata is critical context for machine learning and AI applications."

— Mitchell Buckley, Application Scientist at CDD Vault

Dr. Xiong Liu expands this diagnostic to the enterprise level, emphasizing that technical data inconsistency represents only half the challenge; the more complex undertaking involves establishing clear cross-functional data ownership. In large pharmaceutical companies, data flows continuously from a sprawling array of internal platforms and external contract research organizations (CROs), each utilizing distinct formats and metadata schemas. Even a technically sound enterprise data lake rapidly becomes difficult to govern or scale when cross-departmental teams dispute data ownership responsibilities.

Liu recommends a practical sequencing test for R&D leadership teams before approving capital for new AI use cases: evaluate whether the underlying data required by the proposed model can be readily pooled and queried across disparate host systems, and confirm whether a designated owner is assigned to each dataset. If an organization cannot definitively answer yes to both questions, funding a more sophisticated machine learning model is premature. The immediate priority must be resolving data provenance, ownership governance, and inter-system connectivity.

The operational contrast between siloed and unified enterprises is stark. In a fragmented organization, a researcher investigating a lead optimization series must manually track down scattered assay results across disparate platforms managed by separate project teams, reconciling inconsistent spreadsheet files by hand before an exploratory model can ingest the data. This manual bottleneck frequently leads to redundant experimentation, with teams unknowingly repeating assays previously conducted by adjacent departments. Conversely, in a unified environment, a researcher queries a single interconnected workspace, instantly surfacing relevant historical data across all related research programs, thereby accelerating decision-making and eliminating redundant resource expenditure.

Enforcing Metadata Rigor to Guarantee Reproducible Outputs

Overcoming structural data silos provides only partial remediation; raw access to pooled data is insufficient if the underlying records lack consistent, standardized annotation. Buckley emphasizes that two research laboratories can store technically comparable assay results, but if one team annotates a compound’s biological activity using proprietary terminology while another utilizes a different nomenclature—or omits critical experimental parameters entirely—the resulting combined dataset becomes unreliable for downstream model training and scientific reproducibility.

How to Build the Unified Data Foundation Drug Discovery AI Depends On - Emerj Artificial Intelligence Research

Consequently, metadata annotation quality functions as essential trust infrastructure rather than a mere documentation afterthought. A machine learning model trained on inconsistently annotated data will inherently generate predictive outputs that even the originating scientists cannot fully explain, defend to institutional review boards, or justify to regulatory authorities.

Buckley outlines three foundational data management habits required to maintain rigorous metadata standards:

  1. Enforcing standardized, controlled vocabularies and ontologies across all biological and chemical data entry points at the point of creation.
  2. Embedding comprehensive contextual metadata—such as assay conditions, cell line lineage, and instrument parameters—directly alongside primary numerical results.
  3. Establishing automated validation gates that prevent unannotated or malformed datasets from migrating into enterprise-wide analytical repositories.

By treating metadata standardization as an operational gate rather than an ad-hoc cleanup task, research operations teams can ensure that AI systems ingest defensible, traceable data inputs. Liu notes that prior to implementing these rigorous standards, a promising discovery insight generated within one therapeutic program typically remained isolated within that specific team. Following standardization, that same empirical finding transforms into valuable, accessible context for parallel drug discovery programs and enterprise-wide AI models, maximizing the cumulative return on every laboratory experiment.

Foundation-First Sequencing to Drive Compounding Adoption

A central theme emerging from industry leaders is the necessity of strategic sequencing: organizations must build robust data foundations before attempting to scale advanced computational applications. Dr. Liu encapsulates this principle in a guiding directive for enterprise executives:

"My suggestion is build the foundation once and scale by adoption. The foundation means data foundations, the semantic layer, AI-ready data. Adoption means showing that it’s delivering promise for business decision-making. Once you have that information, you usually get the green light to scale your AI systems."

— Xiong Liu, Director of Data Science and AI at Novartis

Liu illustrates this sequencing strategy through genomics research. Rather than engineering custom data pipelines for each distinct disease indication, an enterprise can ingest related data types—such as single-cell transcriptomics data originating from multiple independent programs—into a centralized repository governed by a unified semantic layer. Once this foundational architecture is established, individual research teams can query the repository based on their specific biological context without rebuilding foundational data plumbing for each new project.

Furthermore, this unified architecture enables closed-loop learning. When screening data, empirical lab results, and computational model outputs reside on a shared data foundation, wet-lab discoveries automatically feed back into the algorithms that generated the initial lead predictions, establishing a continuous, iterative optimization cycle rather than a linear, one-way handoff.

Dr. Bunin observes a corresponding compounding effect at the software platform level. When a research scientist defines a clean data structure for a single assay experiment, that structure propagates automatically across all subsequent data uploads, eliminating repetitive manual formatting—a phenomenon Bunin characterizes as "same song, second verse." He emphasizes that successful organizations consistently anchor their operations around a single, trusted source of truth:

"The first thing is centering on a source of truth for your data, having everybody able to see and use the data. The better organizations will have multiple departments, or even multiple organizations, working as one, so no time is lost. You treat your partners as intelligent as the scientists you’re working with directly. That’s the first foundational thing."

— Barry Bunin, CEO and President at CDD Vault

Mitchell Buckley notes that early-stage biotechnology companies often possess an operational agility advantage in this regard. Lacking extensive legacy IT infrastructure, emerging biotechs can design their preclinical data stacks intentionally for AI compatibility from day one, avoiding the arduous retrofitting processes that encumber larger, established pharmaceutical enterprises.

Navigating Culture and Governance for Enterprise-Wide AI Scale

Even impeccably engineered data foundations will stall without appropriate governance frameworks capable of sheptarding initiatives from small-scale pilots to full enterprise deployment. Dr. Liu delineates two distinct layers of governance that executive leadership must evaluate independently to secure organizational alignment and budgetary approval:

"For larger-scale AI systems, we need governance at different layers. One is basic risk and IT governance, where the data has to be secured and de-identified. Then there is scientific and functional governance: are we building the right AI, and is it actually working? Those are the questions we have to answer for leadership and budget owners before they can give us the green light."

— Xiong Liu, Director of Data Science and AI at Novartis

Liu points out that AI proposals frequently encounter administrative resistance because project teams present purely technical achievements without accompanying business value sizing. Decision-makers require explicit clarity regarding both compliance security and measurable research acceleration to justify capital allocation.

Buckley connects this administrative rigor back to organizational trust. Long-term value creation in computational drug discovery relies on the confidence that bench scientists and executive leadership place in underlying datasets and predictive models, extending far beyond raw algorithmic sophistication. Maintaining human-in-the-loop oversight ensures transparency and prevents teams from becoming trapped in endless cycles of re-litigating pilot projects.

Ultimately, Dr. Bunin locates the bedrock of this trust within executive and cultural behavior. Dismantling long-standing departmental silos—and bridging the collaborative gap between internal research groups and external industry partners—requires deliberate cultural leadership that no software platform or governance policy can substitute for on its own.

Broader Implications and Future Outlook

The collective insights shared by leadership at CDD Vault and Novartis underscore a transformative paradigm shift currently underway across the life sciences sector. As artificial intelligence transitions from a futuristic novelty into an indispensable driver of preclinical R&D, the competitive advantage in drug discovery will increasingly belong not to organizations wielding the most complex algorithms, but to those that have methodically cultivated clean, standardized, and interoperable data foundations.

For pharmaceutical executives, biotech founders, and informatics leaders, the strategic roadmap is clear: prioritize foundational data hygiene, establish rigorous metadata standards as operational gates, implement multi-layered governance frameworks, and foster a collaborative culture that bridges the historic divide between experimentalists and computational scientists. By executing these foundational steps deliberately, organizations can break free from the prohibitive timelines and high attrition rates that have historically defined drug discovery, paving the way for predictable, scalable, and compounding innovation in the years ahead.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button