From Data Silos to Scalable AI: Overcoming the Foundational Barriers in Modern Drug Discovery

Drug discovery remains one of the most notoriously protracted and capital-intensive endeavors in enterprise research and development. According to workshop proceedings published by the National Academies of Sciences, Engineering, and Medicine, shepherding a single pharmaceutical therapy from initial benchtop concept to formal Food and Drug Administration (FDA) approval typically spans a grueling 10 to 15 years. Compounding this challenge, data published in the Journal of the American Medical Association (JAMA) indicates that roughly 90 percent of drug candidates entering clinical development ultimately fail before reaching patients. Many of these expensive failures occur only after organizations have sunk years of capital and resources into a singular, unyielding research direction.
A significant portion of these systemic delays and inflated expenditures can be traced directly to archaic data management practices. Across the biotechnology and pharmaceutical landscapes, discovery data is frequently fragmented, distributed across an array of proprietary platforms, inconsistent file formats, and siloed organizational departments. Researchers from the University of Pennsylvania have formally identified these persistent barriers as major impediments to efficient biomedical research, data sharing, reproducibility, and scientific reuse. In everyday laboratory practice, this fragmentation manifests as high-throughput screening and assay data locked within isolated systems, minimal interoperability between disparate research software, and multidisciplinary teams working concurrently from conflicting data sources.
These acute data integration and standardization hurdles are mirrored in broader industry literature. A landmark data science roadmap authored by a specialized working group from the Structural Genomics Consortium, published in Nature Communications, argues that centralized architectures, standardized terminologies, and fluidly connected research workflows are absolute prerequisites for generating machine learning-ready datasets. As enterprise organizations aggressively expand the integration of artificial intelligence (AI) and machine learning (ML) across their R&D pipelines, the underlying quality, data governance, traceability, and ultimate accessibility of scientific data have grown exponentially more critical.
Reinforcing these industry-wide observations, academic researchers from the University of Maryland, Baltimore County, and the University of Illinois Chicago analyzed official FDA workshop perspectives regarding AI in drug development. They concluded that successful, scalable AI adoption is fundamentally predicated on robust data management protocols and fit-for-purpose datasets capable of supporting reliable algorithmic modeling and satisfying rigorous regulatory scrutiny.
Industry Perspectives on Scaling Enterprise AI
To address these pressing challenges, media platform and advisory firm Emerj recently convened a series of expert discussions confronting a central question facing nearly every modern R&D enterprise: What structural transformations are required to transition artificial intelligence from isolated, localized pilot victories to comprehensive, organization-wide adoption in drug discovery?
In an in-depth internal leadership interview, Barry Bunin, CEO and President of CDD Vault, drew upon two decades of historical data-unification successes and failures observed across the global pharma and biotech sectors. In a complementary webinar titled "Building AI-Ready Foundations for Drug Discovery," Xiong Liu, Director of Data Science and AI at Novartis, and Mitchell Buckley, Application Scientist at CDD Vault, examined these exact operational friction points from inside a major enterprise R&D environment.
The insights synthesized from these industry leaders provide a strategic roadmap for organizations striving to scale artificial intelligence across complex preclinical and clinical drug discovery programs.
The Chronology and Evolution of Research Informatics
To understand the current imperative for AI-ready data foundations, it is valuable to examine the chronological progression of enterprise informatics over the past twenty years. In the early 2000s, high-throughput screening and combinatorial chemistry generated unprecedented volumes of chemical and biological data, prompting many pharmaceutical firms to invest heavily in disparate, departmental data silos. While these localized systems solved immediate record-keeping needs for specific assay groups or chemistry teams, they inadvertently created infrastructural walls.
By the 2010s, as cloud computing matured, organizations attempted to unify these silos through massive "data lakes." However, without rigorous semantic standardization, these data lakes often devolved into unstructured "data swamps," where data was centralized physically but remained functionally disconnected. Entering the mid-2020s, the explosive rise of deep learning and generative AI models exposed the severe limitations of these legacy repositories. Modern AI algorithms require not just large volumes of data, but meticulously cleaned, annotated, and context-rich inputs. Consequently, the industry has shifted its primary focus away from algorithmic novelty and toward fundamental data architecture and metadata governance.
The Architectural Dilemma: Unified Discovery Data vs. Departmental Silos
According to Mitchell Buckley of CDD Vault, fragmented data infrastructure is the primary hurdle any drug discovery organization must clear before artificial intelligence can deliver tangible commercial or scientific value. When disparate databases across individual research teams, specialized programs, or geographically isolated sites lack interoperability, pooling the massive datasets required to train and validate machine learning models becomes virtually impossible.
Barry Bunin traces this friction even deeper, pointing to the historical cultural and operational divide between experimental scientists—who generate wet-lab biological and chemical results—and computational scientists, who model those outcomes using algorithms. Without an integrated infrastructure designed around the distinct daily workflows of both groups, organizations forfeit what Bunin defines as the "economics of specialization": the compounding efficiency gained when biologists, medicinal chemists, and data scientists build seamlessly upon each other’s incremental work.
Historically, this persistent disconnect has bred mutual skepticism within research organizations. Computational modelers have occasionally oversold algorithmic predictions, leaving experimentalists holding the bag on multi-year synthesis and assay projects when predictions fail to materialize in the wet lab. As Bunin notes, there has historically been a significant degree of mistrust, hype, and mutual misunderstanding between computational and experimental teams. Bridging this gap is fundamentally an organizational challenge as much as it is a technical one.
Elaborating on this theme, Xiong Liu of Novartis expands the problem to the enterprise scale, emphasizing that technical inconsistencies represent only half the battle; resolving cross-functional data ownership is frequently much harder. Data routinely flows into an enterprise from a diverse array of internal departments and external contract research organizations (CROs), each utilizing distinct formats. Even a technically sound data lake quickly becomes ungovernable when departmental teams dispute ownership of shared datasets.
Liu advocates for a practical sequencing test before R&D leaders commit capital to new AI initiatives. Rather than authorizing technology purchases based on hype, leaders should evaluate whether the data required by a prospective model can be seamlessly pooled and queried across existing systems, and whether a definitive owner has been assigned to each dataset. If the answer is negative, the immediate priority should not be procuring a more advanced machine learning model, but rather establishing clear data ownership and connective tissue between legacy systems.
Metadata Rigor as Trust Infrastructure

Solving structural data silos only brings an organization halfway to its goal. As Mitchell Buckley emphasizes, raw access to pooled data is insufficient if the underlying information lacks consistent annotation. Two independent laboratories can store technically similar assay results, but if one team labels a compound’s biological activity using different criteria or omits critical metadata explaining how the result was generated, the combined dataset becomes toxic for both scientific reproducibility and downstream AI training.
Consequently, Buckley argues that metadata annotation quality functions as core trust infrastructure rather than a mere documentation afterthought. A machine learning model trained on inconsistently annotated data will produce opaque outputs that even its creators cannot fully defend to internal stakeholders, external collaborators, or regulatory bodies like the FDA.
To mitigate this risk, Buckley outlines three mandatory operational disciplines for research teams:
- Enforcing standardized ontological vocabularies across all experimental data entry points.
- Mandating comprehensive contextual metadata capture, including assay conditions, instrument parameters, and chemical purity metrics.
- Treating metadata validation as a strict gatekeeping mechanism prior to database ingestion.
Liu reinforces this perspective by noting that before these rigorous disciplines are implemented within an enterprise, a promising chemical lead discovered in one program typically remains trapped within that specific team. No other department can confidently interpret the nuances of its generation. Once a standardized semantic layer is established, however, that exact same experimental result transforms into valuable, actionable context for entirely separate therapeutic programs and predictive models across the wider organization, multiplying the return on investment for every individual experiment.
Foundation-First Sequencing for Compounding Adoption
Articulating a core strategic principle for enterprise leaders, Xiong Liu advises organizations to build their technological foundations once and scale upward through demonstrated adoption.
"My suggestion is build the foundation once and scale by adoption," Liu explains. "The foundation means data foundations, the semantic layer, AI-ready data. Adoption means showing that it’s delivering promise for business decision-making. Once you have that information, you usually get the green light to scale your AI systems."
Liu illustrates this methodology using examples from genomics and transcriptomics. Instead of constructing bespoke data ingestion pipelines for every newly targeted disease area, an enterprise should consolidate related data types—such as single-cell omics data originating from multiple independent research programs—into a unified environment governed by a consistent semantic layer. Once this stable foundation is operational, specialized teams can query data based on their specific biological context without rebuilding fundamental plumbing.
Furthermore, this architecture enables closed-loop learning. When high-throughput screening data, wet-lab validation results, and algorithmic model outputs reside on a shared, interoperable foundation, empirical discoveries feed directly back into the models that generated the original hypotheses. This establishes a virtuous, iterative cycle rather than a fragmented, one-way handoff.
Barry Bunin observes an identical compounding effect from a software platform perspective. When a scientist defines a clean data structure for a single assay experiment, that structural template should carry forward automatically for all future data uploads, avoiding repetitive manual entry. Bunin attributes long-term industry success to a fundamental baseline habit: "The first thing is centering on a source of truth for your data, having everybody able to see and use the data. The better organizations will have multiple departments, or even multiple organizations, working as one, so no time is lost."
Culture and Governance as Gates to Enterprise-Scale AI
Even the most sophisticated data foundation will stall without a governance framework capable of shepherding a proposal from a localized pilot project to a secure enterprise-wide deployment. Liu divides enterprise governance into two distinct layers that organizational leaders frequently confuse, often to the detriment of their funding proposals.
"For larger-scale AI systems, we need governance at different layers," Liu states. "One is basic risk and IT governance, where the data has to be secured and de-identified. Then there is scientific and functional governance: are we building the right AI, and is it actually working? Those are the questions we have to answer for leadership and budget owners before they can give us the green light."
Liu observes that R&D proposals frequently fail to secure executive sponsorship because project teams present only the technical and algorithmic layers while omitting rigorous value sizing. Without clear quantitative metrics addressing the business case, budget owners lack the necessary confidence to approve large-scale deployments.
Mitchell Buckley connects this administrative rigor back to organizational trust. Research teams that extract sustained, measurable value from artificial intelligence are invariably those whose scientists and executive leadership genuinely trust the underlying data integrity and model interpretability. Keeping human domain experts explicitly in the loop by design ensures that algorithmic outputs are consistently cross-examined against empirical biological reality.
Implications for the Future of Preclinical R&D
The convergence of insights from industry leaders at CDD Vault and Novartis highlights a definitive strategic shift within pharmaceutical research and development. As the industry moves past the initial wave of AI hype, enterprise success is increasingly decoupled from raw algorithmic complexity and tied directly to foundational data hygiene.
Organizations that prioritize semantic standardization, rigorous metadata annotation, cross-departmental data ownership, and dual-layered governance will be uniquely positioned to compress drug discovery timelines and mitigate the extraordinarily high clinical attrition rates that have plagued the sector for decades. Conversely, enterprises that continue to rely on fragmented data silos and retrofitted IT infrastructures risk perpetual pilot stagnation, unable to scale their digital innovations into approved therapeutics that reach patients in need.







