How Pharmaceutical Leaders Are Overcoming Data Silos to Scale AI in Drug Discovery

Drug discovery remains one of the most resource-intensive and protracted endeavors in enterprise research and development. Industry benchmarks indicate that developing a single U.S. Food and Drug Administration (FDA)-approved therapeutic typically spans 10 to 15 years. Compounding this challenge, approximately 90 percent of drug candidates entering clinical development ultimately fail to reach patients. These failures frequently materialize after years of substantial financial investment and scientific effort dedicated to a singular research hypothesis.
A significant proportion of these delays and financial losses stem from structural inefficiencies in how organizations manage research data. Across the biotechnology and pharmaceutical sectors, discovery data is routinely partitioned across disparate platforms, incompatible formats, and isolated business units. Researchers from the University of Pennsylvania have highlighted these persistent fragmentation issues as core barriers to efficient biomedical research, data sharing, and scientific reproducibility. In practice, this manifests as experimental assay data residing in siloed software systems, minimal interoperability between proprietary research platforms, and multidisciplinary teams operating from conflicting datasets.
The Roadmap to AI-Ready Infrastructure
Recognizing these systemic bottlenecks, life sciences organizations are increasingly turning toward artificial intelligence and machine learning to accelerate candidate identification and preclinical testing. However, the efficacy of these computational models relies entirely on the quality of the underlying data infrastructure.
A data science roadmap published in Nature Communications by a working group from the Structural Genomics Consortium emphasizes that centralized architectures, standardized vocabularies, and interconnected research workflows are prerequisites for generating AI-ready datasets. As pharmaceutical enterprises expand their reliance on artificial intelligence, data governance, traceability, accessibility, and quality control have transitioned from administrative backwaters to central strategic priorities.
Parallel perspectives published in academic literature by researchers from the University of Maryland, Baltimore County, and the University of Illinois Chicago underscore that successful integration of artificial intelligence in drug development depends heavily on rigorous data management practices. To secure regulatory confidence and ensure reliable model outputs, data must be fit for its intended computational purpose.
To address these pressing industry challenges, media and research organizations have increasingly convened industry leaders to discuss how R&D executives can successfully transition artificial intelligence from isolated proof-of-concept pilots to enterprise-wide adoption. Recent discussions featuring executives and data scientists from Collaborative Drug Discovery (CDD Vault) and Novartis have shed light on the structural prerequisites required to achieve scalable AI integration.
Chronology of Enterprise Data Transformation
The evolution of informatics in drug discovery has undergone distinct phases over the past two decades, reflecting shifts in both computational capability and data volume:
- The Early 2000s: Discovery informatics was largely characterized by local, desktop-based databases and isolated electronic laboratory notebooks (ELNs). Data sharing between departments relied heavily on manual exports, static spreadsheets, and ad-hoc file transfers.
- The 2010s: The rise of cloud computing and high-throughput screening technologies led to an explosion in data volume. Organizations began centralizing information into early enterprise data warehouses. However, semantic inconsistencies and inter-departmental silos remained prevalent, limiting the utility of cross-functional datasets.
- The Early to Mid-2020s: Generative artificial intelligence and advanced machine learning models emerged as powerful tools for predictive toxicology and molecular design. Organizations quickly realized that sophisticated algorithms could not compensate for poorly integrated, non-standardized training data, prompting a renewed focus on data foundations, semantic layering, and governance frameworks.
- The Present Era: Industry leaders are actively decoupling infrastructure investments from flashy model deployment, prioritizing foundational data governance, cross-functional data ownership, and metadata standardization as non-negotiable prerequisites for scalable AI deployment.
Insights from Industry Leaders
During recent leadership panels, veteran informatics experts unpacked the friction points hindering enterprise AI adoption. Barry Bunin, CEO and President of CDD Vault, drew upon two decades of data-unification experiences across the global pharmaceutical landscape. Bunin, who founded CDD in 2004 after serving as an entrepreneur-in-residence at Eli Lilly, emphasized that historical divides between experimental scientists and computational modelers have frequently bred organizational skepticism.
"There’s been a lot of mistrust and hype and misunderstanding in the past between computational and experimental teams," Bunin noted, highlighting that bridging this gap is as much an organizational challenge as it is a technical one.

Echoing these observations from a large enterprise perspective, Xiong Liu, Director of Data Science and AI at Novartis, detailed the complexities of managing data ownership across massive research organizations. Having previously led AI and data science initiatives at Eli Lilly and served as a founding member of Novartis’s global AI Innovation Lab, Liu stressed that technical inconsistencies represent only half the hurdle. The more arduous task involves resolving data ownership boundaries across internal departments and external contract research organizations (CROs).
Mitchell Buckley, Application Scientist and Head of Partnerships at CDD Vault, pointed to fragmented data infrastructure as the primary obstacle confronting discovery teams. According to Buckley, when proprietary databases fail to communicate seamlessly, organizations forfeit the ability to pool historical data—a critical input required to train and validate robust machine learning architectures.
Furthermore, Buckley emphasized that raw data access is insufficient without rigorous annotation. Two separate laboratories may generate structurally similar assays, but if inconsistent nomenclature or omitted metadata characterizes the records, the combined dataset becomes unreliable for downstream computational models.
Strategic Frameworks for Scalable Implementation
To successfully navigate the transition toward AI-driven drug discovery, industry veterans recommend a structured, sequencing-first approach rather than opportunistic software acquisitions.
1. Establishing a Single Source of Truth
Enterprise R&D organizations must consolidate disparate departmental silos into unified, accessible environments. Bunin advocates for centering research operations around a definitive source of truth where internal teams and external partners collaborate seamlessly. By treating external collaborators with the same technical integration standards as internal staff, organizations eliminate redundant experimentation and accelerate discovery timelines.
2. Prioritizing Metadata Standards as a Gate
Rather than treating metadata annotation as a retrospective cleanup task, leading organizations enforce strict semantic standards as an operational gate. Datasets must pass consistency checks against shared ontologies before being ingested into machine learning pipelines. This rigorous approach prevents models from learning on corrupted or unexplainable data inputs.
3. Foundation-First Sequencing
Liu underscores the importance of building robust data foundations and semantic layers before attempting enterprise-wide scaling. Rather than constructing bespoke data pipelines for every isolated disease program, organizations should aggregate related data types into centralized repositories with standardized access layers. Once this foundation is established, individual research teams can query the system based on contextual parameters without continuously rebuilding backend data plumbing.
4. Dual-Layer Governance Models
To secure executive buy-in and budgetary approval, data science leaders must articulate governance through two distinct frameworks. First, basic risk and IT governance must ensure data security, privacy compliance, and de-identification. Second, scientific and functional governance must rigorously evaluate whether the AI application is architecturally sound and commercially viable. Proposals that clearly delineate technical security from quantitative business value are significantly more likely to transition from pilot phases to broad deployment.
Broader Implications for the Biotechnology Sector
The ongoing shift toward AI-ready infrastructure carries profound implications for the economics of drug development. By eliminating data silos, enforcing rigorous metadata standards, and fostering cultural trust between computational and experimental scientists, life sciences organizations can significantly reduce the attrition rates that have historically plagued preclinical research.
While early-stage biotechnology firms often possess an agility advantage due to a lack of legacy IT infrastructure, large pharmaceutical enterprises are increasingly demonstrating that legacy systems can be successfully modernized through deliberate architectural sequencing. Ultimately, the organizations that succeed in deploying scalable artificial intelligence will not be those with the most complex algorithms, but those that have successfully built the foundational data discipline required to support them.







