Skip to main content

Authors: – Savita Jayaram (Senior Scientific Manager II, Bioinformatics )

In brief: AI-driven multi-omics orchestration closes the target discovery gap by automating the integration of spatial transcriptomics, scRNA-seq, CITE-seq, and other biological data layers — using AI agents for dataset curation, knowledge graph synthesis, and hypothesis acceleration — so that expert scientists can focus on the biological interpretation and translational judgment required to convert molecular signals into high-confidence, druggable therapeutic targets.

As computational biologists and bioinformatics leaders, we routinely find ourselves awash in petabytes of high-resolution data. We capture the cellular architecture of diseased tissues through spatial transcriptomics, map single-cell heterogeneities via scRNA-seq, and track surface proteins with CITE-seq (Ståhl et al., 2016 [1]; Stoeckius et al., 2017 [2]). Yet, a fundamental question remains: How efficiently are we converting these complex, multi-modal layers into high-confidence, druggable targets?

Traditionally, the pipeline from raw data to a therapeutic target is fragmented. Bioinformaticians build pipelines, data engineers manage infrastructure, and scientists contextualize molecular signals. To truly drive innovation and accelerate R&D digital transformation, we must transition from manual, reactive analysis to automated, intelligent data orchestration.

The fragmentation described here — between data generation, pipeline engineering, and biological interpretation — is one of the defining structural challenges in modern pharmaceutical bioinformatics. Excelra’s blog on Building Omics Data Assets: Key Factors to Keep in Mind examines how organizations can build the infrastructure foundations that make multi-omics orchestration scalable — addressing the data standardization, storage, and retrieval challenges that sit upstream of the AI integration described in this article.

The power of the Multi-Omic lens

Human disease does not operate in a single biological dimension. Chronic inflammatory, fibrotic, and oncological diseases are inherently multi-modal (Hasin et al., 2017 [3]).

Each modality provides only a partial view of disease biology. Bulk RNA-seq captures averaged signals across heterogeneous tissues, while microarrays and single-cell technologies introduce their own technical limitations, including probe bias, dropouts, and ambient RNA contamination. By integrating complementary modalities, we can mitigate individual limitations and construct a more complete representation of disease mechanisms (Argelaguet et al., 2020 [4]).

By integrating multiple molecular layers, we can reconstruct a more comprehensive molecular architecture of disease progression. Spatial analyses reveal how immune cells infiltrate diseased tissues and interact within localized inflammatory niches. In IBD, single-cell and spatial studies have uncovered previously unrecognized cellular interactions that contribute to chronic mucosal inflammation and tissue remodeling (Smillie et al., 2019 [5]). This granular view is where the next generation of precision medicine lives.

The spatial transcriptomics capability described here — mapping gene expression to specific tissue regions to understand how immune cells infiltrate diseased tissue — is directly applied in Excelra’s work on cancer target profiling. Our case study on Advanced Target Profiling in Cancer: A Spatial and Single-Cell Transcriptomics Approach demonstrates how integrating spatial and single-cell data layers enables target prioritization decisions that would be invisible to bulk sequencing approaches.

Enter the AI scout: Automating scientific informatics

Data integration is only half the battle. The real digital transformation happens when we leverage AI for automating data ingestion, quality control, metadata harmonization, and literature contextualization. AI reduces operational burden, allowing scientists to focus on biological interpretation and hypothesis generation.

Imagine an R&D ecosystem where the initial stages of target identification are augmented by AI agents built specifically for scientific informatics:

  • Automated Curation: AI systems can automatically identify, harmonize, and prioritize datasets from repositories such as NCBI GEO based on disease signatures and treatment-response phenotypes.
  • Knowledge Graph Synthesis: AI-enabled knowledge systems can continuously cross-reference emerging multi-omics signatures with scientific literature, clinical evidence, and public databases. This enables rapid assessment of whether molecular signals align with known biology, clinical phenotypes, or historical drug-development outcomes (Santos et al., 2022 [6]).
  • Hypothesis Acceleration: Advances in ML and generative AI are increasingly demonstrating their potential to accelerate biological discovery, from protein structure prediction to biomarker prioritization and target identification (Jumper et al., 2021 [7]; Schneider et al., 2020 [8]).

By automating data ingestion, quality control, metadata harmonization, and initial literature contextualization, we can compress the timeline from data generation to biological hypothesis from months to days.

The AI-accelerated biomarker prioritization described under Hypothesis Acceleration is directly demonstrated in Excelra’s case study on Personalised Psoriasis Treatment: How AI Accelerated Biomarker Discovery — where machine learning applied to integrated omics datasets identified treatment response biomarker signatures that compressed the discovery timeline from years to weeks. For the underlying knowledge graph and data intelligence infrastructure that enables this kind of rapid cross-referencing, see Excelra’s blog on GOBIOM’s Biomarker Database for Enabling Precision Medicine.

Expert scientific insights

Yet, the true bottleneck in precision medicine is not data generation. The bottleneck is the deliberate, rigorous human orchestration required to cross-reference, integrate, and distil disparate molecular layers into high-confidence, druggable targets.

Multi-omics technologies reveal correlations at scale, but causality and therapeutic relevance still require expert interpretation. The most impactful discoveries emerge when computational insights are combined with deep domain expertise and translational understanding (Smillie et al., 2019 [5]; Hasin et al., 2017 [3]).

To drive genuine innovation within our R&D pipelines, we must look beyond individual datasets and embrace a unified, multi-omics framework designed and executed by expert scientists.

This principle — that computational insights must be grounded in deep domain expertise — is the same principle underlying Excelra’s approach to drug target dossier development. Our blog on Drug Target Dossier: Target Intelligence for Data-Driven Drug Discovery describes how structured multi-source evidence synthesis — integrating omics data, literature, clinical genetics, and pathway context — is assembled into decision-ready target intelligence packages that combine computational scale with expert scientific interpretation.

Leading the culture shift in R&D

Implementing this vision requires more than deploying Python scripts, R Shiny visualization portals, or advanced cloud architectures. It requires a mindset shift.

As scientific managers and technology leaders, our role is to bridge the gap between technical execution and executive strategy. We must foster collaboration across bioinformatics, data engineering, and translational science.

The path forward

Organizations that successfully combine AI-driven automation with multi-omics integration will shorten target discovery cycles, improve confidence in target prioritization, and reduce costly downstream attrition (Schneider et al., 2020 [8]). The competitive advantage in modern drug discovery will belong to organizations that can transform complex biological signals into actionable therapeutic insights with greater speed, precision, and scientific rigor. The therapeutic breakthroughs of tomorrow are hidden in the cross-talk between our data layers today.

AI-driven multi-omics orchestration, as described in this article, is the integration of spatial transcriptomics, scRNA-seq, CITE-seq, and other multi-modal biological data layers using AI agents that automate dataset curation, knowledge graph cross-referencing, and hypothesis acceleration — enabling expert scientists to focus their analytical capacity on the biological interpretation and translational judgment that converts computational signals into high-confidence, druggable targets. As this article concludes: the therapeutic breakthroughs of tomorrow are hidden in the cross-talk between our data layers today.

Excelra’s bioinformatics team combines AI-driven data orchestration with deep multi-omics domain expertise to support target discovery, disease landscape analysis, and precision medicine programmes. To explore how Excelra’s capabilities can accelerate your R&D digital transformation, visit our Bioinformatics services page.

 

References

  1. Ståhl PL, et al. (2016). Visualization and analysis of gene expression in tissue sections by spatial transcriptomics. Science. https://doi.org/10.1126/science.aaf2403
  2. Stoeckius M, et al. (2017). Simultaneous epitope and transcriptome measurement in single cells (CITE-seq). Nature Methods. https://doi.org/10.1038/nmeth.4380
  3. Hasin Y, Seldin M, Lusis A. (2017). Multi-omics approaches to disease. Genome Biology. https://doi.org/10.1186/s13059-017-1215-1
  4. Argelaguet R, et al. (2020). MOFA+: A statistical framework for comprehensive integration of multi-modal single-cell data. Genome Biology. https://doi.org/10.1186/s13059-020-02015-1
  5. Smillie CS, et al. (2019). Intra- and inter-cellular rewiring of the human colon during ulcerative colitis. Cell. https://doi.org/10.1016/j.cell.2019.06.029
  6. Santos A, et al. (2022). Knowledge graphs for drug discovery and translational research. Drug Discovery Today. [Citation could not be independently verified — please confirm source/DOI]
  7. Jumper J, et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature. https://doi.org/10.1038/s41586-021-03819-2
  8. Schneider P, et al. (2020). Rethinking drug design in the artificial intelligence era. Nature Reviews Drug Discovery. https://doi.org/10.1038/s41573-019-0050-3

What is AI-driven multi-omics orchestration in drug discovery?

AI-driven multi-omics orchestration in drug discovery is the systematic integration of multiple biological data modalities — including spatial transcriptomics, scRNA-seq, CITE-seq, bulk RNA-seq, proteomics, and metabolomics — using AI agents that automate the data ingestion, quality control, metadata harmonization, and literature contextualization steps that traditionally required months of manual bioinformatics work. The ‘orchestration’ framing distinguishes this from simple data integration: it describes an active, AI-coordinated workflow where automated curation agents identify and harmonize datasets from repositories like NCBI GEO, knowledge graph systems cross-reference molecular signals against published biology and clinical evidence, and hypothesis acceleration tools use ML and generative AI to prioritize targets. The goal is to compress the timeline from data generation to biological hypothesis from months to days — while preserving the human expert oversight required to distinguish meaningful biological signals from computational artefacts.

Why is multi-omics integration necessary for drug target discovery?

Multi-omics integration is necessary for drug target discovery because human diseases — particularly chronic inflammatory, fibrotic, and oncological conditions — do not operate in a single biological dimension. Each individual data modality captures only a partial view of disease biology: bulk RNA-seq produces averaged signals across heterogeneous tissue, masking the cell-type-specific mechanisms that drive pathology; single-cell sequencing reveals cellular heterogeneity but loses spatial context; spatial transcriptomics maps expression to tissue location but with lower depth per cell. By integrating these complementary modalities, researchers can reconstruct a more complete molecular architecture of disease progression — identifying how specific cell populations interact within tissue microenvironments, which molecular pathways are genuinely activated versus averaging artefacts, and which targets are causally implicated rather than merely correlated. In IBD, for example, multi-omics integration uncovered cellular interaction mechanisms driving chronic mucosal inflammation that no single modality could have identified independently.

What are knowledge graphs and how are they used in AI drug discovery?

Knowledge graphs in AI drug discovery are structured computational networks that connect biological entities — genes, proteins, pathways, disease phenotypes, clinical outcomes, drugs, and mechanisms of action — as nodes with explicitly typed relationships as edges. Unlike traditional databases, knowledge graphs are designed for traversal and inference: an AI system can navigate from a genomic variant observed in multi-omics data, through pathway relationships, to known clinical phenotypes and historical drug development outcomes, assessing whether a novel molecular signal aligns with established biology or represents a genuinely new mechanism. In the multi-omics target discovery context, AI-enabled knowledge graph systems continuously cross-reference emerging molecular signatures against scientific literature, clinical evidence, and public biomedical databases — enabling rapid assessment of the therapeutic relevance of computational findings. This capability transforms the literature review step from a weeks-long manual process into a near-real-time automated analysis.

What is the difference between spatial transcriptomics and scRNA-seq for target discovery?

Spatial transcriptomics and single-cell RNA sequencing (scRNA-seq) provide complementary perspectives that, together, are more valuable for target discovery than either alone. scRNA-seq profiles the transcriptome of individual cells at single-cell resolution, revealing the distinct molecular states of different cell populations within a sample — including rare cell types and transitional states that bulk sequencing would obscure. However, it requires tissue dissociation, which removes information about where in the tissue each cell was located. Spatial transcriptomics preserves the tissue architecture: it maps gene expression data to specific spatial coordinates within a tissue section, revealing how different cell types are organized, how they interact across tissue zones, and how disease processes manifest in specific anatomical regions. For target discovery, spatial transcriptomics is particularly valuable for understanding how immune cells infiltrate diseased tissue, where therapeutic targets are expressed relative to disease pathology, and which cellular niches drive the disease mechanisms the therapy needs to address.

Why does drug target discovery still require human expert interpretation despite AI advances?

Drug target discovery still requires human expert interpretation because multi-omics technologies generate correlations at scale but cannot independently establish causality or therapeutic relevance — the two judgments that determine whether a computational finding becomes a drug development program. A knowledge graph can link a genomic variant to a protein to a pathway to a disease phenotype with high statistical confidence, but determining whether that association reflects a causal disease mechanism, a protective adaptive response, or an epiphenomenon requires biological expertise that current AI systems cannot supply reliably. Similarly, translational relevance — whether a target that drives disease in a mouse model or a cell line will also be relevant in the human disease context, be druggable by available modalities, and be safe to modulate — requires cross-referencing with patient data, clinical trial history, and mechanistic understanding that AI can surface but humans must evaluate. The most impactful discoveries emerge when AI automation handles scale and speed while expert scientists handle interpretation and judgment.

How is AlphaFold being used in AI-driven drug target discovery?

AlphaFold, the deep learning system that predicts protein structure from amino acid sequence with near-experimental accuracy (Jumper et al., 2021), has transformed drug target discovery by making high-quality structural information available for proteins that lacked experimentally determined structures. Before AlphaFold, structural data was available for only a fraction of the human proteome — limiting structure-based drug design to well-characterized protein families. AlphaFold predictions, now available through the European Bioinformatics Institute for essentially the entire human proteome, allow target discovery teams to evaluate the druggability of novel targets identified through multi-omics analysis — assessing whether a target has accessible binding pockets, predicting how mutations affect protein structure in patient subgroups, and enabling rational drug design for previously undruggable proteins. In the multi-omics target discovery workflow, AlphaFold functions as a downstream validation layer: once a target is computationally prioritized from omics signals, structural prediction can assess whether it is structurally amenable to therapeutic intervention.

Ready to Accelerate Target Discovery with AI-Driven Multi-Omics?

Excelra combines spatial transcriptomics, single-cell omics, AI-driven data orchestration, and expert scientific interpretation to accelerate drug target discovery and disease landscape analysis. Whether you are integrating heterogeneous multi-omics datasets, building a knowledge graph for target prioritization, or implementing AI automation across your R&D bioinformatics workflows, our team is ready to collaborate.