Lipid nanoparticles (LNPs) are the most widely used delivery vehicles for gene and cell-based therapeutics, prized for their high encapsulation efficiency, tissue penetration, and low immunogenicity. But the data describing them, formulation composition, preparation conditions, physicochemical characterization, and biological outcomes, is scattered across thousands of patents and publications in inconsistent formats. Excelra was engaged to convert this fragmented literature into a single, structured dataset that could feed directly into machine learning models for formulation design and optimization.
Our client
Our client is a biotech company based in the United States working on next-generation cell and gene-based therapeutics. As part of their discovery efforts, the team was investigating novel lipids and lipid combinations that could be deployed in new LNP-based delivery systems, and wanted to apply machine learning to guide that search.
Client’s challenge
The client held a library of close to 1,000 patents and scientific articles referencing LNP formulations, but the information inside them was unusable in its native form:
- Formulation details were buried across tables, figures, and body text, with no consistent structure from one publication to the next.
- Composition, preparation process, characterization, and biological readouts each followed different reporting conventions depending on the source.
- Manually reviewing and reconciling this volume of literature in-house would have been slow, inconsistent, and difficult to scale.
- Without a harmonized dataset, the client’s data science team had no reliable foundation on which to train predictive models for formulation design.
Client’s goals
Excelra and the client agreed on a clear set of objectives for the engagement:
- Identify unique formulations: surface every distinct LNP formulation referenced across the shared literature, avoiding duplication.
- Capture data comprehensively: extract relevant information from all tables, figures, and text, not just abstracts or summaries.
- Structure the full formulation lifecycle: cover composition, preparation process, physicochemical characterization, and in vitro/in vivo outcomes in one schema.
- Preserve chemical context: retain structural information for the lipids involved so the client could explore the available chemical space.
- Deliver ML-ready output: produce a dataset structured and clean enough to plug directly into predictive model training.
Our Approach
Excelra’s scientific data curation team worked through the client’s full document set using a defined screening and extraction workflow rather than ad hoc review. Curators first triaged the roughly 1,000 shared documents to identify which patents and publications actually described relevant LNP formulations, then extracted a consistent set of data fields, publication identifiers, study type, cell line or tissue, API/payload dosage, SMILES notation, physicochemical properties, assay methods, route of administration, activity type, and process parameters such as preparation method and duration, from each one.
Every record then passed through a dedicated quality control step before being finalized, so that inconsistencies introduced by manual extraction across ~1,000 source documents were caught and corrected rather than carried into the deliverable.
Our solution
The result was a single structured database spanning ten interconnected dimensions of LNP science, formulation screening, composition, characterization, stability, production methods, payload details, safety and toxicity, in vitro studies, in vivo studies, and drug-lipid interactions, so that formulations could be compared and analyzed on a common basis for the first time.
fig 1: Key data dimensions captured across the LNP database
The finished dataset was delivered in Excel format, giving the client’s data science team a clean, harmonized foundation they could load directly into their internal AI/ML platform to train predictive models for LNP design and optimization.
Conclusion
By replacing scattered, inconsistently reported literature with one harmonized dataset, Excelra gave the client a foundation for data-driven formulation design that would have been impractical to build in-house at the same speed and consistency.
- ~1000 patents & publications screened
- 10 structured data dimensions captured
- 1 unified, ML- ready dataset delivered
The curated dataset gave the client’s team a consistent, analysis-ready view of the LNP formulation landscape, freeing scientists from manual literature review and putting a validated data foundation behind their AI/ML-driven search for novel lipid delivery systems.
