Skip to main content

Authors: – Nishanth Kandepedu (General Manager, Products, Scientific Products)

In brief: GOSTAR™ Large Molecules has expanded with a new AI-assisted antibody patent curation dataset containing 4,116 monoclonal antibody records extracted from 425 patents across 288 protein targets. By combining Large Language Model (LLM)-assisted sequence extraction with expert scientific validation, the dataset delivers scalable, high-quality biologics intelligence for antibody discovery, competitive intelligence, and AI/ML applications.

A new patent-derived dataset spanning 4,000+ antibody records across 425 patents and 288 protein targets, with 82% of AI-extracted sequences achieving ≥99% accuracy, complemented by expert validation and curation.

Overview

Monoclonal antibodies have become one of the fastest-growing therapeutic modalities in drug discovery, with patent literature serving as one of the earliest sources of novel antibody sequences, target information, engineering strategies, and biologics innovation. While these patents provide valuable scientific insights, extracting structured antibody data is challenging because sequence information and biological annotations are often presented in inconsistent formats across complex patent documents.

To address this challenge, Excelra has expanded GOSTAR™ Large Molecules with an AI-assisted patent curation framework that combines LLM-driven extraction with expert scientific validation. This approach accelerates the identification and standardization of antibody sequences while maintaining the data quality required for biologics research. The resulting dataset comprises 4,116 monoclonal antibody records from 425 patents, covering 288 protein targets and 3,212 unique VH-VL antibody pairs, enabling pharmaceutical researchers, IP teams, and AI/ML scientists to access reliable, patent-derived antibody intelligence at scale.

At Excelra, GOSTAR™ Large Molecules has long provided expertly curated biologics intelligence to support pharmaceutical research. Building on this foundation, we developed a Large Language Model (LLM)-assisted curation framework that allows rapid extraction and standardization of antibody datasets from patent literature, accelerating curation while maintaining scientific oversight and data quality.

For researchers new to GOSTAR™ Large Molecules or evaluating it alongside public databases, see Excelra’s blog on GOSTAR™: The Largest Online Medicinal Chemistry Intelligence Database for an overview of the GOSTAR™ platform philosophy — and our case study on Unveiling the Molecular Code: Antibody Sequence Mining and Target Affinity Analysis for the expert-curated predecessor dataset that this new AI-assisted expansion builds upon.

A new AI-Assisted antibody dataset

Our previously published antibody datasets within GOSTAR™ Large Molecules were built through expert scientific curation, including project-specific datasets developed to support specialised customer requirements. While the earlier dataset established a strong foundation for antibody intelligence, the rapidly growing volume of antibody patent literature called for a more scalable approach to data extraction and curation.

To address this challenge, we developed an AI-assisted curation pipeline capable of identifying antibody sequences and standardizing associated biological metadata from patents.

Using this framework, we generated a new patent-derived monoclonal antibody dataset comprising:

Dataset Metric Count
Patents processed 425
Records extracted 4,116
Unique heavy-chain (VH) sequences 2,785
Unique light-chain (VL) sequences 1,814
Unique VH-VL antibody pairs 3,212
Protein targets covered 288
Target-antibody associations 3,236

Table 1. GOSTAR™ Large Molecules AI-assisted antibody dataset coverage

This dataset represents an AI-assisted extraction effort designed to complement the existing GOSTAR™ Large Molecules antibody content and accelerate the incorporation of newly disclosed antibody sequences into the platform.

From patent documents to searchable biological intelligence

Capturing antibody sequences is only one aspect of biologics curation. To enable meaningful searching and downstream analysis, biological entities must also be standardized and linked to widely accepted reference identifiers.

For each antibody, the workflow captures and standardizes:

  • Variable heavy (VH) and light (VL) chain sequences
  • Antibody or clone names
  • Protein target names
  • Target synonyms
  • Gene symbols
  • UniProt accession numbers
  • Biological source names
  • Patent metadata
  • Evidence linking each annotation back to the source document

By normalising these entities, antibodies directed against the same target can be consistently identified even when different patents use alternative nomenclature or organism descriptions. This level of standardization enables target-centric analyses and facilitates integration with other datasets.

The standardization approach described here — linking every antibody record to UniProt accession numbers, gene symbols, and source document evidence — is what makes the dataset interoperable with other biologics and genomics tools. For context on how structured, analysis-ready biologics datasets integrate with AI and ML workflows in drug discovery, see Excelra’s case study on Structured and Analysis-Ready Data for AI/ML-Based Drug Discovery — demonstrating the data architecture principles that underpin GOSTAR™ Large Molecules’ approach to biologics intelligence.

Coverage across therapeutically relevant targets

The AI-assisted dataset spans 288 unique protein targets, reflecting the breadth of antibody research across oncology, immunology, inflammation, and other therapeutic areas.

The most represented targets include:

Target
Cluster of Differentiation 276 (CD276/B7-H3)
Delta-Like Ligand 3 (DLL3)
T-Cell Surface Glycoprotein CD3 Epsilon Chain (CD3ε)
Interleukin-4 Receptor Subunit Alpha (IL-4Rα)
5′-Nucleotidase (CD73/NT5E)
Interleukin-31 (IL-31)
B-Cell Activating Factor (BAFF/TNFSF13B)
Transferrin Receptor Protein 1 (TFRC/CD71)
Butyrophilin Subfamily 3 Member A1 (BTN3A1)
Cluster of Differentiation 40 (CD40)

Table 2. Top 10 protein targets in the AI-assisted antibody dataset

These biological targets represent active areas of antibody research, particularly within immuno-oncology and immune modulation, while also highlighting the diversity of therapeutic programmes represented in recent patent literature.

Beyond individual sequences, the dataset allows researchers to explore target-specific antibody landscapes, sequence diversity, competitive patent activity, and emerging areas of therapeutic innovation.

The immuno-oncology targets at the top of this list — CD276/B7-H3, DLL3, CD3ε — reflect the same therapeutic priority areas driving bispecific and T-cell engager development in the industry. For competitive intelligence on how these targets are positioned in the broader drug discovery landscape, Excelra’s Drug Target Dossier: Target Intelligence for Data-Driven Drug Discovery blog describes how multi-source target dossiers — combining patent intelligence, clinical data, SAR data, and mechanistic evidence — support go/no-go target decisions. The GOSTAR™ Large Molecules antibody dataset contributes the patent sequence and competitive coverage layer to such analyses.

Evaluating the AI-Assisted extraction framework

Developing an AI-assisted extraction workflow requires more than generating structured outputs; the extracted data must also meet the quality standards expected for scientific curation.

To assess sequence-level extraction quality, 5,007 antibody sequences were compared with their curator-validated sequences.

Quality Metric Result
Sequences evaluated 5,007
Sequences with ≥99% accuracy 4,087
Proportion with ≥99% accuracy 82%

Table 3. Accuracy of AI-assisted antibody sequence extraction

Overall, 82% of extracted sequences showed ≥99% sequence identity with the curator-validated sequences.

Sequences that did not meet the ≥99% identity threshold underwent further curator review. These included sequences from complex disclosures, ambiguous sequence presentations, or patent-specific formatting that required expert scientific assessment.

AI assists. Scientific expertise validates.

AI-assisted extraction allows antibody patent data to be processed at a speed and scale that would be difficult to achieve through manual extraction alone. But extracting a sequence is only one part of building a high-quality antibody database.

Expert curators review the extracted information, resolve ambiguous disclosures, validate biological relevance, standardize target and biological annotations, and maintain data consistency. This combination brings together the speed and scale of AI-assisted extraction with the scientific expertise and curation standards behind GOSTAR™ Large Molecules.

Value to GOSTAR™ Large Molecules Users

Antibody discovery teams: compare or analyze antibody sequences and benchmark the sequence diversity reported for a target against in-house programmes.

IP and portfolio teams: monitor antibody activity across targets, patents and competing initiatives across companies.

Data scientists / ML teams: access a high-quality, diverse set of sequence and target data for training and validating models.

The ‘AI assists, scientific expertise validates’ model described here is central to how GOSTAR™ Large Molecules maintains the quality standard that researchers and AI/ML teams depend on. For biologics teams evaluating GOSTAR™ Large Molecules for portfolio augmentation or competitive intelligence, see Excelra’s case study on Portfolio Augmentation for a Potential Biologic Drug — demonstrating how GOSTAR™ biologics intelligence has been applied to identify and evaluate biologic candidates against commercially active targets.

Expanding GOSTAR™ large molecules

The antibody landscape continues to evolve rapidly, driven by advances in multi-specific antibodies, antibody engineering, and novel therapeutic targets. As patent disclosures increase in both volume and complexity, scalable curation approaches become increasingly important.

The AI-assisted antibody dataset described here represents the next step in expanding GOSTAR™ Large Molecules. The dataset complements the existing expert-curated antibody content by accelerating the incorporation of newly published patent-derived antibody data while maintaining the scientific quality and biological standardisation expected by researchers.

The AI-assisted antibody patent curation framework described in this article represents a defined workflow comprising three components: an LLM-assisted extraction pipeline that identifies VH/VL sequences and standardizes biological metadata from heterogeneous patent documents, an accuracy evaluation layer that benchmarks extracted sequences against curator-validated references (82% at ≥99% identity in the current dataset), and an expert scientific curation stage that resolves ambiguous disclosures and validates all biological annotations. The result — 4,116 antibody records across 425 patents and 288 protein targets — demonstrates that AI-assisted curation and scientific expertise are complementary, not competing, capabilities. GOSTAR™ Large Molecules will continue expanding this dataset as the volume and complexity of antibody patent disclosures increases.

To access GOSTAR™ Large Molecules or to request a dataset demonstration for your antibody discovery, IP intelligence, or ML training data requirements, visit Excelra’s Monoclonal Antibody Curation service page or our Biologics — The Biotech Drugs Transforming Medicine blog for broader context on how antibody biologics intelligence integrates with modern drug discovery workflows.

What is GOSTAR™ Large Molecules and what antibody data does it contain?

GOSTAR™ Large Molecules is Excelra’s biologics intelligence database, providing curated antibody sequence, target, and patent data to support antibody discovery, competitive intelligence, and AI/ML model development. The database combines expert-curated antibody content — built from years of scientific curation for pharmaceutical and biotech clients — with a new AI-assisted patent-derived dataset comprising 4,116 monoclonal antibody records extracted from 425 patents, covering 288 unique protein targets and 3,212 unique VH-VL antibody pairs. For each antibody, the database standardizes variable heavy (VH) and light (VL) chain sequences, protein target names, UniProt accession numbers, gene symbols, biological source information, and patent metadata — with evidence linking every annotation back to its source document. This level of standardization enables researchers to conduct target-centric antibody landscape analyses, sequence diversity assessments, and competitive patent monitoring across therapeutic areas including immuno-oncology, immunology, and inflammation.

How accurate is AI-assisted antibody sequence extraction from patents?

In Excelra’s evaluation of the AI-assisted extraction framework used for GOSTAR™ Large Molecules, 5,007 antibody sequences were compared with curator-validated reference sequences. Of those sequences, 82% — corresponding to 4,087 sequences — showed ≥99% sequence identity with the curator-validated versions. The 18% of sequences that did not meet the ≥99% threshold underwent expert scientific curation review. These sequences typically came from complex patent disclosures, ambiguous sequence presentations, or documents with patent-specific formatting that required expert scientific assessment rather than automated processing. This hybrid approach — AI-assisted extraction combined with expert scientific validation — achieves a quality standard that would be difficult to reach through either approach alone. AI provides the speed and scale needed to process large patent volumes; expert curators provide the scientific judgment needed to handle the ambiguous cases and validate biological annotations that AI cannot resolve reliably.

Which protein targets are most represented in the GOSTAR™ Large Molecules antibody dataset?

The 10 most represented protein targets in the new AI-assisted GOSTAR™ Large Molecules antibody dataset are: CD276 (B7-H3), DLL3 (Delta-Like Ligand 3), CD3ε (T-Cell Surface Glycoprotein CD3 Epsilon Chain), IL-4Rα (Interleukin-4 Receptor Subunit Alpha), CD73/NT5E (5′-Nucleotidase), IL-31 (Interleukin-31), BAFF/TNFSF13B (B-Cell Activating Factor), TFRC/CD71 (Transferrin Receptor Protein 1), BTN3A1 (Butyrophilin Subfamily 3 Member A1), and CD40 (Cluster of Differentiation 40). These targets reflect active areas of antibody research particularly in immuno-oncology and immune modulation. CD276/B7-H3 and DLL3 are among the most active targets in next-generation bispecific and antibody-drug conjugate programmes. CD3ε is a central component in T-cell engager constructs. IL-4Rα is the target of dupilumab, one of the highest-revenue biologics in clinical use. The top-10 target list provides a real-time window into where the industry is investing its antibody discovery programmes.

Why are antibody patents difficult to curate compared to journal literature?

Antibody patents present unique curation challenges that distinguish them from structured scientific databases and even from other types of patent documents. Variable heavy (VH) and light (VL) chain sequences may appear anywhere in the document — in dedicated sequence listings at the back of the patent, in tables within the specification, in figures, or described verbatim within the claim text — and the format used for each varies significantly between patent filers, jurisdictions, and time periods. The same protein target is frequently described using multiple aliases, organism information is inconsistently reported, and the critical biological metadata required for downstream analysis — such as epitope information, affinity values, and engineering modifications — is scattered throughout the specification rather than presented in a structured format. For large-scale antibody competitive intelligence, manually extracting and standardizing this information across hundreds or thousands of patents is both time-consuming and error-prone. LLM-assisted extraction pipelines can identify and parse sequences and metadata at scale, but require expert scientific validation to handle the ambiguous cases that arise from the inherent heterogeneity of patent documents.

How can ML teams use GOSTAR™ Large Molecules for antibody model training?

GOSTAR™ Large Molecules provides ML and data science teams with a high-quality, curated antibody sequence and target dataset that addresses one of the core challenges in antibody ML model development: the availability of validated, non-redundant, paired VH-VL sequence data with accurate target annotations. The new AI-assisted dataset contains 3,212 unique VH-VL antibody pairs linked to 288 protein targets, with each record standardized to UniProt accession numbers and gene symbols. The 82% ≥99% accuracy rate on sequence extraction, combined with expert curation of all ambiguous sequences, means the dataset reflects a quality standard appropriate for use as training or validation data. For teams developing antibody language models, humanization models, affinity prediction models, or sequence diversity analyses, GOSTAR™ Large Molecules provides patent-derived sequence diversity — covering targets and engineering strategies disclosed in patents before they appear in public databases — that complements public resources like SAbDab and the Protein Data Bank.

What is the difference between expert-curated and AI-assisted biologics curation?

Expert-curated biologics curation relies on trained scientific curators who manually read patent documents, identify relevant information, resolve ambiguous descriptions using their domain knowledge, and enter standardized data into a structured database. This approach produces high-quality, scientifically validated records but is limited in throughput by the capacity of the human curation team. AI-assisted biologics curation uses LLMs and other machine learning approaches to automate the extraction and initial standardization of biological data from patent documents — processing at a speed and scale that human-only workflows cannot match. The practical limitation is that AI extraction introduces errors on ambiguous, complex, or unusually formatted disclosures, which require expert scientific review to resolve. The most effective approach — as implemented in GOSTAR™ Large Molecules — is a hybrid model: AI handles initial extraction and standardization at scale, achieving ≥99% sequence accuracy for 82% of sequences in the current evaluation, while expert curators review the remaining 18% of ambiguous cases and validate all biological metadata annotations. This hybrid model maintains the quality standard of expert curation while enabling the throughput needed to keep pace with the growing volume of antibody patent literature.

Access GOSTAR™ Large Molecules Antibody Data

GOSTAR™ Large Molecules now includes a new AI-assisted patent-derived monoclonal antibody dataset: 4,116 records, 425 patents, 288 protein targets, 3,212 unique VH-VL antibody pairs, with 82% of sequences at ≥99% extraction accuracy. Available to antibody discovery teams, IP and portfolio teams, and ML/data science teams requiring high-quality, patent-derived biologics intelligence.