Book a demo

Basecalling & Sequencing Data Patents: Top Companies & Trends 2026

Basecalling & Sequencing Data Patents: Top Companies & Trends 2026
https://www.patsnap.com/resources/blog/rd-blog/basecalling-and-sequencing-data-processing-patent-landscape/ · Patsnap · data cut-off 2026-07-31 · downloaded from the live page
Sequencing Technology · Patent Landscape
Basecalling and Sequencing Data Processing Patents: Who Leads and Where the Field Is Still Open
  • 44.4% concentration at the top. The five leading assignees hold 293 of 660 records in scope — filing here means competing directly with entrenched claim positions, not filling empty space.
  • Filing has kept growing, modestly. From 2021 to 2024, the last fully comparable years, filings rose from 35 to 37 — a +6% gain, not a boom, and not a decline.
  • Bioinformatics claims now rival wet-lab claims. G16B bioinformatics classifications appear on 42.7% of all 660 records, nearly matching C12Q's 58.6% — the contest has moved as much into software as into chemistry.
Get a prior-art report on your approach
660
Published Records
44%
Top-5 Share of All Records
+6%
Filing Growth 2021→2024
US
Leading Jurisdiction

Filing growth compares 2021 (35 records) with 2024 (37) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 660 records in scope (CR5), not by the ranked leaders only.

Published byPatsnap Research··7 min readSourced from Patsnap Eureka
Overview

What this landscape covers

This dataset tracks 660 patent records filed between 2015 and mid-2026 that combine basecalling, sequencing data processing or variant-calling pipeline claims with concerns like signal-to-sequence conversion, compute cost, quality score calibration, reference bias, pipeline versioning or reproducibility. It sits at the intersection of sequencing instrumentation and the software stack that turns raw signal into usable genomic calls, which is why classifications split heavily between C12Q molecular assay claims and G16B bioinformatics claims rather than sitting cleanly in one camp.

Because publication lags filing by roughly 18 months, the 2025 and 2026 figures in any trend view are still filling in and should not be read as a slowdown. The more reliable signal is the 2021-to-2024 window, where filings moved from 35 to 37 records — a modest but real increase in a field that already has a concentrated leadership group.

Filing activity and technology composition, 2015–2026
  1. 1ILLUMINA INC180
  2. 2EDICO GENOME CORP41
  3. 3DIGITAL BIOTECHNOLOGIES INC29
  4. 4OXFORD NANOPORE TECH LTD22
  5. 5COMPLETE GENOMICS INC21
  6. 6CEMM FORSCHUNGZENTRUM FUER MOLEKULARE MEDIZIN GMBH21
  7. 7ARES TRADING SA19
  8. 8PACIFIC BIOSCIENCES OF CALIFORNIA INC18
  9. 9ELEMENT BIOSCIENCES INC18
  10. 10CANCER RESEARCH TECHNOLOGY LTD18
Source: Patsnap Eureka. Assignee ranking and totals. Derived from a Patsnap search on Basecalling and Sequencing Data Processing covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Let an AI agent run this analysis on your own technology

Pick a task. Every answer cites the patents behind it.

10,000 free credits to start
The Numbers

Filing trends and technology composition

Two views of the same 660 records: how filing activity has moved year over year, and which IPC subclasses carry the claims.

Filing activity, 2017–2026

Filings ran from 81 in 2017 to a peak of 96 in 2022, before falling off in the most recent years shown — a pattern consistent with publication lag rather than reduced inventive activity. The comparable 2021-to-2024 window shows a +6% increase, from 35 to 37 records.

Filing activity, 2017–20260255075100812017201820192020202196202220232024202522026Most recent year is partial — publication lag means later filings are not yet visible.

IPC subclass composition

C12Q (measuring and testing involving enzymes or DNA) appears on 58.6% of the 660 records, and G16B (bioinformatics) on 42.7% — the two largest classes by a wide margin. G01N, C12N and G06F each cover a meaningful minority, and A61K, A61P and C07K appear on smaller shares, marking this chiefly as an assay-plus-software field rather than a therapeutics one. Records can carry more than one class, so these shares sum to well over 100%.

IPC subclass compositionC12Q · Measuring & testing involving …38758.6%G16B · Bioinformatics28242.7%G01N · Material analysis & testing12919.5%C12N · Microorganisms & genetic engin…11617.6%G06F · Electric digital data processi…10215.5%A61K · Medicinal preparations436.5%A61P · Therapeutic activity of compou…385.8%C07K · Peptides & proteins304.5%Other24336.8%

Shares are the percentage of the 660 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.

Source: Patsnap Eureka. Filing trend and technology composition. Derived from a Patsnap search on Basecalling and Sequencing Data Processing covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.

Go deeper on Basecalling and Sequencing Data Processing with Eureka

This page is one run against one query. Ask Eureka your own question about basecalling and sequencing data processing and every answer comes back with the patent numbers behind it.

Try Eureka
Key Patents

Most-cited records and a representative claim

Representative Filing
WO2023009758A12023-02-02

WO2023009758A1 — Quality score calibration of basecalling systems (Illumina, Inc., 2023)

ILLUMINA, INC.

A method of generating base calls by a base caller is disclosed. The method includes receiving a plurality of sensor data from a flow cell, identifying a second range such that at least a threshold percentage of the sensor data falls within it, mapping a subset of that data into a third range to produce normalized sensor data, and processing the normalized data in a base caller to call bases.Filed by Illumina, Inc. — illustrates how calibration of raw sensor data ahead of the base-calling step itself has become a patentable chokepoint, separate from the base-calling algorithm.

WO2023009758A1 — patent drawing 1WO2023009758A1 — patent drawing 2
View full filing
Highest-cited records in scope
#Publication no.Patent titleCitations
1US20100331194A1Nanopore sequencing devices and methods442
2US6223127B1Polymorphism detection utilizing clustering analysis337
3US20130157870A1Methods for obtaining a sequence320
4WO2012142611A2Sequencing small amounts of complex nucleic acids315
5US20130124100A1Processing and Analysis of Complex Nucleic Acid Sequence Data234
6WO2013036929A1Methods for obtaining a sequence226
7US6017434AApparatus and method for the generation, separation, detection, and recognition of biopolymer fragments187
8WO2013188831A1Uniquely tagged rearranged adaptive immune receptor genes in a complex gene set163
9WO2013086464A1Markers associated with chronic lymphocytic leukemia prognosis and progression151
10US20140322716A1Uniquely Tagged Rearranged Adaptive Immune Receptor Genes in a Complex Gene Set150

Citation counts favour older records simply because they have had more time to accumulate citations inside the searched corpus — treat this table as a map of influence, not a ranking of current importance.

Each row carries its publication number; clicking a row searches Eureka by that number.

Source: Patsnap Eureka. Citation counts and representative records. Derived from a Patsnap search on Basecalling and Sequencing Data Processing covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Run it yourself

Put your own technology through the same analysis

 
Where to run it
Fastest

Eureka on the web

When you want the answer in the next five minutes.

The agent works the prompt against patents and technical literature, citing every source.

Run your analysis now →
For builders

MCP server & REST API

When it has to run inside your own pipeline.

Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.

Browse MCP servers →
Insights

What the data means for filing strategy

Three patterns stand out once concentration, growth and class composition are read together.

Concentration
44.4% of 660
held by the top 5 assignees

The leadership group is narrow and stable

The leading assignee alone accounts for 180 records, with the fifth-place holder at 21 — a steep drop-off that marks a small group of sequencing platform makers as the entities whose claims any new filer must design around first.

Based on the 660-record assignee ranking
Growth
+6%
2021 → 2024 filings

Growth is real but unspectacular

Filings moved from 35 in 2021 to 37 in 2024, the last years comparable without publication-lag distortion. That is expansion, not a surge — a field being built on steadily rather than one attracting a rush of new entrants.

2021–2024 comparison, publication-lag adjusted
Composition
58.6% vs 42.7%
C12Q vs G16B share of 660 records

Software claims are catching up to assay claims

C12Q still leads as the largest class, but G16B bioinformatics classifications now sit close behind at 42.7% of records — a sign that pipeline, calibration and computational claims are being filed nearly as heavily as the underlying molecular methods.

IPC subclass shares, records in scope
Collaboration
10 pairs
co-assignee pairings identified

Cross-filing is limited but concentrated

Only ten co-assignee pairs appear across the dataset, and the strongest of them link a single sequencing company with its own software subsidiary or link a diagnostics company to research-institute partners — collaboration here is closer than it is broad.

Strongest pairs by shared filing count
Eureka AI Agent
Looking for what nobody has claimed yet?

Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to basecalling and sequencing data processing, with the prior art for and against each one.

Find the white space →
Source: Patsnap Eureka. Co-assignee relationships and derived observations. Derived from a Patsnap search on Basecalling and Sequencing Data Processing covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Players

Who is filing, and who has slowed

The ranked leaders are dominated by sequencing instrument makers and their bioinformatics arms, but recent-year momentum tells a different story than the cumulative totals do.

Leader
180 records
cumulative filings

A single leader holds a wide lead

The top-ranked assignee's cumulative total is more than eight times the fifth-place holder's 21 records, reflecting a long history of filing across both instrumentation and downstream data processing claims.

Cumulative record count, full ranking
Momentum
-100% YoY
several leaders, latest year

Cumulative leaders are quiet in the latest year

Several of the largest historical filers show zero filings in the most recent year with year-over-year drops recorded at -100%, which is consistent with publication lag rather than withdrawal from the field — but it does mean the visible pipeline from these firms looks thinner than their portfolios suggest.

Recent-year filing counts by assignee
Long tail
100 companies
in the full ranking

A broad base of single- and few-filing entrants

Beyond the concentrated top group, the ranking runs to 100 companies, many holding only a handful of records — research institutes, diagnostics firms and smaller instrument makers each staking a narrow claim rather than building a broad portfolio.

Full assignee ranking, 660-record scope
🔍
Under-claimed technical branches
Sub-areas that show up thinly relative to the volume of core basecalling and pipeline claims — early filing positions may still be available here.
Reference-bias correction methodsPipeline versioning for reproducibilityQuality-score recalibration across chemistriesCross-platform variant call harmonizationCompute-cost-aware pipeline scheduling
Rank all filers by momentum →
Recent-year filing momentum by assignee
AssigneeRecent yearYoY
Ares Trading SA1
Illumina, Inc.0-100%
Edico Genome Corp0-100%
Adaptive Biotechnologies Corp0
Oxford Nanopore Technologies Ltd0-100%
CeMM Forschungszentrum für Molekulare Medizin GmbH0
Cancer Research Technology Ltd0-100%
Pacific Biosciences of California, Inc.0
Source: Patsnap Eureka. Assignee-level momentum. Derived from a Patsnap search on Basecalling and Sequencing Data Processing covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
What's Next

Where to take this analysis

The dataset points to specific next steps depending on whether the goal is freedom-to-operate, portfolio positioning or identifying open claim space.

Map claims against the leading five assignees

With 44.4% of records held by five companies, any new filing in basecalling or calibration should first be checked against their specific claim language, not just their existence in the field.

Explore assignee claims in Eureka

Watch the bioinformatics class for new entrants

G16B's climb to 42.7% of records suggests software-side filers may be growing relative to instrument makers — worth tracking separately from the hardware-heavy C12Q class.

Track G16B filings in Eureka

Test the under-claimed branches for freedom-to-operate

Reference-bias correction and pipeline versioning show thinner claim density than core basecalling — a targeted search can confirm whether a first-mover position is still open.

Run a white-space search in Eureka
Source: Patsnap Eureka. Forward-looking reading of the same dataset. Derived from a Patsnap search on Basecalling and Sequencing Data Processing covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
FAQ

Common questions on this landscape

Answers are grounded in the same dataset. Derived from a Patsnap search on Basecalling and Sequencing Data Processing covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Research Basecalling and Sequencing Data Processing in depth with Eureka

Go past this page: query the whole basecalling and sequencing data processing corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.

Try Eureka

Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.

Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.

Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.

Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.

Help us improve this page

Found incorrect or outdated information? Let us know and we'll get it fixed.