Basecalling & Sequencing Data Patents: Top Companies & Trends 2026
- 44.4% concentration at the top. The five leading assignees hold 293 of 660 records in scope — filing here means competing directly with entrenched claim positions, not filling empty space.
- Filing has kept growing, modestly. From 2021 to 2024, the last fully comparable years, filings rose from 35 to 37 — a +6% gain, not a boom, and not a decline.
- Bioinformatics claims now rival wet-lab claims. G16B bioinformatics classifications appear on 42.7% of all 660 records, nearly matching C12Q's 58.6% — the contest has moved as much into software as into chemistry.
Filing growth compares 2021 (35 records) with 2024 (37) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 660 records in scope (CR5), not by the ranked leaders only.
What this landscape covers
This dataset tracks 660 patent records filed between 2015 and mid-2026 that combine basecalling, sequencing data processing or variant-calling pipeline claims with concerns like signal-to-sequence conversion, compute cost, quality score calibration, reference bias, pipeline versioning or reproducibility. It sits at the intersection of sequencing instrumentation and the software stack that turns raw signal into usable genomic calls, which is why classifications split heavily between C12Q molecular assay claims and G16B bioinformatics claims rather than sitting cleanly in one camp.
Because publication lags filing by roughly 18 months, the 2025 and 2026 figures in any trend view are still filling in and should not be read as a slowdown. The more reliable signal is the 2021-to-2024 window, where filings moved from 35 to 37 records — a modest but real increase in a field that already has a concentrated leadership group.
Let an AI agent run this analysis on your own technology
Pick a task. Every answer cites the patents behind it.
Filing trends and technology composition
Two views of the same 660 records: how filing activity has moved year over year, and which IPC subclasses carry the claims.
Filing activity, 2017–2026
Filings ran from 81 in 2017 to a peak of 96 in 2022, before falling off in the most recent years shown — a pattern consistent with publication lag rather than reduced inventive activity. The comparable 2021-to-2024 window shows a +6% increase, from 35 to 37 records.
IPC subclass composition
C12Q (measuring and testing involving enzymes or DNA) appears on 58.6% of the 660 records, and G16B (bioinformatics) on 42.7% — the two largest classes by a wide margin. G01N, C12N and G06F each cover a meaningful minority, and A61K, A61P and C07K appear on smaller shares, marking this chiefly as an assay-plus-software field rather than a therapeutics one. Records can carry more than one class, so these shares sum to well over 100%.
Shares are the percentage of the 660 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Basecalling and Sequencing Data Processing with Eureka
This page is one run against one query. Ask Eureka your own question about basecalling and sequencing data processing and every answer comes back with the patent numbers behind it.
Try EurekaMost-cited records and a representative claim
WO2023009758A1 — Quality score calibration of basecalling systems (Illumina, Inc., 2023)
A method of generating base calls by a base caller is disclosed. The method includes receiving a plurality of sensor data from a flow cell, identifying a second range such that at least a threshold percentage of the sensor data falls within it, mapping a subset of that data into a third range to produce normalized sensor data, and processing the normalized data in a base caller to call bases.Filed by Illumina, Inc. — illustrates how calibration of raw sensor data ahead of the base-calling step itself has become a patentable chokepoint, separate from the base-calling algorithm.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20100331194A1 | Nanopore sequencing devices and methods | 442 |
| 2 | US6223127B1 | Polymorphism detection utilizing clustering analysis | 337 |
| 3 | US20130157870A1 | Methods for obtaining a sequence | 320 |
| 4 | WO2012142611A2 | Sequencing small amounts of complex nucleic acids | 315 |
| 5 | US20130124100A1 | Processing and Analysis of Complex Nucleic Acid Sequence Data | 234 |
| 6 | WO2013036929A1 | Methods for obtaining a sequence | 226 |
| 7 | US6017434A | Apparatus and method for the generation, separation, detection, and recognition of biopolymer fragments | 187 |
| 8 | WO2013188831A1 | Uniquely tagged rearranged adaptive immune receptor genes in a complex gene set | 163 |
| 9 | WO2013086464A1 | Markers associated with chronic lymphocytic leukemia prognosis and progression | 151 |
| 10 | US20140322716A1 | Uniquely Tagged Rearranged Adaptive Immune Receptor Genes in a Complex Gene Set | 150 |
Citation counts favour older records simply because they have had more time to accumulate citations inside the searched corpus — treat this table as a map of influence, not a ranking of current importance.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the data means for filing strategy
Three patterns stand out once concentration, growth and class composition are read together.
The leadership group is narrow and stable
The leading assignee alone accounts for 180 records, with the fifth-place holder at 21 — a steep drop-off that marks a small group of sequencing platform makers as the entities whose claims any new filer must design around first.
Growth is real but unspectacular
Filings moved from 35 in 2021 to 37 in 2024, the last years comparable without publication-lag distortion. That is expansion, not a surge — a field being built on steadily rather than one attracting a rush of new entrants.
Software claims are catching up to assay claims
C12Q still leads as the largest class, but G16B bioinformatics classifications now sit close behind at 42.7% of records — a sign that pipeline, calibration and computational claims are being filed nearly as heavily as the underlying molecular methods.
Cross-filing is limited but concentrated
Only ten co-assignee pairs appear across the dataset, and the strongest of them link a single sequencing company with its own software subsidiary or link a diagnostics company to research-institute partners — collaboration here is closer than it is broad.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to basecalling and sequencing data processing, with the prior art for and against each one.
Who is filing, and who has slowed
The ranked leaders are dominated by sequencing instrument makers and their bioinformatics arms, but recent-year momentum tells a different story than the cumulative totals do.
A single leader holds a wide lead
The top-ranked assignee's cumulative total is more than eight times the fifth-place holder's 21 records, reflecting a long history of filing across both instrumentation and downstream data processing claims.
Cumulative leaders are quiet in the latest year
Several of the largest historical filers show zero filings in the most recent year with year-over-year drops recorded at -100%, which is consistent with publication lag rather than withdrawal from the field — but it does mean the visible pipeline from these firms looks thinner than their portfolios suggest.
A broad base of single- and few-filing entrants
Beyond the concentrated top group, the ranking runs to 100 companies, many holding only a handful of records — research institutes, diagnostics firms and smaller instrument makers each staking a narrow claim rather than building a broad portfolio.
| Assignee | Recent year | YoY |
|---|---|---|
| Ares Trading SA | 1 | — |
| Illumina, Inc. | 0 | -100% |
| Edico Genome Corp | 0 | -100% |
| Adaptive Biotechnologies Corp | 0 | — |
| Oxford Nanopore Technologies Ltd | 0 | -100% |
| CeMM Forschungszentrum für Molekulare Medizin GmbH | 0 | — |
| Cancer Research Technology Ltd | 0 | -100% |
| Pacific Biosciences of California, Inc. | 0 | — |
Where to take this analysis
The dataset points to specific next steps depending on whether the goal is freedom-to-operate, portfolio positioning or identifying open claim space.
Map claims against the leading five assignees
With 44.4% of records held by five companies, any new filing in basecalling or calibration should first be checked against their specific claim language, not just their existence in the field.
Explore assignee claims in EurekaWatch the bioinformatics class for new entrants
G16B's climb to 42.7% of records suggests software-side filers may be growing relative to instrument makers — worth tracking separately from the hardware-heavy C12Q class.
Track G16B filings in EurekaTest the under-claimed branches for freedom-to-operate
Reference-bias correction and pipeline versioning show thinner claim density than core basecalling — a targeted search can confirm whether a first-mover position is still open.
Run a white-space search in EurekaCommon questions on this landscape
One assignee holds 180 of the 660 records in scope, well ahead of the fifth-place holder at 21. The top five assignees together account for 293 records, or 44.4% of all records in scope, meaning nearly half of the field's patent activity sits with a small group of sequencing platform and bioinformatics companies. Any new entrant should expect to design claims around this group's existing positions rather than around the field in general.
Filings grew modestly from 35 records in 2021 to 37 in 2024, a +6% increase over that span, which is the most reliable comparison window available. Numbers for 2025 and 2026 appear lower, but that reflects publication lag of roughly 18 months rather than an actual drop in filing activity. Treat any apparent decline in the final one to two years of a trend chart with caution.
C12Q covers measuring and testing methods involving enzymes or DNA and appears on 58.6% of the 660 records, making it the largest single class — largely the molecular and assay side of sequencing. G16B covers bioinformatics specifically and appears on 42.7% of records, reflecting the computational pipeline, calibration and variant-calling side. Many records carry both classifications, which is why the two shares add to more than 100% rather than splitting the field into separate halves.
Based on the technology composition, branches like reference-bias correction, pipeline versioning for reproducibility, and cross-platform variant call harmonization show up less densely than core basecalling and quality-score claims. These are process-adjacent problems rather than core signal-to-sequence conversion, which is why they may carry lighter claim density even as the core field grows. A targeted freedom-to-operate search is the only reliable way to confirm an opening before filing there.
Not necessarily. The most-cited records in this dataset, such as those describing nanopore sequencing devices and clustering-based polymorphism detection, are older filings that have simply had more years to accumulate citations within the searched corpus. Citation count is a better signal of historical influence on how the field developed than of which claims matter most to a filing decision today. Recent, lower-cited filings can carry equally binding claims for current freedom-to-operate purposes.
Research Basecalling and Sequencing Data Processing in depth with Eureka
Go past this page: query the whole basecalling and sequencing data processing corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.