Machine Learning Patents: Who Leads, Where the Gaps Are 2026
- 14.7% of the field sits with a handful of filers. the top five assignees combined account for 100,848 of 685,824 records in scope, with a long tail of single- and low-filing entrants behind them.
- Filing has plateaued, not fallen. 2021-to-2024 volume moved from 2,871 to 2,948 records, a +3% span across the last three years that can be treated as complete before publication lag distorts the count.
- Core AI classification carries the density; adjacent classes stay thin. G06N holds 1.8% of all records while classes like G16H healthcare informatics and G06Q business-process applications sit at 0.3% each — claim space that is open, not settled.
Filing growth compares 2021 (2,871 records) with 2024 (2,948) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 685,824 records in scope (CR5), not by the ranked leaders only.
What the machine learning patent record actually shows
The search set spans 685,824 published records filed against machine learning models, training datasets, feature vectors and model inference from 2015 through the 2026 cut-off. Filing concentration is real but not extreme: the top 5 assignees hold 14.7% of all records in scope, and the top 10 hold 20.5% — meaning roughly four-fifths of the corpus is distributed across a long tail of companies filing far less densely than the leaders. Publication lags filing by roughly 18 months, so the most recent one or two years in any trend chart will always look lighter than the underlying filing activity actually was.
Reading the IPC composition alongside the assignee ranking gives a fuller picture than either alone. Dense classes like G06N and G06F show where claim language is already crowded; thinner classes such as G16H healthcare informatics or G06Q business-process applications show where the same underlying model-training and inference techniques have not yet been claimed as heavily against a specific vertical.
Filing trend and technology composition
Two views of the same 685,824-record corpus: filing volume by year, and the IPC subclasses that structure where those filings actually claim their inventions.
Filing volume, 2017–2026
Annual filings ran from 460 in 2017 to a peak of 3,499 in 2023. The 2021-to-2024 span, the last window unaffected by publication lag, moved from 2,871 to 2,948 records — a +3% increase rather than a decline. Counts for 2025 and 2026 are still filling in and should not be read as a slowdown.
Technology composition by IPC subclass
G06N (computing based on AI models) is the densest single class at 1.8% of all 685,824 records, followed by G06F electric digital data processing at 0.9%. Image and recognition classes (G06V, G06T, G06K) each sit near 0.3-0.4%, and application-specific classes like G16H healthcare informatics and G06Q business data processing trail at 0.3%. Because a record can carry multiple IPC codes, these shares add up to more than 100% and should be read against the full record total, not against each other.
Shares are the percentage of the 685,824 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Machine Learning Patent Landscape with Eureka
This page is one run against one query. Ask Eureka your own question about machine learning patent landscape and every answer comes back with the patent numbers behind it.
Try EurekaA representative filing and the most-cited prior art
US11948050B2 — Caching of machine learning model training parameters
The patent covers caching a parameter of a machine learning model during training with a given dataset, then reusing that cached parameter for a subsequent training run. Caching can occur after each of multiple training iterations, and a given cached iteration is identified using a key built from a hash of the training dataset and a hash of the model parameter.Assignee: EMC IP Holding Company LLC · Published 2024-04-02


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20150379430A1 | Efficient duplicate detection for machine learning data sets | 854 |
| 2 | US20180018590A1 | Distributed Machine Learning Systems, Apparatus, and Methods | 762 |
| 3 | US20190209022A1 | Wearable electronic device and system for tracking location and identifying changes in salient indicators of … | 625 |
| 4 | US20150379429A1 | Interactive interfaces for machine learning model evaluations | 596 |
| 5 | US20170032281A1 | System and Method to Facilitate Welding Software as a Service | 576 |
| 6 | US20190236598A1 | Systems, methods, and apparatuses for implementing machine learning models for smart contracts using distribu… | 397 |
| 7 | US20170124487A1 | Systems, methods, and apparatuses for implementing machine learning model training and deployment with a roll… | 397 |
| 8 | US20160358099A1 | Advanced analytical infrastructure for machine learning | 359 |
| 9 | US20180341248A1 | Real-time adaptive control of additive manufacturing processes using machine learning | 336 |
| 10 | US20200175352A1 | Structure defect detection using machine learning algorithms | 314 |
Citation counts reflect influence within this searched corpus and skew toward older filings; they are not a measure of current commercial importance.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the data means for filing and freedom-to-operate decisions
Three patterns worth acting on rather than just noting.
Leadership is real but not a chokehold
With the top five assignees holding 14.7% of all records and the top ten holding 20.5%, no single filer controls the field. A freedom-to-operate review needs to look past the leaders to the long tail, where most of the remaining volume actually sits.
Volume has held, not collapsed
The 2021-to-2024 window, the last stretch not distorted by publication lag, shows filings moving from 2,871 to 2,948 — modest but positive growth. Recent-year drops at the very top of the ranking reflect lag in the data, not a real pullback by those filers.
Core AI classification is dense; verticals are not
G06N carries the heaviest claim density in the corpus. Vertical applications such as healthcare informatics (G16H) and business-process data handling (G06Q) sit at roughly a sixth of that share, suggesting the underlying training and inference techniques are far more claimed than their applied use in those domains.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to machine learning patent landscape, with the prior art for and against each one.
Leading assignees and where activity is shifting
The ranking covers 100 companies returned by the data endpoint — not a curated top-50 or top-100 list, but the full ranked set for this search.
A single filer sets the pace by volume
The leading assignee's record count is more than double the fifth-place figure of 13,113, showing a steep drop-off even within the top tier before the ranking flattens into a longer tail.
Top-ranked filers show sharp recent-year drops — read with caution
Several of the most active historical assignees show large year-on-year declines in the latest year of data. Given the roughly 18-month publication lag, this is far more likely to reflect incomplete recent-year publication than an actual pullback in filing.
Joint filing is limited and concentrated
Only ten co-assignee pairs appear in the corpus, and the strongest of them link a single major filer to its own regional subsidiary or to individual named inventors, rather than reflecting broad industry collaboration.
| Assignee | Recent year | YoY |
|---|---|---|
| Capital One Services, LLC | 8 | -62% |
| Google LLC | 7 | -68% |
| Microsoft Technology Licensing, LLC | 2 | -89% |
| Oracle International Corporation | 2 | -91% |
| International Business Machines Corporation | 1 | -94% |
| Qualcomm Incorporated | 1 | -98% |
| Adobe Inc. | 1 | -75% |
| SAP SE | 1 | -86% |
Turning this landscape into a filing or FTO decision
The figures above describe where claim density sits today. Deciding what to do with that requires drilling into specific claim language, not just aggregate counts.
Check freedom-to-operate against the dense classes
G06N and G06F carry the heaviest claim density in this corpus. Before filing in core model-training or inference claims, run a targeted search against the leading assignees' active families in these classes.
Explore G06N filings in EurekaScope a first claim in the thinner verticals
Healthcare informatics and business-process applications of machine learning show markedly lower claim density than the core computing classes. That gap is where a narrowly drafted first claim is more likely to clear prior art.
Map white space in EurekaTrack momentum, not just rank
Recent-year YoY figures for top-ranked assignees are affected by publication lag and should not be read as a slowdown. Re-run momentum analysis once the current filing year is fully published.
Set up monitoring in EurekaCommon questions on the machine learning patent landscape
The assignee ranking for this corpus is led by a single company with 32,436 records, more than double the fifth-place figure of 13,113. The top five assignees combined hold 14.7% of all 685,824 records in scope, and the top ten hold 20.5%. That leaves roughly four-fifths of filings spread across a long tail of companies filing at much lower volume, so leadership at the top does not translate into control of the whole field.
Filing volume grew modestly rather than exploding or collapsing: the 2021-to-2024 window, the last stretch that can be treated as complete, moved from 2,871 to 2,948 records, a +3% increase. The peak year on record so far is 2023 at 3,499 filings. Figures for 2025 and 2026 look lower only because publication typically lags actual filing by around 18 months, not because filing activity has genuinely dropped.
G06N, the class for computing arrangements based on AI models, is the densest at 1.8% of all 685,824 records in scope, followed by G06F electric digital data processing at 0.9%. Image and recognition-related classes such as G06V, G06T and G06K each account for roughly 0.3-0.4% of records. Application-specific classes like G16H healthcare informatics and G06Q business-process data handling trail at around 0.3%, indicating those verticals are less densely claimed than the core computing classes.
The clearest gaps sit in applied verticals rather than the core algorithmic classes: healthcare informatics (G16H) and business-process applications (G06Q) each hold about a sixth of the claim density seen in G06N. Industrial and cross-modal applications, such as welding-process ML control or cross-modal image-to-text training pipelines, also show thinner coverage relative to the core classes. These are areas where the underlying training and inference techniques are well established but their applied claims in a specific vertical are not yet crowded.
US11948050B2, assigned to EMC IP Holding Company LLC and published 2024-04-02, covers caching a parameter of a machine learning model during training and reusing that cached parameter in a subsequent training run. The caching can occur after each of multiple training iterations, and a specific cached iteration is identified using a key derived from a hash of the training dataset and a hash of the model parameter. It is a training-infrastructure patent focused on parameter reuse and iteration identification, not on model architecture or a specific application domain.
Research Machine Learning Patent Landscape in depth with Eureka
Go past this page: query the whole machine learning patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.