Foundation Models Patents: Who Leads, Where the Gaps Are 2026
- Filings grew 265% from 2021 to 2024, climbing from 538 to 1,963 records before the most recent, still-incomplete years understate the true count.
- The top 10 assignees hold just 23.8% of all 6,757 records, leaving a long tail of single- and few-filing entrants across universities, integrators and component vendors.
- G06N and G06F together anchor the field, at 41.7% and 37.5% of records respectively, while healthcare informatics (G16H) sits at a comparatively open 9.8%.
Filing growth compares 2021 (538 records) with 2024 (1,963) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 6,757 records in scope (CR5), not by the ranked leaders only.
What the foundation models patent record actually shows
Foundation models — pretrained, general-purpose models adapted downstream via fine-tuning or inference — sit at the intersection of training data engineering and model architecture claims. The dataset behind this page spans 6,757 published records filed between 2015 and the 2026 cut-off, drawn from a search string that pairs foundation-model terminology with the practical mechanics of training: dataset construction, model parameters, feature vectors and inference. That pairing matters: a patent claiming a pretrained model in the abstract but not touching training data, parameters or inference in the claims falls outside this scope, so the corpus is weighted toward implementation-level filings rather than purely conceptual ones.
Reading the numbers requires one caveat throughout: publication lags filing by roughly 18 months, so 2025 and 2026 figures are still filling in and should not be read as a slowdown. Every trend, concentration and technology-share figure below uses the same 6,757-record denominator unless stated otherwise.
Let an AI agent run this analysis on your own technology
Pick a task. Every answer cites the patents behind it.
Filing trends and technology composition
Two views of the same 6,757 records: how filing activity has moved year over year, and which IPC subclasses the claims actually sit in.
A sharp filing ramp, then a data-lag dip
Annual filings rose from 26 records in 2017 to a peak of 1,963 in 2024 — a 265% increase between 2021 and 2024 alone. The 268 records recorded for 2026 and the lower 2025 count are not a real contraction; they reflect publication lag, not falling filing activity.
Claims concentrate in AI computing and data processing
G06N (AI-based computing) appears in 41.7% of records and G06F (digital data processing) in 37.5%, confirming that most filings claim either the model architecture or the data-handling pipeline around it. Image and video recognition (G06V, 16.5%) and image data processing (G06T, 15.4%) form a substantial secondary cluster, while healthcare informatics (G16H, 9.8%), business data processing (G06Q, 7.2%), speech and audio (G10L, 6.5%) and digital transmission (H04L, 6.2%) each cover a smaller but non-trivial share — note these add to more than 100% because records can carry multiple IPC classes.
Shares are the percentage of the 6,757 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Foundation Models Patent Landscape with Eureka
This page is one run against one query. Ask Eureka your own question about foundation models patent landscape and every answer comes back with the patent numbers behind it.
Try EurekaThe most-cited records and a representative recent filing
WO2025132949A1 — Dataset augmentation (Helsing GmbH, 2025-06-26)
The application claims a computer-implemented method for generating a training dataset: starting from an initial training dataset, applying a plurality of augmentation functions to produce augmented datasets, running a pre-trained ML model over each to generate feature vectors, and comparing those feature vectors against the feature vectors of the initial dataset.The comparison step against the original dataset's feature vectors is the operative limitation — it is what a designer would need to route around rather than the augmentation step itself.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20110137636A1 | Context aware back-transliteration and translation of names and common phrases using web resources | 377 |
| 2 | US20240386015A1 | Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search | 374 |
| 3 | US20200125956A1 | Application Development Platform and Software Development Kits that Provide Comprehensive Machine Learning Se… | 244 |
| 4 | US20180300576A1 | Semi-automatic labelling of datasets | 201 |
| 5 | US20170011279A1 | Latent embeddings for word images and their semantics | 200 |
| 6 | US20220124543A1 | Graph neural network and reinforcement learning techniques for connection management | 197 |
| 7 | US20180293988A1 | Method and system of speaker recognition using context aware confidence modeling | 191 |
| 8 | US20170357896A1 | Content embedding using deep metric learning algorithms | 179 |
| 9 | US20200202171A1 | Systems and methods for rapidly building, managing, and sharing machine learning models | 173 |
| 10 | US20240046318A1 | Social network with network-based rewards | 157 |
Citation counts favour older records simply because they have had more time to accumulate citations within this searched corpus — treat them as a signal of influence on the field, not of current technical importance.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the filing pattern signals
Three read-throughs from the trend, concentration and citation data that matter for a filing or freedom-to-operate decision.
The ramp is real, the tail end is not yet visible
Filings rose from 538 in 2021 to 1,963 in 2024. Because publication lags filing by about 18 months, the lower counts recorded for 2025 and 2026 reflect incomplete data, not a cooling field.
No single owner controls the field
The ten most active assignees account for 23.8% of all records, and the top five for 17.9%. The remaining three-quarters of filings sit with a long tail of single- and few-filing entrants, which keeps the field open to new entrants on most claim territory.
Two subclasses carry the bulk of the claim density
G06N and G06F together describe the great majority of filings, meaning core model-architecture and data-pipeline claim space is already dense. Adjacent subclasses like G16H and G10L carry far fewer records and are correspondingly less crowded.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to foundation models patent landscape, with the prior art for and against each one.
Who is filing, and where momentum is slowing
The ranked leaders span large platform vendors, enterprise software firms and component makers, but recent-year momentum figures show even the most active assignees pulling back from their peak filing years — consistent with publication lag rather than strategic withdrawal.
The top filer leads by a wide single margin
The leading assignee holds 381 records against a fifth-place count of 120 and a tenth-place count of 67 — a steep drop-off after the leader rather than a tight cluster at the top.
Even the most active filers show sharp latest-year declines
Every assignee tracked for recent-year momentum shows a steep year-over-year drop in the latest year — from -40% to -96% — which is the expected shape of the data-lag effect on very recent filings, not a real pullback in R&D investment.
Co-filing is limited and mostly intra-group
Only ten co-assignee pairs appear in the data, and the strongest pairs link a parent company with its own subsidiary or regional IP arm rather than independent organisations, suggesting cross-company joint filing is still rare in this field.
| Assignee | Recent year | YoY |
|---|---|---|
| Google LLC | 8 | -77% |
| Oracle International Corporation | 6 | -40% |
| Microsoft Technology Licensing, LLC | 5 | -79% |
| NVIDIA Corporation | 4 | -96% |
| Accenture Global Solutions Limited | 4 | -85% |
| International Business Machines Corporation | 2 | -93% |
| Robert Bosch GmbH (Germany) | 0 | -100% |
| Samsung Electronics Co., Ltd. (Korea) | 0 | -100% |
Where to take this analysis
The dataset points to three practical next steps depending on whether the goal is freedom-to-operate, white space identification, or tracking a specific competitor.
Check freedom-to-operate against the dense core
Before filing claims that touch G06N or G06F territory — model architecture or the data pipeline around it — run a focused clearance search, since these two subclasses already cover the large majority of records in scope.
Explore claim-level search in EurekaTest claim language against the under-claimed branches
Healthcare informatics, speech/audio and business-process applications of pretrained models show meaningfully lower filing density than the core, which is where narrower, well-drafted claims are more likely to issue cleanly.
Draft and stress-test claims in EurekaTrack the leading assignees' next filings
With filing concentrated among a small number of leaders but momentum figures distorted by publication lag, ongoing monitoring is more reliable than a single snapshot for catching a competitor's next move.
Set up assignee monitoring in EurekaCommon questions about foundation model patents
This dataset identifies 6,757 published records matching foundation-model terminology combined with training-data, parameter, feature-vector or inference language, filed between 2015 and the 2026 data cut-off. Annual filings rose from 26 in 2017 to a peak of 1,963 in 2024. The lower counts shown for 2025 and 2026 reflect publication lag of roughly 18 months rather than an actual decline in filing activity, so the true 2025-2026 totals will rise as more records publish.
The ranked leaders include large platform vendors, enterprise software companies and component makers, with the leading assignee holding 381 records compared with 120 at fifth place and 67 at tenth. The top 10 assignees together account for 23.8% of all 6,757 records, and the top 5 for 17.9%, which means the great majority of filings sit outside the leading group in a long tail of smaller filers. This is a moderately concentrated field, not one dominated by a single owner.
By IPC subclass, AI-based computing (G06N) appears in 41.7% of records and general digital data processing (G06F) in 37.5%, making these the two densest claim areas. Image and video recognition (G06V) and image data processing (G06T) form a secondary cluster at 16.5% and 15.4% respectively. Healthcare informatics, business data processing, speech/audio and digital transmission each cover a smaller single-digit-to-high-single-digit share, and because records can carry multiple IPC classes these shares sum to more than 100%.
Yes, particularly in subclasses with comparatively low filing density relative to the crowded core, such as healthcare informatics applications of pretrained models, speech and audio compression for edge inference, and business-process fine-tuning with audit or governance features. The core claim territory around model architecture and training-data pipelines (G06N and G06F) is dense and likely to carry significant prior art. Narrower claims targeting specific downstream applications of foundation models are more likely to find open space than broad architecture or training-method claims.
Filings grew 265% between 2021 and 2024, from 538 records to 1,963, which is the clearest complete-year growth signal in the dataset. Because publication typically lags filing by around 18 months, the 2025 and 2026 figures are still incomplete and should not be read as the field slowing down. Treat 2024 as the most recent year that gives a reliable picture of actual filing volume.
Research Foundation Models Patent Landscape in depth with Eureka
Go past this page: query the whole foundation models patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.