Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →Filing growth compares 2021 (859 records) with 2024 (5,329) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 31,128 records in scope (CR5), not by the ranked leaders only.
The dataset spans 31,128 published records filed against large language model and related speech/language-representation claims between 2015 and the 2026 cut-off. Filing activity was modest through the late 2010s and accelerated sharply from 2021 onward, tracking the point at which transformer-based systems moved from research demonstrations into product roadmaps. Publication lags filing by roughly 18 months, so the most recent one to two years in any chart will always look thinner than the underlying filing activity actually was.
Coverage is broad rather than narrow: claims range from core model training and inference mechanics through speech recognition, image/video recognition and business-process applications. That breadth means the field is less a single contest over one claim type and more a set of adjacent fights — training methodology, inference efficiency, multimodal grounding, and domain-specific deployment — each with its own density and its own gaps.
Pick a task. Every answer cites the patents behind it.
Three views of the same 31,128-record dataset: how filing volume moved over time, how records classify across the IPC subclasses that govern the technology, and where those filings actually land jurisdictionally.
Annual filings moved from 344 in 2017 to a peak of 5,329 in 2024, a +520% rise across just 2021 to 2024. The 2025 and 2026 figures sit far lower, but that reflects publication lag, not a real slowdown — treat the last one to two years as still filling in rather than as evidence of a falling trend.
G06F (electric digital data processing) touches 31.9% of records and G06N (AI-model computing) 20.1%, together anchoring the core technical claims. G10L (speech and audio) covers 18.0%, reflecting how much of the corpus traces back to speech-interface applications. Business-process (G06Q), transmission (H04L) and healthcare informatics (G16H) each sit under 6%, marking them as narrower, less contested territory relative to the core.
Shares are the percentage of the 31,128 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
This page is one run against one query. Ask Eureka your own question about large language models patent landscape and every answer comes back with the patent numbers behind it.
Try EurekaA computer-implemented method, computer program product and system for tuning large language models: pairs of textual prompts and ground-truth labels are received, a data selection scoring function is created by repurposing one or more reward functions to compute similarity between prompts and labels, a training dataset is selected from those pairs using that scoring function, and the large language model is tuned accordingly.Filed by International Business Machines Corporation, published 2025-07-17.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20020135618A1 | System and method for multi-modal focus detection, referential ambiguity resolution and mood classification u… | 881 |
| 2 | US20060149558A1 | Synchronized pattern recognition source data processed by manual or automatic means for creation of shared sp… | 827 |
| 3 | US6070140A | Speech recognizer | 781 |
| 4 | US20110060587A1 | Command and control utilizing ancillary information in a mobile voice-to-speech application | 770 |
| 5 | US5477451A | Method and system for natural language translation | 754 |
| 6 | US7702508B2 | System and method for natural language processing of query answers | 720 |
| 7 | US20210090694A1 | Data based cancer research and treatment systems and methods | 712 |
| 8 | US20020184373A1 | Conversational networking via transport, coding and control conversational protocols | 696 |
| 9 | US20090144609A1 | NLP-based entity recognition and disambiguation | 691 |
| 10 | US6964023B2 | System and method for multi-modal focus detection, referential ambiguity resolution and mood classification u… | 663 |
Citation counts favour older filings simply because they have had longer to accumulate citations inside this searched corpus — read them as a signal of influence on later work, not as a ranking of current technical importance.
Each row carries its publication number; clicking a row searches Eureka by that number.
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →Read together, the ranking, the technology split and the citation table point to a field with a real leadership cluster, a long tail of single-filing entrants, and prior art that is older than the current wave of transformer-specific claims.
The top 5 assignees combine for 6,359 of 31,128 records in scope, and the top 10 add up to 27.3%. That leaves nearly three-quarters of the corpus spread across a long tail of filers, including single-filing entrants — the field has recognisable leaders but no single gatekeeper.
Filings rose from 859 in 2021 to a peak of 5,329 in 2024. Because publication lags filing by around 18 months, 2025 and 2026 figures understate real activity and should not be read as a slowdown.
G06F and G06N together anchor the corpus's technical core, while G10L shows how much of the field still runs through speech-interface applications rather than pure language modelling. G16H, healthcare informatics, covers only 3.4% of records — a narrow slice relative to the overall field.
The most-cited records in this corpus are speech recognition and multi-modal input filings predating the current LLM wave, not recent transformer-specific claims. That is expected in a searched corpus — older filings have had more time to accumulate citations — but it means citation rank is a poor proxy for which claims matter to a filing decision today.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to large language models patent landscape, with the prior art for and against each one.
The ranked leaders combine large diversified technology filers with speech-and-communications specialists. Recent-year momentum figures for several leaders show sharp year-on-year declines, but that pattern is consistent with publication lag rather than a real pullback from the field.
The top-ranked assignee holds 1,947 records against a fifth-place figure of 726 and a tenth-place figure of 293 — a steep drop-off after the leader that flattens quickly through the rest of the ranked set.
Several of the largest filers show latest-year counts far below prior years, with year-on-year declines in the -80% to -93% range. Given the roughly 18-month publication lag, these are provisional figures still filling in rather than evidence that leaders are exiting the field.
Only 10 co-assignee pairs appear in the dataset, and the strongest links pair a large filer with either an individual inventor-assignee or an affiliated subsidiary rather than with an unrelated competitor. Cross-company joint filing is not yet a notable pattern in this field.
| Assignee | Recent year | YoY |
|---|---|---|
| Google LLC | 22 | -85% |
| NVIDIA Corporation | 15 | -91% |
| Microsoft Technology Licensing, LLC | 8 | -93% |
| Samsung Electronics Co., Ltd. | 7 | -88% |
| Amazon Technologies, Inc. | 5 | -81% |
| Oracle International Corporation | 5 | -93% |
| Qualcomm Incorporated | 2 | -97% |
| International Business Machines Corporation | 0 | -100% |
The dataset points to a field with clear leaders but wide open claim space beyond the core. Turning that into a filing or freedom-to-operate decision means going deeper on the specific sub-areas that matter to your roadmap.
Run your draft claims against the training-dataset, model-parameter and inference-mechanism language that defines this field's densest classes before you file, to see where you sit relative to existing art.
Explore in EurekaHealthcare informatics and business-process integration are thin today but will not stay that way if filing volume keeps climbing at anything like the 2021-2024 rate.
Set up monitoring in EurekaWith 27.3% of records held by the top 10 assignees, a targeted clearance search against those specific portfolios is more useful than a broad keyword sweep.
Run a clearance search in EurekaThe ranked dataset shows a clear leader holding 1,947 records, well ahead of the fifth-ranked assignee at 726 and the tenth-ranked at 293. The top 5 assignees combined hold 6,359 records, 20.4% of the 31,128 records in scope, and the top 10 add up to 27.3%. That leaves the majority of filings spread across a long tail of smaller and single-filing entities, so no one company controls the field outright.
Filing volume rose from 859 records in 2021 to a peak of 5,329 in 2024, a growth of 520% over that three-year span. 2024 is the most recent year that can be read as a complete annual figure. Counts for 2025 and 2026 appear much lower in the raw data, but that reflects the roughly 18-month lag between filing and publication rather than an actual slowdown in filing activity.
The core technical claims cluster in G06F (electric digital data processing, 31.9% of records) and G06N (computing based on AI models, 20.1%), with G10L (speech and audio analysis/synthesis) covering a further 18.0%. Narrower application classes such as G06Q (business processes), H04L (digital transmission), G06T (image processing) and G16H (healthcare informatics) each cover under 6% of records. Because a single record can carry multiple IPC classes, these shares add up to more than 100% and should not be summed into a single total.
The technology composition data shows healthcare informatics (G16H, 3.4% of records) and business-process integration (G06Q, 5.8%) are far less densely claimed than the core data-processing and AI-model classes. Co-assignee data also shows only 10 identified collaboration pairs across the whole dataset, suggesting cross-company joint development claims remain rare. Both patterns point to sub-areas where claim space is comparatively open relative to the core training and inference mechanics that the largest filers already occupy densely.
The most-cited records in this corpus, including several speech-recognition filings from the 1990s and 2000s, predate the current transformer-based wave of LLM patents by years or decades. That is a structural feature of citation counting inside any searched corpus: older filings have simply had more time to accumulate citations from later work. A high citation count signals historical influence on the field, not that a claim is more technically important today than a recent, less-cited filing.
Go past this page: query the whole large language models patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company's registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.