Speculative Decoding Patents: Who Leads, Where the Gaps Are 2026
- One assignee dominates the ranking. the leader holds 53 of the records tracked, against a single filing at fifth place and at tenth — a steep drop rather than a gradual tail.
- Filing only recently took off. activity was essentially zero through 2017 and peaked at 35 records in 2024, so the whole field's patent history is compressed into a few years.
- Momentum is cooling across the board. every assignee with prior-year volume shows a year-over-year drop in the most recent filing year, including the leader at -93%.
What speculative decoding patents actually cover
Speculative decoding speeds up autoregressive generation by having a smaller draft model propose several tokens ahead, which a larger target model then verifies in a single pass. The patent claims in this dataset cluster around three mechanics: how the draft proposes tokens, how the verifier accepts or rejects them without distorting the output distribution, and how the whole scheme behaves once it meets production batching. The search scope spans 62 published records from 2015 through the 2026 cut-off, concentrated almost entirely in the back half of that window.
Because publication lags filing by roughly 18 months, the 2025 and 2026 counts in the trend chart understate real filing activity for those years — the 2024 peak of 35 is the most reliable recent read on how fast the field was moving before the data cut-off.
Filing trend and technology composition
The record set is small enough that a handful of assignees and two IPC subclasses account for nearly all of it — read the shares below against the 62-record denominator, not against each other.
From zero to a 2024 peak
Filings were flat at zero in 2017 and climbed to a peak of 35 records in 2024, the last complete year before publication lag starts to understate the count. The 2026 figure of 2 is a partial year and should not be read as a slowdown on its own.
Two subclasses cover almost every record
G06N (computing arrangements based on AI models) appears on 80.6% of the 62 records and G06F (electric digital data processing) on 79.0%. Because records can carry both classes, these figures overlap rather than sum, and together they confirm that nearly every filing in scope treats the technology as a model-execution problem rather than, say, a hardware-accelerator one.
Shares are the percentage of the 62 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Speculative Decoding and Draft Models with Eureka
This page is one run against one query. Ask Eureka your own question about speculative decoding and draft models and every answer comes back with the patent numbers behind it.
Try EurekaThe most-cited records in the corpus
| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20240320433A1 | Speculative decoding in autoregressive generative artificial intelligence models | 18 |
| 2 | US20240354346A1 | Speculative decoding in autoregressive generative artificial intelligence models | 12 |
| 3 | US20250245430A1 | Efficient speculative decoding in autoregressive generative artificial intelligence models | 7 |
| 4 | US20240354345A1 | Speculative decoding in autoregressive generative artificial intelligence models | 5 |
| 5 | US20250245530A1 | Adaptive length speculative decoding in autoregressive generative artificial intelligence models | 4 |
| 6 | US12229192B2 | Speculative decoding in autoregressive generative artificial intelligence models | 4 |
| 7 | US20250231989A1 | Speculative decoding in autoregressive generative artificial intelligence models | 2 |
| 8 | US20260065048A1 | Self-speculative decoding using forecasted embeddings in autoregressive generative artificial intelligence mo… | 2 |
| 9 | US20260093960A1 | Large language model inferencing acceleration techniques | 2 |
| 10 | WO2024220144A1 | Speculative decoding in autoregressive generative artificial intelligence models | 2 |
Citation counts inside a searched corpus favour older publications, since they have had more time to accumulate citations — treat this as a signal of influence on the field's early claim language, not a ranking of current technical importance.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the numbers say about the field's shape
Three patterns stand out once the ranking, the trend and the filing offices are read together.
One assignee, then a cliff
The leading assignee holds 53 of the records in the ranking; fifth place holds only 1 and tenth place also holds 1. That is not a long tail so much as a single dominant filer surrounded by a scatter of one-off entrants, which changes how a freedom-to-operate review should be scoped — most of the claim risk sits with one party.
A field that arrived late and fast
Filing volume was at zero in 2017 and reached its recorded peak of 35 records in 2024. The compressed timeline means most of the prior art a new filer has to clear was written in the last two or three complete years, not built up gradually over a decade.
Every tracked assignee is filing less
The leading assignee's latest-year filings dropped 93% year over year, and every other assignee with prior activity shows a 100% year-over-year drop to zero. Only one newer entrant shows any latest-year filing at all. Read this alongside the publication-lag point: some of the apparent drop is filings not yet published rather than filings not made.
US and PCT dominate the filing offices
The United States receives the largest single share of filings at 16, with WIPO (PCT) close behind at 12; India, Europe, Israel and Singapore each sit in single digits. That split suggests most applicants are still deciding on downstream national coverage rather than having already committed to a broad multi-jurisdiction strategy.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to speculative decoding and draft models, with the prior art for and against each one.
Who is filing, and where the claim space is still open
The ranking is dominated by one filer, with the rest of the field made up of smaller corporate and individual applicants who show up once or twice each.
A single dominant filer
One assignee accounts for the large majority of the ranked records, with related-title continuation filings appearing among the most-cited documents in the corpus. Its latest-year filings dropped 93% year over year, which may reflect a maturing internal portfolio rather than reduced interest in the technology.
A thin scatter of single-digit filers
Below the leader, corporate names in the ranking include semiconductor and telecom-adjacent filers alongside a domestic research institute, each holding only a handful of records. Co-assignee pairings in the dataset show the leader collaborating with individual named inventors rather than with other companies.
One filer still adding records
Vellore Institute of Technology is the only assignee in the momentum data showing a filing in the latest year without a year-over-year decline, making it the one name in the ranking worth watching for continued activity as other filers pull back.
| Assignee | Recent year | YoY |
|---|---|---|
| Qualcomm Incorporated | 1 | -93% |
| VELLORE INSITUTE OF TECH | 1 | — |
| Advanced Micro Devices (AMD) | 0 | -100% |
| Huawei Technologies Co., Ltd. | 0 | -100% |
| Shandong Inspur Scientific Research Institute Co., Ltd. | 0 | -100% |
| Beijing Yijing Information Technology Co., Ltd. | 0 | -100% |
| ZHANG WENHAO | 0 | — |
| MARZOLLO MICHELE | 0 | — |
Where to take this analysis
The dataset points to a field with one dominant filer, a recent filing peak, and several adjacent claim areas that remain thin.
Scope a freedom-to-operate check around the leader
With 53 of the ranked records held by one assignee, a targeted FTO review of that portfolio's claim language on acceptance-rate and verification mechanics is likely more useful than a broad field-wide search.
Explore assignee claims in EurekaWatch the 2025-2026 publication window
Because publication lags filing by around 18 months, filings made in 2024-2025 are still arriving in the record. Re-running the trend query in a few months will sharpen the read on whether the 2024 peak was a high point or a plateau.
Track filing trends in EurekaDraft around the under-claimed branches
Multi-draft ensembling, batching-aware thresholds and hardware-aware draft sizing show little claim density in this corpus relative to the core verification mechanism, making them candidates for a first-mover filing.
Search white space in EurekaCommon questions on speculative decoding patents
Most claims in this corpus cover the interaction between a draft model that proposes multiple tokens and a verifier that accepts or rejects them in a single forward pass. Common claim elements include how the acceptance rate is computed, how the scheme preserves the original output distribution, and how draft length is adjusted adaptively. Several of the most-cited records in this dataset use near-identical titles, which suggests continuation-style filing around a core mechanism rather than isolated one-off inventions.
The assignee ranking behind this landscape is dominated by a single company holding 53 of the records tracked, with the next closest filers holding only one record each at fifth and tenth place. That is a steep concentration rather than a gradual tail, so most of the claim risk in this space sits with one portfolio. Smaller filers include semiconductor and telecom-adjacent companies and at least one research institute, each appearing only a handful of times.
Filing activity was effectively at zero as of 2017 and climbed to a recorded peak of 35 records in 2024, the last complete year before publication lag distorts the count. That means the bulk of the prior art in this field was written within the last two to three complete years rather than accumulated over a long period. Readers should treat the 2025 and 2026 figures as undercounts, since publication typically lags filing by around 18 months.
Yes, based on the IPC composition and the claim density around the leading assignee's portfolio, several adjacent branches remain thinly claimed, including multi-draft-model ensembling, batching-aware acceptance thresholds, and hardware-aware draft sizing. These sit next to the dense claim territory around basic draft-and-verify mechanics rather than inside it. A first claim in one of these branches would need to specify the exact interaction left unaddressed by the core verification patents, such as how acceptance thresholds shift under variable batch sizes.
Every assignee with prior-year filing activity in this dataset shows a year-over-year decline in the latest year, including a 93% drop for the leading filer and 100% drops for several others down to zero. Part of this is real: portfolios built around an initial technical wave often slow once the core claim space is staked out. But part of it is a data artefact, since publication lags filing by roughly 18 months and the most recent year in any trend is always undercounted. The one exception, Vellore Institute of Technology, still shows a filing in the latest year without a recorded decline.
Research Speculative Decoding and Draft Models in depth with Eureka
Go past this page: query the whole speculative decoding and draft models corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.