Speculative Decoding Patents: Who Leads, Trends & Gaps 2026
- 68.0% concentration. The top 5 of 80 ranked assignees hold 100 of the 147 records in scope — a field with a narrow, well-defended core.
- +2900% in three years. Filings rose from 2 in 2021 to 60 in 2024, the fastest build-up phase in this dataset before publication lag sets in.
- G06N and G06F dominate. 70.1% and 67.3% of records respectively sit in AI-model and digital-data-processing classes, leaving transmission and speech-specific claims comparatively thin.
Filing growth compares 2021 (2 records) with 2024 (60) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 147 records in scope (CR5), not by the ranked leaders only.
What speculative decoding patents actually cover
Speculative decoding accelerates autoregressive generation by having a smaller draft model propose several tokens ahead, then having the target model verify and accept or reject them in a single pass. The patent activity in this dataset centres on the mechanics that make that trade-off work in production: acceptance-rate tuning, draft length selection, tree-based drafting, distribution preservation, and the memory overhead of running two models side by side. These are the claim elements that determine whether a speculative-decoding implementation is fast enough, cheap enough, and faithful enough to the original model’s output distribution to ship.
The 147 records in scope span 2015 through the first partial year of 2026, but the practical history is much shorter: filing activity was negligible before 2021 and only became dense from 2023 onward, tracking the broader commercial push toward large-language-model inference serving.
Let an AI agent run this analysis on your own technology
Pick a task. Every answer cites the patents behind it.
Filing trends and technology composition
The trend line and IPC breakdown below are drawn from the same 147-record corpus used throughout this page. Because publication typically lags filing by around 18 months, the most recent one or two years will fill in further as pending applications publish.
From near-zero to a 2024 peak
Filings held at 2 in 2021 and climbed to 60 in 2024 — a +2900% rise over that span — before the 2025-2026 figures, still incomplete, show the expected drop-off from publication lag rather than a genuine slowdown.
Concentrated in two AI/data-processing classes
G06N (AI models) and G06F (digital data processing) each cover roughly two-thirds of all records, confirming that most claims are being written as computing and model-architecture inventions rather than as communications or speech-specific ones; H04L, G07B, G10L, G06Q, G06K and H04W each account for a small minority of records.
Shares are the percentage of the 147 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Speculative Decoding Methods with Eureka
This page is one run against one query. Ask Eureka your own question about speculative decoding methods and every answer comes back with the patent numbers behind it.
Try EurekaRepresentative and most-cited filings
Speculative decoding in autoregressive generative artificial intelligence models (US20250231989A1)
The application describes receiving multiple candidate token sets generated from a first generative model, then using a second generative model together with recursive adjustment of a target distribution to select among those candidates before outputting the chosen set — a distribution-preserving verification step applied to draft-and-verify inference.Filed by Qualcomm; published 2025-07-17.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US6430543B1 | Controlled acceptance mail fraud detection system | 54 |
| 2 | US20220116411A1 | Deobfuscating and decloaking web-based malware with abstract execution | 40 |
| 3 | US20130159724A1 | Method And Apparatus For A Scalable And Secure Transport Protocol For Sensor Data Collection | 38 |
| 4 | US20240320433A1 | Speculative decoding in autoregressive generative artificial intelligence models | 18 |
| 5 | US8935533B2 | Method and apparatus for a scalable and secure transport protocol for sensor data collection | 15 |
| 6 | US20250021761A1 | Accelerating inferencing in generative artificial intelligence models | 14 |
| 7 | US20240354346A1 | Speculative decoding in autoregressive generative artificial intelligence models | 12 |
| 8 | WO2024118603A1 | Methods and systems for fast inference from machine learning models | 12 |
| 9 | CN119761316A | 投机解码优化方法、电子设备和存储介质 | 10 |
| 10 | US20240320529A1 | Multi-stage watermarking of a digital object generated by a machine learning model | 9 |
Citation counts reflect influence within the searched corpus and skew toward older records; they are not a measure of current commercial relevance.
Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the numbers mean for a filing decision
Three patterns stand out once the assignee ranking, technology composition and filing trend are read together.
A narrow core holds most of the claim space
Five assignees out of 80 ranked account for 100 of the 147 records in scope. That level of concentration means new entrants are more likely to find open ground in adjacent implementation details than in the core draft-verify mechanism itself.
The build-up phase is recent and steep
The +2900% rise from 2021 to 2024 marks this as a technology still in its early commercial consolidation, not a mature field with settled claim boundaries. Expect continued filing once the 2025-2026 figures finish publishing.
Claims cluster in model architecture, not transmission
G06N and G06F together dominate the IPC mix, while H04L, G10L and H04W each cover a small slice of records. Filers treating speculative decoding as a networking or speech-pipeline problem, rather than a model-architecture one, are working in comparatively open territory.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to speculative decoding methods, with the prior art for and against each one.
Who is filing, and where activity is slowing
The assignee ranking runs 80 companies deep, counted in records, with a steep drop from the leader to the rest of the field.
One filer sets the pace
The leading assignee's record count is nearly twenty times the fifth-place total of 4, an unusually steep drop-off even for a concentrated field.
Recent activity has cooled from its peak
The leading assignee filed 3 records in the latest year, down 85% year-on-year, while several other named assignees show zero filings in the latest year. Given publication lag, this reads as incomplete recent data rather than confirmed retreat.
Co-filing is limited and concentrated
Only 10 co-assignee pairs appear in the dataset, and the strongest links all involve the same leading assignee paired with different named individuals — a pattern of internal inventor credit rather than cross-company joint development.
| Assignee | Recent year | YoY |
|---|---|---|
| Qualcomm Incorporated | 3 | -85% |
| Unifabrix Ltd. | 0 | -100% |
| Google LLC | 0 | -100% |
| Pitney Bowes Inc. | 0 | — |
| Palo Alto Networks Inc. | 0 | — |
| DeepMind Technologies Limited | 0 | — |
| Beijing SiliconFlow Technology Co., Ltd. | 0 | -100% |
| Advanced Micro Devices, Inc. (AMD) | 0 | -100% |
Where to take this next
The dataset points to specific next steps for teams deciding where to file or where to watch.
Check the white space chips against your own claims
The under-claimed sub-areas listed above are drawn directly from the IPC composition gaps in this corpus, not from guesswork. Run a candidate claim against each before assuming it is open.
Explore in EurekaTrack the leading assignee's next filings
With momentum data still incomplete for the most recent year, the leading assignee's next published applications will be the clearest signal of whether the 2024 peak was cyclical or a permanent step up.
Set up monitoring in EurekaCommon questions about speculative decoding patents
Highly concentrated at the top: five assignees out of 80 ranked hold 100 of the 147 records in scope, or 68.0% of all records. The leading assignee alone holds 76 records, far ahead of the fifth-place holder at 4. This means the core draft-and-verify mechanism is claimed densely by a small group, while the remaining 75 ranked assignees each hold only a handful of records.
Filing was negligible through 2017-2020 and stayed at just 2 records in 2021. It then rose sharply to a peak of 60 records in 2024, a +2900% increase over that three-year span. 2025 and 2026 figures appear lower in the raw trend, but that reflects the roughly 18-month lag between filing and publication rather than an actual slowdown.
The large majority sit in IPC class G06N (AI-based computing), covering 70.1% of the 147 records, and G06F (electric digital data processing), covering 67.3%. Smaller shares touch digital transmission (H04L), speech and audio processing (G10L), and business-process or ticketing classes (G06Q, G07B), each in single-digit percentages of the total. Because records can carry multiple classes, these shares sum to more than 100%.
The application, filed by Qualcomm and published 2025-07-17, covers receiving multiple candidate token sets from a first generative model and using a second generative model with recursive adjustment of a target distribution to select among those candidates before output. In practice this is a distribution-preserving verification step layered onto draft-and-verify inference. Anyone implementing a similar acceptance-and-resampling mechanism in a commercial inference stack should review this filing's claim scope closely before shipping a comparable design.
There is meaningful white space outside the dominant G06N/G06F core. Sub-areas such as speech-synthesis-specific draft models, wireless-link-specific tree-based verification, and memory-overhead reduction for dual-model serving each show comparatively few records in this dataset. That does not guarantee those areas are unclaimed elsewhere, but relative to the 147-record corpus mapped here, they are far less crowded than the core acceptance-rate and draft-length mechanisms.
Research Speculative Decoding Methods in depth with Eureka
Go past this page: query the whole speculative decoding methods corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.