Book a demo

Speculative Decoding Patents: Who Leads, Trends & Gaps 2026

Speculative Decoding Patents: Who Leads, Trends & Gaps 2026
https://www.patsnap.com/resources/blog/rd-blog/speculative-decoding-methods-patent-landscape/ · Patsnap · data cut-off 2026-07-31 · downloaded from the live page
Patent Landscape · AI Inference Serving
Speculative decoding patents: mapping who owns draft-and-verify inference acceleration
  • 68.0% concentration. The top 5 of 80 ranked assignees hold 100 of the 147 records in scope — a field with a narrow, well-defended core.
  • +2900% in three years. Filings rose from 2 in 2021 to 60 in 2024, the fastest build-up phase in this dataset before publication lag sets in.
  • G06N and G06F dominate. 70.1% and 67.3% of records respectively sit in AI-model and digital-data-processing classes, leaving transmission and speech-specific claims comparatively thin.
Get a prior-art report on your approach
147
Published Records
68%
Top-5 Share of All Records
+2900%
Filing Growth 2021→2024
US
Leading Jurisdiction

Filing growth compares 2021 (2 records) with 2024 (60) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 147 records in scope (CR5), not by the ranked leaders only.

Published byPatsnap Research··6 min readSourced from Patsnap Eureka
Overview

What speculative decoding patents actually cover

Speculative decoding accelerates autoregressive generation by having a smaller draft model propose several tokens ahead, then having the target model verify and accept or reject them in a single pass. The patent activity in this dataset centres on the mechanics that make that trade-off work in production: acceptance-rate tuning, draft length selection, tree-based drafting, distribution preservation, and the memory overhead of running two models side by side. These are the claim elements that determine whether a speculative-decoding implementation is fast enough, cheap enough, and faithful enough to the original model’s output distribution to ship.

The 147 records in scope span 2015 through the first partial year of 2026, but the practical history is much shorter: filing activity was negligible before 2021 and only became dense from 2023 onward, tracking the broader commercial push toward large-language-model inference serving.

Filing activity and technology composition, 2017–2026
  1. 1QUALCOMM INC76
  2. 2UNIFABRIX LTD10
  3. 3GOOGLE LLC5
  4. 4PITNEY BOWERS INC5
  5. 5PALO ALTO NETWORKS INC4
  6. 6GDM HOLDING LLC3
  7. 7Beijing SiliconFlow Technology Co., Ltd.3
  8. 8SHANDONG INSPUR SCI RES INST CO LTD2
  9. 9SANG XIAOJUN2
  10. 10GEHLHAAR JEFFREY BAGINSKY2
Source: Patsnap Eureka. Assignee ranking and totals. Derived from a Patsnap search on Speculative Decoding Methods covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Let an AI agent run this analysis on your own technology

Pick a task. Every answer cites the patents behind it.

10,000 free credits to start
The Data

Filing trends and technology composition

The trend line and IPC breakdown below are drawn from the same 147-record corpus used throughout this page. Because publication typically lags filing by around 18 months, the most recent one or two years will fill in further as pending applications publish.

From near-zero to a 2024 peak

Filings held at 2 in 2021 and climbed to 60 in 2024 — a +2900% rise over that span — before the 2025-2026 figures, still incomplete, show the expected drop-off from publication lag rather than a genuine slowdown.

From near-zero to a 2024 peak015304560020172018201920202021202220236020242025132026Most recent year is partial — publication lag means later filings are not yet visible.

Concentrated in two AI/data-processing classes

G06N (AI models) and G06F (digital data processing) each cover roughly two-thirds of all records, confirming that most claims are being written as computing and model-architecture inventions rather than as communications or speech-specific ones; H04L, G07B, G10L, G06Q, G06K and H04W each account for a small minority of records.

Concentrated in two AI/data-processing classesG06N · Computing based on AI models10370.1%G06F · Electric digital data processi…9967.3%H04L · Digital information transmissi…117.5%G07B · Ticket & fare-registering53.4%G10L · Speech & audio analysis/synthe…32.0%G06Q · Business, commerce & admin dat…21.4%G06K · Data recognition & presentation10.7%H04W · Wireless communication networks10.7%

Shares are the percentage of the 147 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.

Source: Patsnap Eureka. Filing trend and technology composition. Derived from a Patsnap search on Speculative Decoding Methods covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.

Go deeper on Speculative Decoding Methods with Eureka

This page is one run against one query. Ask Eureka your own question about speculative decoding methods and every answer comes back with the patent numbers behind it.

Try Eureka
Key Patents

Representative and most-cited filings

Representative filing
US20250231989A12025-07-17

Speculative decoding in autoregressive generative artificial intelligence models (US20250231989A1)

QUALCOMM INCORPORATED

The application describes receiving multiple candidate token sets generated from a first generative model, then using a second generative model together with recursive adjustment of a target distribution to select among those candidates before outputting the chosen set — a distribution-preserving verification step applied to draft-and-verify inference.Filed by Qualcomm; published 2025-07-17.

US20250231989A1 — patent drawing 1US20250231989A1 — patent drawing 2
View full filing
Most-cited records in this corpus
#Publication no.Patent titleCitations
1US6430543B1Controlled acceptance mail fraud detection system54
2US20220116411A1Deobfuscating and decloaking web-based malware with abstract execution40
3US20130159724A1Method And Apparatus For A Scalable And Secure Transport Protocol For Sensor Data Collection38
4US20240320433A1Speculative decoding in autoregressive generative artificial intelligence models18
5US8935533B2Method and apparatus for a scalable and secure transport protocol for sensor data collection15
6US20250021761A1Accelerating inferencing in generative artificial intelligence models14
7US20240354346A1Speculative decoding in autoregressive generative artificial intelligence models12
8WO2024118603A1Methods and systems for fast inference from machine learning models12
9CN119761316A投机解码优化方法、电子设备和存储介质10
10US20240320529A1Multi-stage watermarking of a digital object generated by a machine learning model9

Citation counts reflect influence within the searched corpus and skew toward older records; they are not a measure of current commercial relevance.

Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.

Source: Patsnap Eureka. Citation counts and representative records. Derived from a Patsnap search on Speculative Decoding Methods covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Run it yourself

Put your own technology through the same analysis

 
Where to run it
Fastest

Eureka on the web

When you want the answer in the next five minutes.

The agent works the prompt against patents and technical literature, citing every source.

Run your analysis now →
For builders

MCP server & REST API

When it has to run inside your own pipeline.

Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.

Browse MCP servers →
Insights

What the numbers mean for a filing decision

Three patterns stand out once the assignee ranking, technology composition and filing trend are read together.

Concentration
68.0% / top 5
share of 147 records

A narrow core holds most of the claim space

Five assignees out of 80 ranked account for 100 of the 147 records in scope. That level of concentration means new entrants are more likely to find open ground in adjacent implementation details than in the core draft-verify mechanism itself.

Top 10 combined reach 76.2% of all records.
Momentum
2 → 60
filings, 2021 to 2024

The build-up phase is recent and steep

The +2900% rise from 2021 to 2024 marks this as a technology still in its early commercial consolidation, not a mature field with settled claim boundaries. Expect continued filing once the 2025-2026 figures finish publishing.

2024 is the last year treated as complete for trend purposes.
Technology mix
70.1% G06N
of 147 records

Claims cluster in model architecture, not transmission

G06N and G06F together dominate the IPC mix, while H04L, G10L and H04W each cover a small slice of records. Filers treating speculative decoding as a networking or speech-pipeline problem, rather than a model-architecture one, are working in comparatively open territory.

Classes overlap; shares are given against the 147-record total, not against each other.
Eureka AI Agent
Looking for what nobody has claimed yet?

Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to speculative decoding methods, with the prior art for and against each one.

Find the white space →
Source: Patsnap Eureka. Co-assignee relationships and derived observations. Derived from a Patsnap search on Speculative Decoding Methods covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Players

Who is filing, and where activity is slowing

The assignee ranking runs 80 companies deep, counted in records, with a steep drop from the leader to the rest of the field.

Leader
76 records
leading assignee

One filer sets the pace

The leading assignee's record count is nearly twenty times the fifth-place total of 4, an unusually steep drop-off even for a concentrated field.

Fifth place holds 4 records; tenth place holds 2.
Momentum
-85% YoY
leading assignee, latest year

Recent activity has cooled from its peak

The leading assignee filed 3 records in the latest year, down 85% year-on-year, while several other named assignees show zero filings in the latest year. Given publication lag, this reads as incomplete recent data rather than confirmed retreat.

Momentum figures use the most recent available filing year, which is still partial.
Collaboration
10 pairs
co-assignee pairs

Co-filing is limited and concentrated

Only 10 co-assignee pairs appear in the dataset, and the strongest links all involve the same leading assignee paired with different named individuals — a pattern of internal inventor credit rather than cross-company joint development.

No pair in this dataset exceeds 2 shared records.
🔍
Under-claimed sub-areas worth checking before filing
Based on the technology composition, these branches carry comparatively few records relative to the core G06N/G06F claims.
Tree-based draft token verification over wireless linksSpeech-synthesis-specific draft models (G10L)Ticket/fare-system inference acceleration (G07B)Memory-overhead reduction for dual-model servingBusiness-process applications of draft-verify (G06Q)
Rank all filers by momentum →
Recent-year filing momentum by assignee
AssigneeRecent yearYoY
Qualcomm Incorporated3-85%
Unifabrix Ltd.0-100%
Google LLC0-100%
Pitney Bowes Inc.0
Palo Alto Networks Inc.0
DeepMind Technologies Limited0
Beijing SiliconFlow Technology Co., Ltd.0-100%
Advanced Micro Devices, Inc. (AMD)0-100%
Source: Patsnap Eureka. Assignee-level momentum. Derived from a Patsnap search on Speculative Decoding Methods covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
What's Next

Where to take this next

The dataset points to specific next steps for teams deciding where to file or where to watch.

Check the white space chips against your own claims

The under-claimed sub-areas listed above are drawn directly from the IPC composition gaps in this corpus, not from guesswork. Run a candidate claim against each before assuming it is open.

Explore in Eureka

Track the leading assignee's next filings

With momentum data still incomplete for the most recent year, the leading assignee's next published applications will be the clearest signal of whether the 2024 peak was cyclical or a permanent step up.

Set up monitoring in Eureka
Source: Patsnap Eureka. Forward-looking reading of the same dataset. Derived from a Patsnap search on Speculative Decoding Methods covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
FAQ

Common questions about speculative decoding patents

Answers are grounded in the same dataset. Derived from a Patsnap search on Speculative Decoding Methods covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Research Speculative Decoding Methods in depth with Eureka

Go past this page: query the whole speculative decoding methods corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.

Try Eureka

Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.

Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.

Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.

Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.

Help us improve this page

Found incorrect or outdated information? Let us know and we'll get it fixed.