Book a demo

Speculative Decoding Patents: Who Leads, Where the Gaps Are 2026

Speculative Decoding Patents: Who Leads, Where the Gaps Are 2026
https://www.patsnap.com/resources/blog/rd-blog/speculative-decoding-and-draft-models-patent-landscape/ · Patsnap · data cut-off 2026-07-31 · downloaded from the live page
AI Inference Patent Landscape
Speculative decoding patents: mapping draft-model verification and wall-clock speedup claims
  • One assignee dominates the ranking. the leader holds 53 of the records tracked, against a single filing at fifth place and at tenth — a steep drop rather than a gradual tail.
  • Filing only recently took off. activity was essentially zero through 2017 and peaked at 35 records in 2024, so the whole field's patent history is compressed into a few years.
  • Momentum is cooling across the board. every assignee with prior-year volume shows a year-over-year drop in the most recent filing year, including the leader at -93%.
Get a prior-art report on your approach
62
Published Records
US
Leading Jurisdiction
10
Active Filers Ranked
2024
Peak Filing Year

Published byPatsnap Research··7 min readSourced from Patsnap Eureka
Field Overview

What speculative decoding patents actually cover

Speculative decoding speeds up autoregressive generation by having a smaller draft model propose several tokens ahead, which a larger target model then verifies in a single pass. The patent claims in this dataset cluster around three mechanics: how the draft proposes tokens, how the verifier accepts or rejects them without distorting the output distribution, and how the whole scheme behaves once it meets production batching. The search scope spans 62 published records from 2015 through the 2026 cut-off, concentrated almost entirely in the back half of that window.

Because publication lags filing by roughly 18 months, the 2025 and 2026 counts in the trend chart understate real filing activity for those years — the 2024 peak of 35 is the most reliable recent read on how fast the field was moving before the data cut-off.

Filing activity, 2017-2026
  1. 1Qualcomm Incorporated53
  2. 2Advanced Micro Devices (AMD)2
  3. 3Huawei Technologies Co., Ltd.2
  4. 4Shandong Inspur Scientific Research Institute Co., Ltd.1
  5. 5Beijing Yijing Information Technology Co., Ltd.1
  6. 6ZHANG WENHAO1
  7. 7VELLORE INSITUTE OF TECH1
  8. 8MARZOLLO MICHELE1
  9. 9LV SHAOCHUN1
  10. 10LI ZHIGUO1
Source: Patsnap Eureka. Assignee ranking and totals. Derived from a Patsnap search on Speculative Decoding and Draft Models covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Filing Data

Filing trend and technology composition

The record set is small enough that a handful of assignees and two IPC subclasses account for nearly all of it — read the shares below against the 62-record denominator, not against each other.

From zero to a 2024 peak

Filings were flat at zero in 2017 and climbed to a peak of 35 records in 2024, the last complete year before publication lag starts to understate the count. The 2026 figure of 2 is a partial year and should not be read as a slowdown on its own.

From zero to a 2024 peak01020304002017201820192020202120222023352024202522026Most recent year is partial — publication lag means later filings are not yet visible.

Two subclasses cover almost every record

G06N (computing arrangements based on AI models) appears on 80.6% of the 62 records and G06F (electric digital data processing) on 79.0%. Because records can carry both classes, these figures overlap rather than sum, and together they confirm that nearly every filing in scope treats the technology as a model-execution problem rather than, say, a hardware-accelerator one.

Two subclasses cover almost every recordG06N · Computing based on AI models5080.6%G06F · Electric digital data processi…4979.0%

Shares are the percentage of the 62 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.

Source: Patsnap Eureka. Filing trend and technology composition. Derived from a Patsnap search on Speculative Decoding and Draft Models covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.

Go deeper on Speculative Decoding and Draft Models with Eureka

This page is one run against one query. Ask Eureka your own question about speculative decoding and draft models and every answer comes back with the patent numbers behind it.

Try Eureka
Key Patents

The most-cited records in the corpus

Highest-cited published records
#Publication no.Patent titleCitations
1US20240320433A1Speculative decoding in autoregressive generative artificial intelligence models18
2US20240354346A1Speculative decoding in autoregressive generative artificial intelligence models12
3US20250245430A1Efficient speculative decoding in autoregressive generative artificial intelligence models7
4US20240354345A1Speculative decoding in autoregressive generative artificial intelligence models5
5US20250245530A1Adaptive length speculative decoding in autoregressive generative artificial intelligence models4
6US12229192B2Speculative decoding in autoregressive generative artificial intelligence models4
7US20250231989A1Speculative decoding in autoregressive generative artificial intelligence models2
8US20260065048A1Self-speculative decoding using forecasted embeddings in autoregressive generative artificial intelligence mo…2
9US20260093960A1Large language model inferencing acceleration techniques2
10WO2024220144A1Speculative decoding in autoregressive generative artificial intelligence models2

Citation counts inside a searched corpus favour older publications, since they have had more time to accumulate citations — treat this as a signal of influence on the field's early claim language, not a ranking of current technical importance.

Each row carries its publication number; clicking a row searches Eureka by that number.

Source: Patsnap Eureka. Citation counts and representative records. Derived from a Patsnap search on Speculative Decoding and Draft Models covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Run it yourself

Put your own technology through the same analysis

 
Where to run it
Fastest

Eureka on the web

When you want the answer in the next five minutes.

The agent works the prompt against patents and technical literature, citing every source.

Run your analysis now →
For builders

MCP server & REST API

When it has to run inside your own pipeline.

Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.

Browse MCP servers →
Data Insights

What the numbers say about the field's shape

Three patterns stand out once the ranking, the trend and the filing offices are read together.

Concentration
53 vs 1
leader vs fifth place, records

One assignee, then a cliff

The leading assignee holds 53 of the records in the ranking; fifth place holds only 1 and tenth place also holds 1. That is not a long tail so much as a single dominant filer surrounded by a scatter of one-off entrants, which changes how a freedom-to-operate review should be scoped — most of the claim risk sits with one party.

Assignee ranking, 10 companies
Timing
35 in 2024
peak-year filings

A field that arrived late and fast

Filing volume was at zero in 2017 and reached its recorded peak of 35 records in 2024. The compressed timeline means most of the prior art a new filer has to clear was written in the last two or three complete years, not built up gradually over a decade.

Filing trend, 2017-2026
Momentum
-93% YoY
leader's latest-year change

Every tracked assignee is filing less

The leading assignee's latest-year filings dropped 93% year over year, and every other assignee with prior activity shows a 100% year-over-year drop to zero. Only one newer entrant shows any latest-year filing at all. Read this alongside the publication-lag point: some of the apparent drop is filings not yet published rather than filings not made.

Recent-year momentum by assignee
Geography
16 US filings
of the receiving offices tracked

US and PCT dominate the filing offices

The United States receives the largest single share of filings at 16, with WIPO (PCT) close behind at 12; India, Europe, Israel and Singapore each sit in single digits. That split suggests most applicants are still deciding on downstream national coverage rather than having already committed to a broad multi-jurisdiction strategy.

Receiving offices
Eureka AI Agent
Looking for what nobody has claimed yet?

Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to speculative decoding and draft models, with the prior art for and against each one.

Find the white space →
Source: Patsnap Eureka. Co-assignee relationships and derived observations. Derived from a Patsnap search on Speculative Decoding and Draft Models covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Competitive Landscape

Who is filing, and where the claim space is still open

The ranking is dominated by one filer, with the rest of the field made up of smaller corporate and individual applicants who show up once or twice each.

Leader
53 records
of the ranked total

A single dominant filer

One assignee accounts for the large majority of the ranked records, with related-title continuation filings appearing among the most-cited documents in the corpus. Its latest-year filings dropped 93% year over year, which may reflect a maturing internal portfolio rather than reduced interest in the technology.

Assignee ranking
Mid-tier filers
1 record
at fifth and tenth place

A thin scatter of single-digit filers

Below the leader, corporate names in the ranking include semiconductor and telecom-adjacent filers alongside a domestic research institute, each holding only a handful of records. Co-assignee pairings in the dataset show the leader collaborating with individual named inventors rather than with other companies.

Co-assignee pairs, 7 total
New entrant
1 in latest year
only positive momentum

One filer still adding records

Vellore Institute of Technology is the only assignee in the momentum data showing a filing in the latest year without a year-over-year decline, making it the one name in the ranking worth watching for continued activity as other filers pull back.

Recent-year momentum
🔍
Under-claimed branches worth checking before filing
These sit adjacent to the dense claim territory around draft-model verification and acceptance-rate tuning.
multi-draft-model ensemblingbatching-aware acceptance thresholdsoutput-distribution fidelity proofshardware-aware draft sizingtree-structured speculative sampling
Rank all filers by momentum →
Momentum is slowing across every tracked assignee
AssigneeRecent yearYoY
Qualcomm Incorporated1-93%
VELLORE INSITUTE OF TECH1
Advanced Micro Devices (AMD)0-100%
Huawei Technologies Co., Ltd.0-100%
Shandong Inspur Scientific Research Institute Co., Ltd.0-100%
Beijing Yijing Information Technology Co., Ltd.0-100%
ZHANG WENHAO0
MARZOLLO MICHELE0
Source: Patsnap Eureka. Assignee-level momentum. Derived from a Patsnap search on Speculative Decoding and Draft Models covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
What's Next

Where to take this analysis

The dataset points to a field with one dominant filer, a recent filing peak, and several adjacent claim areas that remain thin.

Scope a freedom-to-operate check around the leader

With 53 of the ranked records held by one assignee, a targeted FTO review of that portfolio's claim language on acceptance-rate and verification mechanics is likely more useful than a broad field-wide search.

Explore assignee claims in Eureka

Watch the 2025-2026 publication window

Because publication lags filing by around 18 months, filings made in 2024-2025 are still arriving in the record. Re-running the trend query in a few months will sharpen the read on whether the 2024 peak was a high point or a plateau.

Track filing trends in Eureka

Draft around the under-claimed branches

Multi-draft ensembling, batching-aware thresholds and hardware-aware draft sizing show little claim density in this corpus relative to the core verification mechanism, making them candidates for a first-mover filing.

Search white space in Eureka
Source: Patsnap Eureka. Forward-looking reading of the same dataset. Derived from a Patsnap search on Speculative Decoding and Draft Models covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
FAQ

Common questions on speculative decoding patents

Answers are grounded in the same dataset. Derived from a Patsnap search on Speculative Decoding and Draft Models covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Research Speculative Decoding and Draft Models in depth with Eureka

Go past this page: query the whole speculative decoding and draft models corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.

Try Eureka

Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.

Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.

Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.

Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.

Help us improve this page

Found incorrect or outdated information? Let us know and we'll get it fixed.