Book a demo

Text-to-Speech Patents: Who Leads, Where the Gaps Are 2026

Text-to-Speech Patents: Who Leads, Where the Gaps Are 2026
https://www.patsnap.com/resources/blog/rd-blog/text-to-speech-and-voice-synthesis-patent-landscape/ · Patsnap · data cut-off 2026-07-31 · downloaded from the live page
Patent Landscape · Speech & NLP
Text-to-Speech and Voice Synthesis Patents: Filing Trends and Leading Assignees
  • Concentration is moderate, not extreme. the top 5 assignees hold 39.1% of the 87 records in scope, and the top 10 hold 59.8% — a leader, but a real field behind it.
  • Filing has already peaked. the trend hit 9 records in 2024 after a flat 2022 midpoint, and none of the most active recent assignees added a filing in the latest year.
  • Almost everything sits in one IPC subclass. 97.7% of records carry G10L, while G06F, G06N and H04M appear in far fewer — a sign that adjacent computing and telephony framing is still thin.
Get a prior-art report on your approach
87
Published Records
39%
Top-5 Share of All Records
+50%
3-Yr Growth (lag-adjusted)
US
Leading Jurisdiction
Published byPatsnap Research··8 min readSourced from Patsnap Eureka
Field Overview

What the patent record shows about text-to-speech and voice synthesis

Text-to-speech and voice synthesis patenting covers everything from waveform concatenation to neural vocoders, with claims typically anchored on prosody control, naturalness measured against MOS scores, and increasingly on voice cloning safeguards and multilingual synthesis. The 87 records in scope span 2015 to the 2026 data cut-off and were filtered against IPC classes covering speech and audio synthesis, AI computing models, and speech analysis specifically.

Filing activity is concentrated in a small set of long-standing speech-technology assignees alongside a longer tail of single- or few-filing entrants, and the receiving-office spread points to the United States and Europe as the primary filing venues, with India, WIPO and Japan behind them.

Filing activity by year, 2017-2026
  1. 1GOOGLE LLC11
  2. 2LERNOUT & HAUSPIE SPEECH PRODS8
  3. 3CERENCE OPERATING CO6
  4. 4CAMB AI INC5
  5. 5CANON KK4
  6. 6IND TECH RES INST4
  7. 7INTERNATIONAL BUSINESS MACHINE CORPORATION4
  8. 8NUANCE COMMUNICATIONS INC4
  9. 9SRC INC3
  10. 10NEC CORP3
Source: Patsnap Eureka. Assignee ranking and totals. Derived from a Patsnap search on Text-to-Speech and Voice Synthesis covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Let an AI agent run this analysis on your own technology

Pick a task. Every answer cites the patents behind it.

10,000 free credits to start
Data & Trends

Filing trend and technology composition

The filing curve and the IPC mix together describe a field that grew steadily into a 2024 peak and is now built almost entirely on one core classification, with only light spillover into adjacent computing and telephony subclasses.

A 2024 peak followed by an unfinished 2026

Filings moved from 0 in 2017 to a flat midpoint of 1 in 2022, then rose to a peak of 9 records in 2024. The 2026 figure of 1 is a partial year under the data cut-off and understates real activity, since publication typically lags filing by around 18 months.

A 2024 peak followed by an unfinished 20260358100201720182019202020212022202392024202512026Most recent year is partial — publication lag means later filings are not yet visible.

G10L dominates; adjacent classes are thin

G10L, speech and audio analysis and synthesis, appears on 97.7% of the 87 records in scope, confirming this is the primary classification for the field. G06F appears on 12.6%, G06N (AI computing models) on 4.6%, H04M (telephonic communication) on 3.4% and G11B (information storage) on 1.1% — each record can carry more than one class, so these shares add up to more than 100%.

G10L dominates; adjacent classes are thinG10L · Speech & audio analysis/synthe…8597.7%G06F · Electric digital data processi…1112.6%G06N · Computing based on AI models44.6%H04M · Telephonic communication33.4%G11B · Information storage (magnetic/…11.1%

Shares are the percentage of the 87 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.

Source: Patsnap Eureka. Filing trend and technology composition. Derived from a Patsnap search on Text-to-Speech and Voice Synthesis covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.

Go deeper on Text-to-Speech and Voice Synthesis with Eureka

This page is one run against one query. Ask Eureka your own question about text-to-speech and voice synthesis and every answer comes back with the patent numbers behind it.

Try Eureka
Key Patents

Representative filing and most-cited prior art

Representative Record
US8407054B22013-03-26

Speech synthesis device, speech synthesis method, and speech synthesis program (US8407054B2, NEC Corporation, 2013-03-26)

NEC CORPORATION

A speech synthesis device selects a central segment from a set of speech segments, generates prosody information from it, then selects non-central segments outside the central segment section based on that prosody information, and finally generates a synthesized speech waveform from the central and non-central segments together.The claim structure ties prosody generation to a specific central-segment selection step before non-central segments are chosen — a sequencing choice that later filers have had to design around rather than copy.

US8407054B2 — patent drawing 1US8407054B2 — patent drawing 2
View full record
Most-cited records in this dataset
#Publication no.Patent titleCitations
1US6665641B1Speech synthesis using concatenation of speech waveforms461
2TW201227715AMulti-lingual text-to-speech synthesis system and method270
3US20040111266A1Speech synthesis using concatenation of speech waveforms201
4US20030078780A1Method and apparatus for controlling a speech synthesis system to provide multiple styles of speech196
5US5029211ASpeech analysis and synthesis system177
6US20200380952A1Multilingual speech synthesis and cross-language voice cloning104
7US6810378B2Method and apparatus for controlling a speech synthesis system to provide multiple styles of speech101
8US20050114137A1Intonation generation method, speech synthesis apparatus using the method and voice server67
9US20130289998A1Realistic Speech Synthesis System66
10US20020120451A1Apparatus and method for providing information by speech55

Citation counts favour older records that have had more time to accumulate citations inside the searched corpus; treat them as a signal of influence on the field, not of current commercial relevance.

Each row carries its publication number; clicking a row searches Eureka by that number.

Source: Patsnap Eureka. Citation counts and representative records. Derived from a Patsnap search on Text-to-Speech and Voice Synthesis covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Run it yourself

Put your own technology through the same analysis

 
Where to run it
Fastest

Eureka on the web

When you want the answer in the next five minutes.

The agent works the prompt against patents and technical literature, citing every source.

Run your analysis now →
For builders

MCP server & REST API

When it has to run inside your own pipeline.

Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.

Browse MCP servers →
Insights

What the numbers mean for a filing decision

Three patterns stand out once the ranking, the trend and the IPC mix are read together: a moderate concentration at the top, a filing curve that has already crested, and a technology mix that is nearly all G10L with limited claims into adjacent computing classes.

Concentration
39.1% / 59.8%
top 5 / top 10 share of 87 records

A leader, not a monopoly

The top 5 assignees combine for 39.1% of the 87 records in scope and the top 10 for 59.8%. That leaves close to 40% of filings spread across a long tail of 23 further ranked assignees plus unranked entrants — room to file without going head-to-head with the leader on every claim.

Based on the full 33-assignee ranking returned for this dataset.
Momentum
9 in 2024
peak filing year

Activity has already crested

Filings rose from 0 in 2017 through a flat 2022 midpoint of 1 to a peak of 9 records in 2024, then dropped back sharply into the partial 2026 year. None of the most recent-year-active assignees logged a filing in the latest year, consistent with a maturing rather than accelerating field.

2026 figure is a partial year under the 2026-07-31 data cut-off.
Technology mix
97.7% G10L
share of 87 records

Nearly everything sits in one subclass

G10L covers 97.7% of records, far ahead of G06F (12.6%), G06N (4.6%), H04M (3.4%) and G11B (1.1%). The gap between G10L and every other class suggests claims framed purely around AI-model computing or telephony integration remain comparatively rare.

Shares sum to more than 100% because records can carry multiple IPC classes.
Eureka AI Agent
Looking for what nobody has claimed yet?

Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to text-to-speech and voice synthesis, with the prior art for and against each one.

Find the white space →
Source: Patsnap Eureka. Co-assignee relationships and derived observations. Derived from a Patsnap search on Text-to-Speech and Voice Synthesis covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Players & White Space

Who is filing, and what is still open

The ranked leaders come from long-established speech-technology and consumer-electronics assignees, with a long tail of smaller or newer entrants including at least one AI-native voice company. The gaps sit less in who is filing than in which claim angles they have left untouched.

Leader
11
records from the top-ranked assignee

One clear leader, then a gradual drop-off

The top-ranked assignee holds 11 records, roughly triple the fifth-place count of 4 and well ahead of tenth place at 3. That gap suggests a genuine leadership position on core synthesis claims rather than a crowded tie at the top.

Counted by patent family across the 87 records in scope.
Long tail
33 ranked
assignees in the ranking

A fragmented field beyond the leaders

The ranking covers 33 assignees, and the drop from 4 records at fifth place to 3 at tenth place shows how quickly counts thin out. Most of the ranked field holds only one or a few records each, which is typical of a technology area still absorbing new entrants alongside its incumbents.

This is the complete ranking returned for the dataset, not a top-50 or top-100 cut.
Venues
US 36 / EPO 20
leading receiving offices

Filing is concentrated in two jurisdictions

The United States (36) and the European Patent Office (20) account for most receiving-office activity, with India (8), WIPO (7), Japan (4) and China (3) trailing well behind. A filing strategy built only around US and EPO coverage would miss a meaningful share of where competitors are already active.

Counts are by receiving office, not by assignee nationality.
🔍
Under-claimed sub-areas worth checking before filing
These branches show light overlap with the core G10L cluster and may offer more open claim space than the dominant synthesis-method filings.
voice cloning consent safeguardsreal-time-factor optimisation for on-device vocodersmultilingual prosody transfertelephony-integrated synthesis (H04M overlap)AI-model-based vocoder architectures (G06N overlap)
Rank all filers by momentum →
Recent-year filing momentum by assignee
AssigneeRecent yearYoY
Google LLC0
International Business Machines Corporation (IBM)0
LERNOUT & HAUSPIE SPEECH PRODS0
CAMB AI INC0
Industrial Technology Research Institute (ITRI)0
Lucent Technologies0
Canon Inc.0
iFLYTEK Co., Ltd.0
Source: Patsnap Eureka. Assignee-level momentum. Derived from a Patsnap search on Text-to-Speech and Voice Synthesis covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
What's Next

Where to take this analysis

The dataset points to specific next steps depending on whether the goal is freedom-to-operate, portfolio benchmarking, or spotting an entry point.

Check freedom-to-operate against the most-cited prior art

The most-cited records in this dataset, several tracing back to concatenative waveform synthesis methods, still shape how later prosody and vocoder claims must be drafted to avoid overlap.

Explore prior art in Eureka

Track the leader's recent filing behaviour

With the top-ranked assignee holding 11 records but no filings in the latest tracked year, watching whether that pace resumes is a useful signal for competitive timing.

Set up assignee tracking in Eureka

Scope claims around the under-claimed branches

Voice cloning safeguards, multilingual prosody transfer and telephony-integrated synthesis show comparatively light overlap with the dominant G10L cluster and may offer clearer claim space.

Draft a claim scope in Eureka
Source: Patsnap Eureka. Forward-looking reading of the same dataset. Derived from a Patsnap search on Text-to-Speech and Voice Synthesis covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
FAQ

Common questions about text-to-speech patents

Answers are grounded in the same dataset. Derived from a Patsnap search on Text-to-Speech and Voice Synthesis covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Research Text-to-Speech and Voice Synthesis in depth with Eureka

Go past this page: query the whole text-to-speech and voice synthesis corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.

Try Eureka

Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.

Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.

Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.

Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.

Help us improve this page

Found incorrect or outdated information? Let us know and we'll get it fixed.