Text-to-Speech Patents: Who Leads, Where the Gaps Are 2026
- Concentration is moderate, not extreme. the top 5 assignees hold 39.1% of the 87 records in scope, and the top 10 hold 59.8% — a leader, but a real field behind it.
- Filing has already peaked. the trend hit 9 records in 2024 after a flat 2022 midpoint, and none of the most active recent assignees added a filing in the latest year.
- Almost everything sits in one IPC subclass. 97.7% of records carry G10L, while G06F, G06N and H04M appear in far fewer — a sign that adjacent computing and telephony framing is still thin.
What the patent record shows about text-to-speech and voice synthesis
Text-to-speech and voice synthesis patenting covers everything from waveform concatenation to neural vocoders, with claims typically anchored on prosody control, naturalness measured against MOS scores, and increasingly on voice cloning safeguards and multilingual synthesis. The 87 records in scope span 2015 to the 2026 data cut-off and were filtered against IPC classes covering speech and audio synthesis, AI computing models, and speech analysis specifically.
Filing activity is concentrated in a small set of long-standing speech-technology assignees alongside a longer tail of single- or few-filing entrants, and the receiving-office spread points to the United States and Europe as the primary filing venues, with India, WIPO and Japan behind them.
Let an AI agent run this analysis on your own technology
Pick a task. Every answer cites the patents behind it.
Filing trend and technology composition
The filing curve and the IPC mix together describe a field that grew steadily into a 2024 peak and is now built almost entirely on one core classification, with only light spillover into adjacent computing and telephony subclasses.
A 2024 peak followed by an unfinished 2026
Filings moved from 0 in 2017 to a flat midpoint of 1 in 2022, then rose to a peak of 9 records in 2024. The 2026 figure of 1 is a partial year under the data cut-off and understates real activity, since publication typically lags filing by around 18 months.
G10L dominates; adjacent classes are thin
G10L, speech and audio analysis and synthesis, appears on 97.7% of the 87 records in scope, confirming this is the primary classification for the field. G06F appears on 12.6%, G06N (AI computing models) on 4.6%, H04M (telephonic communication) on 3.4% and G11B (information storage) on 1.1% — each record can carry more than one class, so these shares add up to more than 100%.
Shares are the percentage of the 87 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Text-to-Speech and Voice Synthesis with Eureka
This page is one run against one query. Ask Eureka your own question about text-to-speech and voice synthesis and every answer comes back with the patent numbers behind it.
Try EurekaRepresentative filing and most-cited prior art
Speech synthesis device, speech synthesis method, and speech synthesis program (US8407054B2, NEC Corporation, 2013-03-26)
A speech synthesis device selects a central segment from a set of speech segments, generates prosody information from it, then selects non-central segments outside the central segment section based on that prosody information, and finally generates a synthesized speech waveform from the central and non-central segments together.The claim structure ties prosody generation to a specific central-segment selection step before non-central segments are chosen — a sequencing choice that later filers have had to design around rather than copy.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US6665641B1 | Speech synthesis using concatenation of speech waveforms | 461 |
| 2 | TW201227715A | Multi-lingual text-to-speech synthesis system and method | 270 |
| 3 | US20040111266A1 | Speech synthesis using concatenation of speech waveforms | 201 |
| 4 | US20030078780A1 | Method and apparatus for controlling a speech synthesis system to provide multiple styles of speech | 196 |
| 5 | US5029211A | Speech analysis and synthesis system | 177 |
| 6 | US20200380952A1 | Multilingual speech synthesis and cross-language voice cloning | 104 |
| 7 | US6810378B2 | Method and apparatus for controlling a speech synthesis system to provide multiple styles of speech | 101 |
| 8 | US20050114137A1 | Intonation generation method, speech synthesis apparatus using the method and voice server | 67 |
| 9 | US20130289998A1 | Realistic Speech Synthesis System | 66 |
| 10 | US20020120451A1 | Apparatus and method for providing information by speech | 55 |
Citation counts favour older records that have had more time to accumulate citations inside the searched corpus; treat them as a signal of influence on the field, not of current commercial relevance.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the numbers mean for a filing decision
Three patterns stand out once the ranking, the trend and the IPC mix are read together: a moderate concentration at the top, a filing curve that has already crested, and a technology mix that is nearly all G10L with limited claims into adjacent computing classes.
A leader, not a monopoly
The top 5 assignees combine for 39.1% of the 87 records in scope and the top 10 for 59.8%. That leaves close to 40% of filings spread across a long tail of 23 further ranked assignees plus unranked entrants — room to file without going head-to-head with the leader on every claim.
Activity has already crested
Filings rose from 0 in 2017 through a flat 2022 midpoint of 1 to a peak of 9 records in 2024, then dropped back sharply into the partial 2026 year. None of the most recent-year-active assignees logged a filing in the latest year, consistent with a maturing rather than accelerating field.
Nearly everything sits in one subclass
G10L covers 97.7% of records, far ahead of G06F (12.6%), G06N (4.6%), H04M (3.4%) and G11B (1.1%). The gap between G10L and every other class suggests claims framed purely around AI-model computing or telephony integration remain comparatively rare.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to text-to-speech and voice synthesis, with the prior art for and against each one.
Who is filing, and what is still open
The ranked leaders come from long-established speech-technology and consumer-electronics assignees, with a long tail of smaller or newer entrants including at least one AI-native voice company. The gaps sit less in who is filing than in which claim angles they have left untouched.
One clear leader, then a gradual drop-off
The top-ranked assignee holds 11 records, roughly triple the fifth-place count of 4 and well ahead of tenth place at 3. That gap suggests a genuine leadership position on core synthesis claims rather than a crowded tie at the top.
A fragmented field beyond the leaders
The ranking covers 33 assignees, and the drop from 4 records at fifth place to 3 at tenth place shows how quickly counts thin out. Most of the ranked field holds only one or a few records each, which is typical of a technology area still absorbing new entrants alongside its incumbents.
Filing is concentrated in two jurisdictions
The United States (36) and the European Patent Office (20) account for most receiving-office activity, with India (8), WIPO (7), Japan (4) and China (3) trailing well behind. A filing strategy built only around US and EPO coverage would miss a meaningful share of where competitors are already active.
| Assignee | Recent year | YoY |
|---|---|---|
| Google LLC | 0 | — |
| International Business Machines Corporation (IBM) | 0 | — |
| LERNOUT & HAUSPIE SPEECH PRODS | 0 | — |
| CAMB AI INC | 0 | — |
| Industrial Technology Research Institute (ITRI) | 0 | — |
| Lucent Technologies | 0 | — |
| Canon Inc. | 0 | — |
| iFLYTEK Co., Ltd. | 0 | — |
Where to take this analysis
The dataset points to specific next steps depending on whether the goal is freedom-to-operate, portfolio benchmarking, or spotting an entry point.
Check freedom-to-operate against the most-cited prior art
The most-cited records in this dataset, several tracing back to concatenative waveform synthesis methods, still shape how later prosody and vocoder claims must be drafted to avoid overlap.
Explore prior art in EurekaTrack the leader's recent filing behaviour
With the top-ranked assignee holding 11 records but no filings in the latest tracked year, watching whether that pace resumes is a useful signal for competitive timing.
Set up assignee tracking in EurekaScope claims around the under-claimed branches
Voice cloning safeguards, multilingual prosody transfer and telephony-integrated synthesis show comparatively light overlap with the dominant G10L cluster and may offer clearer claim space.
Draft a claim scope in EurekaCommon questions about text-to-speech patents
Across the 87 records in this dataset, one assignee leads with 11 records, roughly triple the count held by the fifth-ranked assignee at 4. The top 5 assignees combined account for 39.1% of all records, and the top 10 account for 59.8%, which means well over a third of the field sits with a long tail of smaller filers rather than a small handful of dominant firms. This pattern suggests the field has an established leader but is still open enough for new entrants to carve out claim space.
The trend rose from 0 records in 2017 to a peak of 9 in 2024, after sitting flat at 1 record at the 2022 midpoint. Filing counts drop off after 2024 into the partial 2026 year, and none of the most recently active assignees logged a filing in the latest tracked year. Because publication typically lags filing by around 18 months, the most recent years understate real activity, but the overall shape points to a field that has already passed its filing peak rather than one still accelerating.
The dominant classification is G10L, speech and audio analysis and synthesis, which appears on 97.7% of the 87 records in scope. G06F (electric digital data processing) appears on 12.6%, G06N (AI-based computing models) on 4.6%, H04M (telephonic communication) on 3.4%, and G11B (information storage) on 1.1%. Because a single record can carry multiple IPC classes, these shares add up to more than 100%, and the low overlap with G06N and H04M suggests claims explicitly framed around AI-model architecture or telephony integration are still comparatively rare.
US8407054B2, assigned to NEC Corporation and dated 2013-03-26, claims a speech synthesis device that selects a central speech segment, generates prosody information from it, then selects non-central segments based on both the central segment and that prosody information, before generating a final synthesized waveform. The specific sequencing — prosody generation gated on central-segment selection before non-central segments are chosen — is the part later filers need to check against, since copying that exact order without variation risks overlap. It is a useful reference point for anyone drafting segment-selection or prosody-generation claims in this space.
The clearest gaps sit in areas that overlap only lightly with the dominant G10L cluster: voice cloning consent safeguards, real-time-factor optimisation for on-device vocoders, multilingual prosody transfer, and synthesis claims integrated with telephony systems (H04M) or explicit AI-model architectures (G06N). These branches show meaningfully lower representation than core waveform and prosody synthesis claims, which is where most of the 87 records concentrate. A first claim in one of these areas would need to tie the technical mechanism (for example, a specific consent-verification step in a cloning pipeline) to a measurable output like naturalness MOS or real-time factor, rather than restating the general synthesis pipeline already well covered by existing filings.
Research Text-to-Speech and Voice Synthesis in depth with Eureka
Go past this page: query the whole text-to-speech and voice synthesis corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.