https://www.patsnap.com/resources/blog/rd-blog/large-language-models-patent-landscape-patent-landscape/ · Patsnap · data cut-off 2026-08-31 · downloaded from the live page
Patent Landscape · Artificial Intelligence & Machine Learning
Large Language Model Patents: Who Leads and Where the Field Is Still Open
  • 20.4% concentration at the top. The top 5 assignees hold 6,359 of 31,128 records in scope — a real lead, but not a lock on the field.
  • Filings grew 520% from 2021 to 2024. 859 records in 2021 rose to 5,329 in 2024, the last year the trend can be read as complete.
  • G06F and G06N dominate, G16H barely registers. Digital data processing and AI-model computing classes cover roughly a third and a fifth of records; healthcare informatics sits at just 3.4%.
Get a prior-art report on your approach
31.1K
Published Records
20%
Top-5 Share of All Records
+520%
Filing Growth 2021→2024
US
Leading Jurisdiction

Filing growth compares 2021 (859 records) with 2024 (5,329) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 31,128 records in scope (CR5), not by the ranked leaders only.

Published byPatsnap Research··8 min readSourced from Patsnap Eureka
Overview

What the large language model patent record actually shows

The dataset spans 31,128 published records filed against large language model and related speech/language-representation claims between 2015 and the 2026 cut-off. Filing activity was modest through the late 2010s and accelerated sharply from 2021 onward, tracking the point at which transformer-based systems moved from research demonstrations into product roadmaps. Publication lags filing by roughly 18 months, so the most recent one to two years in any chart will always look thinner than the underlying filing activity actually was.

Coverage is broad rather than narrow: claims range from core model training and inference mechanics through speech recognition, image/video recognition and business-process applications. That breadth means the field is less a single contest over one claim type and more a set of adjacent fights — training methodology, inference efficiency, multimodal grounding, and domain-specific deployment — each with its own density and its own gaps.

Filing volume by year, 2017–2026
  1. 1MICROSOFT TECHNOLOGY LICENSING LLC1,947
  2. 2GOOGLE LLC1,615
  3. 3SAMSUNG ELECTRONICS CO LTD1,070
  4. 4AMAZON TECH INC1,001
  5. 5INTERNATIONAL BUSINESS MACHINE CORPORATION726
  6. 6NVIDIA CORP656
  7. 7NUANCE COMMUNICATIONS INC463
  8. 8QUALCOMM INC399
  9. 9SALESFORCE INC333
  10. 10ORACLE INT CORP293
Source: Patsnap Eureka. Assignee ranking and totals. Derived from a Patsnap search on Large Language Models Patent Landscape covering 2015–2026, data cut-off 2026-08-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Let an AI agent run this analysis on your own technology

Pick a task. Every answer cites the patents behind it.

10,000 free credits to start
The Numbers

Filing trends and technology composition

Three views of the same 31,128-record dataset: how filing volume moved over time, how records classify across the IPC subclasses that govern the technology, and where those filings actually land jurisdictionally.

A ten-fold climb, then a plateau at the edge of visibility

Annual filings moved from 344 in 2017 to a peak of 5,329 in 2024, a +520% rise across just 2021 to 2024. The 2025 and 2026 figures sit far lower, but that reflects publication lag, not a real slowdown — treat the last one to two years as still filling in rather than as evidence of a falling trend.

A ten-fold climb, then a plateau at the edge of visibility01,5003,0004,5006,00034420172018201920202021202220235,329202420254922026Most recent year is partial — publication lag means later filings are not yet visible.

Data processing and speech classes carry the field

G06F (electric digital data processing) touches 31.9% of records and G06N (AI-model computing) 20.1%, together anchoring the core technical claims. G10L (speech and audio) covers 18.0%, reflecting how much of the corpus traces back to speech-interface applications. Business-process (G06Q), transmission (H04L) and healthcare informatics (G16H) each sit under 6%, marking them as narrower, less contested territory relative to the core.

Data processing and speech classes carry the fieldG06F · Electric digital data processi…9,91531.9%G06N · Computing based on AI models6,26020.1%G10L · Speech & audio analysis/synthe…5,61818.0%G06V · Image/video recognition1,8505.9%G06Q · Business, commerce & admin dat…1,8095.8%H04L · Digital information transmissi…1,5094.8%G06T · Image data processing & genera…1,2344.0%G16H · Healthcare informatics1,0613.4%Other5,02616.1%

Shares are the percentage of the 31,128 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.

Source: Patsnap Eureka. Filing trend and technology composition. Derived from a Patsnap search on Large Language Models Patent Landscape covering 2015–2026, data cut-off 2026-08-31. Counts reflect published records only and shift as new filings publish.

Go deeper on Large Language Models Patent Landscape with Eureka

This page is one run against one query. Ask Eureka your own question about large language models patent landscape and every answer comes back with the patent numbers behind it.

Try Eureka
Key Patents

A representative filing and the most-cited prior art

Representative recent filing
US20250232129A12025-07-17

Combining data selection and reward functions for tuning large language models using reinforcement learning

INTERNATIONAL BUSINESS MACHINES CORPORATION

A computer-implemented method, computer program product and system for tuning large language models: pairs of textual prompts and ground-truth labels are received, a data selection scoring function is created by repurposing one or more reward functions to compute similarity between prompts and labels, a training dataset is selected from those pairs using that scoring function, and the large language model is tuned accordingly.Filed by International Business Machines Corporation, published 2025-07-17.

US20250232129A1 — patent drawing 1US20250232129A1 — patent drawing 2
View full filing
Most-cited records in the corpus
#Publication no.Patent titleCitations
1US20020135618A1System and method for multi-modal focus detection, referential ambiguity resolution and mood classification u…881
2US20060149558A1Synchronized pattern recognition source data processed by manual or automatic means for creation of shared sp…827
3US6070140ASpeech recognizer781
4US20110060587A1Command and control utilizing ancillary information in a mobile voice-to-speech application770
5US5477451AMethod and system for natural language translation754
6US7702508B2System and method for natural language processing of query answers720
7US20210090694A1Data based cancer research and treatment systems and methods712
8US20020184373A1Conversational networking via transport, coding and control conversational protocols696
9US20090144609A1NLP-based entity recognition and disambiguation691
10US6964023B2System and method for multi-modal focus detection, referential ambiguity resolution and mood classification u…663

Citation counts favour older filings simply because they have had longer to accumulate citations inside this searched corpus — read them as a signal of influence on later work, not as a ranking of current technical importance.

Each row carries its publication number; clicking a row searches Eureka by that number.

Source: Patsnap Eureka. Citation counts and representative records. Derived from a Patsnap search on Large Language Models Patent Landscape covering 2015–2026, data cut-off 2026-08-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Run it yourself

Put your own technology through the same analysis


  
Where to run it
Fastest

Eureka on the web

When you want the answer in the next five minutes.

The agent works the prompt against patents and technical literature, citing every source.

Run your analysis now →
For builders

MCP server & REST API

When it has to run inside your own pipeline.

Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.

Browse MCP servers →
Insights

What the concentration and citation data mean for filing strategy

Read together, the ranking, the technology split and the citation table point to a field with a real leadership cluster, a long tail of single-filing entrants, and prior art that is older than the current wave of transformer-specific claims.

Concentration
20.4%
of 31,128 records held by top 5

A leadership cluster, not a monopoly

The top 5 assignees combine for 6,359 of 31,128 records in scope, and the top 10 add up to 27.3%. That leaves nearly three-quarters of the corpus spread across a long tail of filers, including single-filing entrants — the field has recognisable leaders but no single gatekeeper.

Ranking is drawn from 100 companies, the full ranked set returned for this dataset.
Growth
+520%
2021 → 2024 filing growth

The curve turned sharply, then the data thins out

Filings rose from 859 in 2021 to a peak of 5,329 in 2024. Because publication lags filing by around 18 months, 2025 and 2026 figures understate real activity and should not be read as a slowdown.

2024 is the most recent year that can be treated as a complete annual figure.
Technology mix
31.9%
of records classed under G06F

Core data-processing claims dominate, healthcare trails

G06F and G06N together anchor the corpus's technical core, while G10L shows how much of the field still runs through speech-interface applications rather than pure language modelling. G16H, healthcare informatics, covers only 3.4% of records — a narrow slice relative to the overall field.

Class shares sum to more than 100% because records can carry multiple IPC classes.
Prior art
881 citations
on the most-cited record

Foundational speech-recognition art still anchors citation counts

The most-cited records in this corpus are speech recognition and multi-modal input filings predating the current LLM wave, not recent transformer-specific claims. That is expected in a searched corpus — older filings have had more time to accumulate citations — but it means citation rank is a poor proxy for which claims matter to a filing decision today.

Citation counts reflect influence inside this corpus, not current technical importance.
Eureka AI Agent
Looking for what nobody has claimed yet?

Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to large language models patent landscape, with the prior art for and against each one.

Find the white space →
Source: Patsnap Eureka. Co-assignee relationships and derived observations. Derived from a Patsnap search on Large Language Models Patent Landscape covering 2015–2026, data cut-off 2026-08-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Players

Who is filing, and where their momentum is heading

The ranked leaders combine large diversified technology filers with speech-and-communications specialists. Recent-year momentum figures for several leaders show sharp year-on-year declines, but that pattern is consistent with publication lag rather than a real pullback from the field.

Leader
1,947
records, top-ranked assignee

A single filer sits well ahead of the field

The top-ranked assignee holds 1,947 records against a fifth-place figure of 726 and a tenth-place figure of 293 — a steep drop-off after the leader that flattens quickly through the rest of the ranked set.

Based on the 100-company ranked set for this dataset.
Momentum
-85% to -93% YoY
latest-year change across several leaders

Recent-year counts look thin — read that as publication lag, not retreat

Several of the largest filers show latest-year counts far below prior years, with year-on-year declines in the -80% to -93% range. Given the roughly 18-month publication lag, these are provisional figures still filling in rather than evidence that leaders are exiting the field.

Momentum figures are latest-year snapshots and will revise upward as later publications land.
Collaboration
10 pairs
co-assignee pairs identified

Co-filing is limited and mostly internal

Only 10 co-assignee pairs appear in the dataset, and the strongest links pair a large filer with either an individual inventor-assignee or an affiliated subsidiary rather than with an unrelated competitor. Cross-company joint filing is not yet a notable pattern in this field.

Strongest pair recorded 17 shared filings.
🔍
Under-claimed branches worth a closer look
These sub-areas sit below the density of the core G06F/G06N/G10L claim clusters, based on the technology composition in this dataset.
Reward-function data selection for fine-tuningMultimodal grounding for LLM inferenceHealthcare-specific language model deployment (G16H)Business-process LLM integration claims (G06Q)Low-power on-device inference architectures
Rank all filers by momentum →
Recent-year filing momentum by assignee
AssigneeRecent yearYoY
Google LLC22-85%
NVIDIA Corporation15-91%
Microsoft Technology Licensing, LLC8-93%
Samsung Electronics Co., Ltd.7-88%
Amazon Technologies, Inc.5-81%
Oracle International Corporation5-93%
Qualcomm Incorporated2-97%
International Business Machines Corporation0-100%
Source: Patsnap Eureka. Assignee-level momentum. Derived from a Patsnap search on Large Language Models Patent Landscape covering 2015–2026, data cut-off 2026-08-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
What's Next

Where to take this analysis

The dataset points to a field with clear leaders but wide open claim space beyond the core. Turning that into a filing or freedom-to-operate decision means going deeper on the specific sub-areas that matter to your roadmap.

Map your own claim language against the corpus

Run your draft claims against the training-dataset, model-parameter and inference-mechanism language that defines this field's densest classes before you file, to see where you sit relative to existing art.

Explore in Eureka

Track the under-claimed branches as they fill in

Healthcare informatics and business-process integration are thin today but will not stay that way if filing volume keeps climbing at anything like the 2021-2024 rate.

Set up monitoring in Eureka

Check freedom-to-operate against the leadership cluster

With 27.3% of records held by the top 10 assignees, a targeted clearance search against those specific portfolios is more useful than a broad keyword sweep.

Run a clearance search in Eureka
Source: Patsnap Eureka. Forward-looking reading of the same dataset. Derived from a Patsnap search on Large Language Models Patent Landscape covering 2015–2026, data cut-off 2026-08-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
FAQ

Common questions about large language model patents

Answers are grounded in the same dataset. Derived from a Patsnap search on Large Language Models Patent Landscape covering 2015–2026, data cut-off 2026-08-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Research Large Language Models Patent Landscape in depth with Eureka

Go past this page: query the whole large language models patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.

Try Eureka

Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.

Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.

Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.

Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company's registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.