Synthetic Data Patents: Top Companies & Filing Trends 2026
- 16.0% concentration at the top. The five leading filers account for 4,013 of the 25,091 records in scope — real concentration, but not a lockout: 84% of filings sit outside that group.
- Filings nearly doubled in three years. Published filings rose from 2,059 in 2021 to 4,071 in 2024, a +98% increase, before the 2025-2026 count understates true activity due to publication lag.
- Momentum is cooling even at the leader. The top filer's most recent-year output fell 78% year-on-year, and every other tracked assignee in the momentum data also declined — a sign publication lag is masking a still-active field, not a sign of retreat.
Filing growth compares 2021 (2,059 records) with 2024 (4,071) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 25,091 records in scope (CR5), not by the ranked leaders only.
What the synthetic data patent record shows
Synthetic data sits at the intersection of generative modelling and data governance: patents in this set claim methods for generating artificial training data, matching statistical distributions, and producing privacy-preserving synthetic samples for downstream model training. The search scope spans 25,091 published records filed or published between 2015 and the 2026 cut-off, drawn from applicants ranging from large AI and semiconductor companies to healthcare, financial-services and imaging firms. Because publication lags filing by roughly 18 months, the most recent one to two years in any trend chart will always look lighter than they eventually turn out to be.
Filing activity is not confined to one industry. Business and commerce data processing, healthcare informatics and digital transmission classes all appear alongside the core AI and imaging classes, which points to synthetic data being adopted as an enabling technique across sectors rather than staying inside a single vertical.
Filing trends and technology composition
The dataset covers 25,091 records published between 2015 and the August 2026 cut-off, ranked by assignee, filing year and IPC subclass.
A near-doubling in three years, then a lag-distorted tail
Published filings climbed from 379 in 2017 to a peak of 4,071 in 2024 — a +98% rise from the 2,059 recorded in 2021 alone. The 2025 and 2026 figures (down to 348 in the partial 2026 year) should be read as incomplete rather than as a downturn, since publication lag has not yet caught up.
AI and imaging classes dominate, but adoption is broad
G06N (AI-based computing) leads at 28.9% of the 25,091 records, followed by G06F (electric digital data processing) at 24.2% and G06T (image data generation) at 19.4%. Because records commonly carry more than one IPC class, these shares add to well over 100%; smaller shares in G06Q (6.0%), H04L (5.4%) and G16H (5.0%) show synthetic data techniques reaching into commerce, networking and clinical informatics filings alongside the core computing classes.
Shares are the percentage of the 25,091 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Synthetic Data Patent Landscape with Eureka
This page is one run against one query. Ask Eureka your own question about synthetic data patent landscape and every answer comes back with the patent numbers behind it.
Try EurekaA representative claim and the most-cited prior art
Synthetic data generation using a classifier — Capital One Services
The disclosure covers a recurrent neural network trained for synthetic data generation: a classifier determines that a sequence of elements corresponds to a token, the network's vocabulary is expanded to include that token, and the modified network is retrained to generate synthetic data using the enlarged vocabulary.Filed by Capital One Services, LLC — published 2021-05-20 as US20210150332A1.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20180165554A1 | Semisupervised autoencoder for sentiment analysis | 793 |
| 2 | US20180018590A1 | Distributed Machine Learning Systems, Apparatus, and Methods | 762 |
| 3 | US20220126864A1 | Autonomous vehicle system | 675 |
| 4 | US20180307859A1 | Systems and methods for enforcing centralized privacy controls in de-centralized systems | 626 |
| 5 | US20170032281A1 | System and Method to Facilitate Welding Software as a Service | 576 |
| 6 | US20180367484A1 | Suggested items for use with embedded applications in chat conversations | 561 |
| 7 | US20180367483A1 | Embedded programs and interfaces for chat conversations | 510 |
| 8 | US20170243028A1 | Systems and Methods for Enhancing Data Protection by Anonosizing Structured and Unstructured Data and Incorpo… | 453 |
| 9 | US5896403A | Dot code and information recording/reproducing system for recording/reproducing the same | 407 |
| 10 | US20180210874A1 | Automatic suggested responses to images received in messages using language model | 400 |
Citation counts favour older filings inside this searched corpus and are a signal of influence, not of current commercial importance.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the filing pattern tells a competitive analyst
Three figures from the ranking and trend data set the shape of the field: moderate concentration, a growth curve still climbing at its last complete year, and IPC coverage that spreads well beyond core AI classes.
Leadership is real but not exclusive
The five leading assignees combine for 4,013 records, 16.0% of all records in scope; the ten leading assignees add up to 21.6%. That leaves the large majority of filings distributed across a long tail of single- and few-filing entrants, which is where freedom-to-operate work usually needs to focus.
The field roughly doubled in three years
Published filings rose from 2,059 in 2021 to a peak of 4,071 in 2024. Because publication lag understates 2025-2026, treat the apparent recent decline as an artefact of timing rather than a real slowdown in filing activity.
AI-model classes lead, but adoption is cross-sector
G06N and G06F together cover the largest share of records, but G06Q, H04L and G16H filings show synthetic data methods being claimed for commerce, networking and healthcare applications rather than staying confined to core model architecture.
Recent-year counts are falling across the board
Every assignee in the recent-momentum data, including the leader, shows a year-on-year decline in the latest tracked year. Given the 18-month publication lag, this pattern is consistent with incomplete recent-year data rather than an actual pullback in R&D investment.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to synthetic data patent landscape, with the prior art for and against each one.
Assignee concentration and where the gaps sit
The ranking covers 100 companies measured by record count, led by a filer with 1,938 records; fifth place holds 333 and tenth place holds 257. That steep early drop-off, followed by a long tail, is typical of a technique adopted broadly rather than owned narrowly.
A single filer well ahead of the field
The leading assignee's record count is more than five times the fifth-place holder's 333, indicating a company that has built synthetic data generation into core infrastructure claims rather than filing opportunistically.
A tight mid-tier band
Records held by the fifth- and tenth-ranked filers sit within roughly a quarter of each other, suggesting a cluster of comparably active filers below the leader rather than a second dominant player.
Co-assignment clusters around a handful of firms
Ten co-assignee pairs appear in the dataset, with the strongest links tied to related corporate entities filing jointly — a pattern typical of large groups splitting ownership across subsidiaries rather than genuine cross-company collaboration.
| Assignee | Recent year | YoY |
|---|---|---|
| NVIDIA Corporation | 62 | -78% |
| Snap Inc. | 8 | -80% |
| Google LLC | 4 | -95% |
| Schlumberger Technology Corporation | 4 | -90% |
| Qualcomm Incorporated | 3 | -96% |
| Capital One Services, LLC | 2 | -92% |
| Microsoft Technology Licensing, LLC | 2 | -90% |
| International Business Machines Corporation | 1 | -90% |
Where to take this analysis
The dataset points to specific next steps depending on whether the goal is freedom-to-operate, licensing strategy or R&D scoping.
Check freedom-to-operate against the leader's portfolio
With one assignee holding 1,938 records against a fifth-place count of 333, any new filing in core generation methods should be checked against that leader's claim scope first.
Run a claim comparison in Eureka →Map the under-claimed branches before filing
Thinner filing density in areas like differential-privacy calibration and cross-modal alignment suggests room for a first-mover claim, but only after confirming the gap holds up against full-text search.
Explore white space in Eureka →Track momentum, not just headline counts
Recent-year declines across all tracked assignees are consistent with publication lag; re-running the trend in six months will separate genuine slowdown from a reporting artefact.
Set up monitoring in Eureka →Frequently asked questions
One assignee leads the ranked field with 1,938 records, well ahead of the fifth-ranked filer's 333 and the tenth-ranked filer's 257. The top five assignees combined hold 4,013 records, 16.0% of the 25,091 records in scope, which means the field is led by one company but not controlled by a small clique. The remaining 84% of filings are spread across a long tail of companies and individual filers, so a freedom-to-operate review needs to look well past the leaderboard.
Published filings rose from 2,059 in 2021 to a peak of 4,071 in 2024, a +98% increase over that span, which is the most recent period that can be treated as complete. Counts for 2025 and 2026 appear lower, but that is expected: publication lags filing by roughly 18 months, so recent years are always undercounted at the point of measurement. Judged on the last complete year, 2024, the field was still expanding, not contracting.
The largest shares sit in G06N (AI-based computing, 28.9% of the 25,091 records), G06F (electric digital data processing, 24.2%) and G06T (image data processing and generation, 19.4%). Smaller but meaningful shares appear in G06Q (business and commerce data processing), H04L (digital information transmission) and G16H (healthcare informatics), showing synthetic data techniques being claimed well outside core AI model architecture. Because a single record can carry several IPC codes, these percentages add to more than 100% and should not be summed into a single total.
The filing, assigned to Capital One Services, LLC and published 2021-05-20, describes training a recurrent neural network for synthetic data generation by using a classifier to detect that a sequence of elements corresponds to a token, then expanding the network's vocabulary to include that token before retraining. It is a specific method for vocabulary-aware sequence generation rather than a broad claim over synthetic data generation as a category. Anyone building token-expansion or classifier-guided sequence generation for synthetic data should read the full claim set before assuming the technique is open.
Filing density is comparatively thin in areas like differential-privacy calibration for synthetic outputs, cross-modal synthetic data alignment, and synthetic data quality scoring metrics relative to the dense core classes in G06N, G06F and G06T. These are not guarantees of an open field — they reflect lower filing volume in the dataset, not a formal clearance search — but they are reasonable starting points for scoping a first claim. Any white-space hypothesis should be verified against full-text prior art before drafting.
The leading assignee by total record count also has the highest recent-year filing volume among tracked companies, though its most recent tracked year shows a 78% year-on-year decline, a pattern shared by every other assignee in the recent-momentum data. This is consistent with the roughly 18-month lag between filing and publication rather than an actual pullback in R&D activity. A clearer read on current activity will emerge once the 2025-2026 filings finish publishing.
Research Synthetic Data Patent Landscape in depth with Eureka
Go past this page: query the whole synthetic data patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.