Data Curation & Governance Patents: Top Companies & Filing Trends 2026
- One assignee alone holds 533 of 2,652 records, and the top 5 filers together account for 35.0% of everything in scope — concentration sits hard at the top of this field.
- Filings grew 82% from 2021 to 2024 (195 to 355), with 2025 still the highest year on record before the usual 18-month publication lag understates the tail.
- G06N (AI-based computing) already touches 22.4% of records, meaning governance and curation claims are increasingly written alongside AI-model language rather than as plain data-management filings.
Filing growth compares 2021 (195 records) with 2024 (355) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 2,652 records in scope (CR5), not by the ranked leaders only.
What this landscape covers
This landscape covers patent families that combine governance-side language — data governance, data stewardship or data curation — with the operational mechanics of moving and shaping that data: data pipelines, metadata records, data validation, schema transformation, metadata management and schema mapping. The search string requires both sides to appear, so the 2,652 records in scope are filings that treat governance as an engineering problem rather than a policy document.
The dataset spans publications from 2015 through the August 2026 cut-off. Because publication trails filing by roughly 18 months, the most recent one to two years understate actual filing activity and should be read as a floor, not a ceiling.
Let an AI agent run this analysis on your own technology
Pick a task. Every answer cites the patents behind it.
Filing trend and technology composition
Two views of the same 2,652 records: how filing volume moved year over year, and which IPC subclasses the claims actually sit in.
Filing trend, 2017–2026
Annual publications rose from 67 in 2017 to a peak of 563 in 2025. The 2021-to-2024 span alone shows 82% growth (195 to 355), and 2026's count of 121 so far is a partial year, not a slowdown signal.
Technology composition by IPC subclass
G06F (electric digital data processing) appears in 75.9% of the 2,652 records, confirming this is fundamentally a data-processing field. G06N (AI-based computing) at 22.4% and H04L (digital transmission) at 19.1% show the two directions claims are extending toward: AI-assisted governance and networked pipeline infrastructure. Because records carry multiple classes, these shares sum past 100% by design.
Shares are the percentage of the 2,652 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Data Curation & Governance Patent Landscape with Eureka
This page is one run against one query. Ask Eureka your own question about data curation & governance patent landscape and every answer comes back with the patent numbers behind it.
Try EurekaA representative claim: curating data before it reaches downstream systems
Curating ambiguous data for use in a data pipeline through interaction with a data source
Methods and systems for curating data by a data manager are disclosed. Data may be curated from various data sources before being provided to downstream consumers that may rely on the trustworthiness of the curated data in order to provide desired computer-implemented services. During the data curation process, data curation resources are used to improve the trustworthiness and/or value of the collected data. However, data curation resources (e.g., data curators, computing resources) may be limited and/or insufficient to perform the data curation process as desired, which may result in unusable and/or uncurated (e.g., untrustworthy) data. Thus, the data may be screened for ambiguous values.Filed by Dell Products L.P., published 2024-12-31 — a recent example of curation logic bound directly into pipeline architecture rather than treated as a separate governance layer.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20210090694A1 | Data based cancer research and treatment systems and methods | 714 |
| 2 | US8200527B1 | Method for prioritizing and presenting recommendations regarding organizaion's customer care capabilities | 667 |
| 3 | US20210192412A1 | Cognitive Intelligent Autonomous Transformation System for actionable Business intelligence (CIATSFABI) | 258 |
| 4 | US20170293356A1 | Methods and Systems for Obtaining, Analyzing, and Generating Vision Performance Data and Modifying Media Base… | 235 |
| 5 | US20230206329A1 | Transaction platforms where systems include sets of other systems | 233 |
| 6 | US10541938B1 | Integration of distributed data processing platform with one or more distinct supporting platforms | 228 |
| 7 | US20220327119A1 | Generating and analyzing a data model to identify relevant data catalog data derived from graph-based data ar… | 215 |
| 8 | US20250259085A1 | Convergent Intelligence Fabric for Multi-Domain Orchestration of Distributed Agents with Hierarchical Memory … | 209 |
| 9 | US20230214925A1 | Transaction platforms where systems include sets of other systems | 205 |
| 10 | US20200026710A1 | Systems and methods for data storage and processing | 187 |
Citation counts favour older filings within the searched corpus — read them as a signal of influence, not of which technology is currently strongest.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the numbers mean for a filing decision
Three patterns stand out once concentration, growth and technology composition are read together.
The top of the field is crowded, the middle is not
With the leader alone holding 533 records and the top 5 combined at 928 (35.0% of all records in scope), core curation and pipeline claims from the leading filers are dense prior art. The drop to 59 records at fifth place and 35 at tenth shows the concentration is steep, not gradual — beyond the top handful, filing density falls off quickly.
Sustained growth through the last complete year
Publications rose from 195 in 2021 to 355 in 2024, a period long enough to rule out a short-lived spike. 2025's count of 563 is the highest on record, though 2026 figures are still filling in under the usual publication lag and should not yet be read as a trend change.
Governance claims are merging with AI-model language
Nearly a quarter of records combine governance or curation subject matter with AI-computing classification, alongside 19.1% touching digital transmission (H04L) and 17.3% touching business-process claims (G06Q). This spread indicates governance patents are no longer filed as a standalone category but woven into AI pipeline and business-system claims.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to data curation & governance patent landscape, with the prior art for and against each one.
Assignee landscape: a steep leader, then a long tail
The ranked assignee list covers 100 companies drawn from the 2,652 records in scope — not a top-50 or top-100 cut, but the full set the data endpoint returns.
One filer's share dwarfs the rest of the ranking
The leading assignee's 533 records is more than nine times the fifth-place count of 59, an unusually steep drop for this kind of ranking. Anyone assessing freedom to operate in core curation or pipeline claims should expect this filer's portfolio to be the first wall to check against.
Recent-year activity is cooling across the tracked filers
Every assignee with recent-year momentum data shows flat or negative year-over-year change, including a filer still posting 19 records in the latest year but down 83% YoY. This pattern is consistent with publication lag rather than an actual pullback in filing — the most recent year is always the most incomplete.
Co-filing is limited and concentrated within corporate families
The ten identified co-assignee pairs are dominated by intra-group relationships — a parent and its subsidiary, or a company and named individual inventors sharing assignment. There is little evidence of cross-company joint filing in this space, suggesting most governance and curation IP is developed and held internally rather than through formal filing partnerships.
| Assignee | Recent year | YoY |
|---|---|---|
| DIGITAL GLOBAL SYSTEMS INC | 19 | -83% |
| Pure Storage Inc | 5 | -85% |
| International Business Machines Corporation (IBM) | 1 | 0% |
| TRUIST BANK | 1 | -91% |
| Microsoft Technology Licensing LLC | 1 | 0% |
| Ab Initio Technology LLC | 0 | -100% |
| Oracle International Corporation | 0 | -100% |
| STRONG FORCE TX PORTFOLIO 2018 LLC | 0 | -100% |
Where to take this next
The trend and concentration data point to specific follow-up work depending on what you are trying to decide.
Map freedom to operate against the leader's portfolio
With one assignee holding 533 of 2,652 records, a claim chart against that portfolio should come before any filing decision in core curation or pipeline mechanics.
Run a freedom-to-operate check in EurekaTrack the under-claimed branches before they fill in
Healthcare informatics and control-system overlaps show thinner claim density than the G06F core — worth a closer prior-art pull before that changes.
Explore white space in EurekaWatch the 2025–2026 filings as they publish
2025 is already the peak year on record; 2026 will keep filling in for another 18 months. Set up monitoring rather than relying on a single snapshot.
Set up filing alerts in EurekaFrequently asked questions
One assignee leads the ranked field with 533 of the 2,652 records in scope, well ahead of the fifth-ranked filer at 59 and the tenth-ranked filer at 35. That gap indicates a genuinely concentrated field rather than a gradual decline from first to tenth place. The top 5 filers combined hold 35.0% of all records, and the top 10 hold 42.9%, so a large share of activity sits with a small number of organizations even though 100 companies appear in the full ranking.
It is growing. Filings rose from 195 in 2021 to 355 in 2024, an increase of 82% over that three-year span, and 2025 is the highest year on record so far at 563. Figures for 2025 and 2026 will keep rising as publications catch up, since patent publication typically lags filing by around 18 months — so treat the most recent one to two years as undercounted rather than as evidence of a slowdown.
Electric digital data processing (G06F) appears in 75.9% of the 2,652 records, making it the dominant classification by a wide margin. AI-based computing (G06N) follows at 22.4%, digital transmission (H04L) at 19.1%, and business/commerce data processing (G06Q) at 17.3%. Because a single record can carry multiple IPC classes, these percentages add up to more than 100%, and they show governance claims increasingly written alongside AI and networked-pipeline language rather than as a standalone category.
The dataset shows comparatively thin claim density in healthcare informatics overlaps (G16H, 7.2% of records), control and regulating systems overlaps (G05B, 3.8%), and image/video recognition overlaps (G06V, 4.5%) compared with the dominant G06F, G06N and G06Q clusters. These are areas where governance and curation mechanics are being applied to a specific vertical or data type rather than claimed generically, and filing density there is measurably lower. That does not guarantee an easy grant, but it does mean less prior art to design around in those specific combinations.
Not extensively. Only 10 co-assignee pairs were identified across the dataset, and the strongest pairs are intra-corporate relationships — a parent company filing alongside its own subsidiary or named inventors — rather than joint filings between unrelated companies. This suggests that most data curation and governance IP in this space is developed and held internally, with formal cross-company filing partnerships remaining rare.
Research Data Curation & Governance Patent Landscape in depth with Eureka
Go past this page: query the whole data curation & governance patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.