Data & Database Patents: Top Companies & Filing Trends 2026
- Filings kept climbing through 2024, up 24% from 1,216 in 2021 to 1,506 in 2024, with a peak of 1,534 in 2023.
- The field is not consolidated at the top, the leading assignee holds 1,212 records while the ranked field's top 10 combine for just 27.8% of all 24,104 records in scope.
- Core data processing dominates the claim space, G06F covers 61.0% of records, more than four times the share of the next largest class, H04L at 15.6%.
Filing growth compares 2021 (1,216 records) with 2024 (1,506) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 24,104 records in scope (CR5), not by the ranked leaders only.
What this landscape covers
This landscape tracks 24,104 published records filed or published between 2015 and mid-2026 that touch data management, database technology and knowledge representation — spanning database systems, data lakehouse and governance architectures, semantic knowledge graphs, vector search, stream processing and data observability. The scope draws a wide net across enterprise software and infrastructure claims rather than a single narrow technique, so the composition figures below matter as much as the headline count.
Because publication lags filing by roughly 18 months, the most recent one to two years in the trend chart will fill in further as later filings publish; treat 2024 as the last complete year and read 2025-2026 as a floor, not a ceiling.
Filing trend and technology composition
Two views of the same 24,104-record dataset: filing activity over time, and how records distribute across IPC subclasses.
Sustained growth through the last complete filing year
Annual filings rose from 884 in 2017 to a peak of 1,534 in 2023. The 2021-2024 span shows +24% growth (1,216 to 1,506), and 2026 figures are still partial, so the apparent tail-off after 2024 reflects publication lag rather than a real decline.
Core computing classes carry the claim density
G06F (electric digital data processing) touches 61.0% of the 24,104 records, with H04L (digital transmission, 15.6%) and G06Q (business/commerce data processing, 13.1%) next. G06N (AI-based computing, 7.9%) and smaller shares in H04W, G16H, H04N and G06T show where data techniques are being claimed alongside wireless, healthcare and imaging applications; because records can carry multiple classes, these shares sum to well over 100%.
Shares are the percentage of the 24,104 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Data, Knowledge & Database Technologies Patent Landscape with Eureka
This page is one run against one query. Ask Eureka your own question about data, knowledge & database technologies patent landscape and every answer comes back with the patent numbers behind it.
Try EurekaMost-cited records and a recent filing
EP4791059A1 — managing a mobile communication network via a semantic knowledge graph
The application describes modeling a mobile communication network and its telecommunication domain knowledge as a semantic knowledge graph representing an ontology of structured concepts, taxonomies, relationships and constraints, configured to support logical inference. Runtime data obtained from the network is then correlated against that graph.Filed by Deutsche Telekom, published 2026-08-12 — one of the most recent records in scope, illustrating how semantic knowledge-graph claims are moving into telecom network operations.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US5892900A | Systems and methods for secure transaction management and electronic rights protection | 6,075 |
| 2 | US6850252B1 | Intelligent electronic appliance system and method | 4,059 |
| 3 | US20090254572A1 | Digital information infrastructure and method | 2,826 |
| 4 | US5862325A | Computer-based communication system and method using metadata defining a control structure | 2,710 |
| 5 | US6571282B1 | Block-based communication in a communication services patterns environment | 2,353 |
| 6 | US6400996B1 | Adaptive pattern recognition based control system and method | 2,342 |
| 7 | US6321158B1 | Integrated routing/mapping information | 2,264 |
| 8 | US20170006135A1 | Systems, methods, and devices for an enterprise internet-of-things application development platform | 1,906 |
| 9 | US5982891A | Systems and methods for secure transaction management and electronic rights protection | 1,903 |
| 10 | US6601233B1 | Business components framework | 1,860 |
Citation counts accumulate over time, so older filings such as US5892900A and US6850252B1 lead this table by virtue of age as much as continuing relevance; treat citation rank as a signal of historical influence, not current importance.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the data signals
Three patterns stand out once filing volume, concentration and citation data are read together.
A leader, then a long tail
The top assignee holds 1,212 records, well ahead of fifth place at 620 and tenth at 431. Yet the top 5 combined account for only 17.7% of all records in scope, and the top 10 for 27.8% — this is a field with a recognisable leader but no dominant bloc controlling claim space.
Momentum built steadily, not explosively
Annual filings grew from 1,216 in 2021 to 1,506 in 2024, a +24% rise, after peaking at 1,534 in 2023. That growth sits on top of a base that had already more than doubled since 2017 (884 filings), pointing to a maturing but still-expanding filing base rather than a short-lived spike.
Recent-year filings from major players have dropped sharply
Several of the most active assignees, including Nvidia, Pure Storage, Snowflake, Oracle and IBM, show steep year-over-year declines in the latest year, ranging from -78% to -100%. Given the roughly 18-month publication lag, this pattern is consistent with recent filings not yet having published rather than an actual pullback in R&D.
Core data processing claims dominate the class mix
G06F covers well over half of all records, with H04L and G06Q as the next largest classes. The smaller shares in G06N, H04W, G16H, H04N and G06T mark where data and knowledge-representation techniques intersect with AI, wireless, healthcare and imaging applications — areas that may be less crowded than the G06F core.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to data, knowledge & database technologies patent landscape, with the prior art for and against each one.
Who is filing, and where the field is still open
The ranked list spans 100 companies pulled from the same 24,104-record dataset, from large enterprise software vendors to storage and cloud infrastructure specialists.
A single leader well ahead of the field
The top-ranked assignee holds 1,212 records, nearly double the fifth-place total of 620. That gap suggests a sustained, broad filing programme rather than a handful of flagship applications.
A gradual drop-off, not a cliff
Fifth place sits at 620 records and tenth at 431 — a gentle slope rather than a sharp fall, indicating a group of similarly active filers behind the leader rather than one runner-up.
Cross-entity filing is concentrated within corporate families
The strongest co-assignee pairs link a single parent company with its national or regional subsidiaries, rather than joint filings between unrelated companies — a pattern typical of multinational IP administration rather than open collaboration.
| Assignee | Recent year | YoY |
|---|---|---|
| Nvidia Corp | 18 | -78% |
| Pure Storage Inc | 10 | -81% |
| Snowflake Inc | 2 | -88% |
| Oracle International Corp | 1 | -96% |
| Solidigm (or similar storage/compute company) | 1 | -96% |
| International Business Machines Corporation (IBM) | 0 | -100% |
| SAP SE | 0 | -100% |
| Amazon Technologies Inc | 0 | -100% |
Where to take this analysis
The dataset points to a field with an identifiable leader, steady growth through 2024, and technology sub-areas that remain comparatively open.
Check freedom-to-operate against the most-cited records
Before filing in core database or data-governance claim areas, review the highest-cited records in this dataset, several of which date to the late 1990s and early 2000s and may still carry live claims.
Explore prior art in EurekaTrack momentum shifts as later years publish
The sharp year-over-year declines among leading filers likely reflect publication lag rather than reduced R&D; revisit assignee momentum once 2025-2026 filings have fully published.
Set up monitoring in EurekaDraft around under-claimed sub-areas
Semantic knowledge-graph applications outside core enterprise software, such as telecom network management, show comparatively light claim density and may offer more room to file a strong first claim.
Search white space in EurekaCommon questions about this landscape
The ranked assignee list covers 100 companies drawn from the 24,104 records in scope, led by an enterprise technology company with 1,212 records. The field is not tightly consolidated: the top 5 assignees combined hold only 17.7% of all records, and the top 10 hold 27.8%, so there is a long tail of active filers behind the leaders. Large enterprise software, cloud infrastructure and storage companies appear throughout the ranking rather than a single dominant player controlling most of the claim space.
Filings grew from 884 in 2017 to a peak of 1,534 in 2023, and the 2021-2024 span shows +24% growth, from 1,216 to 1,506 filings. Figures for 2025 and 2026 appear lower, but that reflects the roughly 18-month lag between filing and publication rather than an actual slowdown. 2024 is the most recent year that can be treated as a complete picture of filing activity.
The search scope spans database systems, data lakehouse architectures, data governance, semantic knowledge graphs, vector search, stream processing and data observability, all filtered against core data-management and knowledge-representation terms. In IPC terms, 61.0% of the 24,104 records fall under G06F (electric digital data processing), with meaningful shares in H04L (digital transmission), G06Q (business data processing) and G06N (AI-based computing). Because records can carry multiple IPC codes, these shares overlap rather than sum to 100%.
Sub-areas such as vector search indexing, data observability pipelines, and semantic knowledge-graph applications outside core enterprise software (like telecom network management, illustrated by EP4791059A1) carry comparatively fewer records than the dominant G06F-classified core. That does not guarantee an easy grant, but it indicates less densely occupied claim space than mainstream database or business-data-processing filings. Anyone drafting in these areas should still run a full freedom-to-operate search against the most-cited records in the dataset.
Citation counts accumulate over time within a searched corpus, so records filed in the late 1990s and early 2000s, such as US5892900A with 6,075 citations, naturally outrank newer filings regardless of current commercial relevance. This is a general property of citation analysis, not a sign that older technology is more important today. Recent filings, like EP4791059A1 from 2026, should be evaluated on their own claims and abstract rather than compared directly against citation counts built up over two decades.
Research Data, Knowledge & Database Technologies Patent Landscape in depth with Eureka
Go past this page: query the whole data, knowledge & database technologies patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.