https://www.patsnap.com/resources/blog/rd-blog/storage-engines-and-data-layout-patent-landscape/ · Patsnap · data cut-off 2026-07-31 · downloaded from the live page
Patent Landscape · Database Systems
Storage Engines and Data Layout: Patents Shaping LSM Trees and Columnar Formats
  • Filing has plateaued, not grown. activity peaked at 50 families in 2023 and the pace through the 2022 midpoint was already flat, not building toward a new wave.
  • China now files more than the US. 161 China-origin records versus 86 US-origin puts the bulk of recent claim-writing activity in Chinese receiving offices.
  • Momentum has cooled at nearly every top filer. assignees including Alibaba, Huawei and VMware show zero families in the latest year, a sign of a settled rather than an active filing race.
Get a prior-art report on your approach
280
Published Records
19%
Top-5 Share of All Records
+28%
3-Yr Growth (lag-adjusted)
CN
Leading Jurisdiction
Published byPatsnap Research··7 min readSourced from Patsnap Eureka
Overview

What this landscape covers

This landscape tracks patent families addressing storage engine architecture and data layout - log-structured merge (LSM) trees, columnar storage, and the compaction, compression and lookup mechanics that determine how a storage engine performs under write and scan load. The search combines title/abstract terms for storage engine, LSM tree and columnar storage with claim-level language on write amplification, compaction, compression ratio, point lookup and scan performance, restricted to the G06F16/G06F12/G06F3 data-processing and storage IPC classes.

Coverage runs from 2015 through the July 2026 data cut-off, with 280 published families forming the ranking base. Because publication trails filing by roughly eighteen months, the 2025 and 2026 counts in any trend line understate real filing activity for those years.

Filing activity and technology composition, 2017-2026
  1. 1ALIPAY (HANGZHOU) INFORMATION TECH CO LTD12
  2. 2VMWARE INC12
  3. 3LEMON INC(GB)10
  4. 4HUAZHONG UNIV OF SCI & TECH10
  5. 5HUAWEI TECH CO LTD10
  6. 6HUAQIAO UNIVERSITY8
  7. 7SAMSUNG ELECTRONICS CO LTD8
  8. 8SK HYNIX INC7
  9. 9ZHEJIANG UNIV6
  10. 10BEIJING OCEANBASE TECHNOLOGY CO LTD6
Source: Patsnap Eureka. Assignee ranking and totals. Derived from a Patsnap search on Storage Engines and Data Layout covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Let an AI agent run this analysis on your own technology

Pick a task. Every answer cites the patents behind it.

10,000 free credits to start
The Data

Filing trends and technology composition

Two views of the same 280-family dataset: how filing volume has moved year over year, and how those families sort across IPC subclasses.

A flat trajectory after a 2023 peak

Filings rose from 7 in 2017 to a peak of 50 in 2023, with the 2022 midpoint already at 42 - meaning most of the growth had already happened well before the peak, and the curve has since leveled rather than continued climbing. Read the final one to two years as undercounts given publication lag.

A flat trajectory after a 2023 peak0132538507201720182019202020212022502023502024202572026Most recent year is partial — publication lag means later filings are not yet visible.

Concentrated in core data-processing classes

All 280 records carry a G06F classification, confirming the search is anchored correctly in core data-processing and storage IPC codes. Secondary classes are thin - 9 records touch G06N (AI-based computing), and single digits appear in coding/transmission classes H03M and H04L - indicating storage-engine work here is rarely filed jointly with AI-model or networking claims.

Concentrated in core data-processing classesG06F · Electric digital data processi…280100.0%G06N · Computing based on AI models93.2%H03M · Coding & code conversion31.1%H04L · Digital information transmissi…31.1%G06K · Data recognition & presentation10.4%G06Q · Business, commerce & admin dat…10.4%

Shares are the percentage of the 280 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.

Source: Patsnap Eureka. Filing trend and technology composition. Derived from a Patsnap search on Storage Engines and Data Layout covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.

Go deeper on Storage Engines and Data Layout with Eureka

This page is one run against one query. Ask Eureka your own question about storage engines and data layout and every answer comes back with the patent numbers behind it.

Try Eureka
Key Patents

Representative and most-cited filings

Representative Filing
US11237744B22022-02-01

US11237744B2 - Method and system for configuring a write amplification factor of a storage engine based on a compaction value associated with a data file

YAHOO ASSETS LLC

The patent discloses a method for tying a storage engine's write amplification factor directly to a compaction value - the minimum number of valid records a data file must hold before a clean-up process triggers. A counter tracks records per file; once it satisfies a threshold, each record is evaluated against a second criterion to decide whether it gets rewritten to another file. The write amplification factor is then derived from the compaction value itself, linking a tunable configuration parameter to a measurable performance outcome.Filed by Yahoo Assets LLC, granted 2022.

US11237744B2 — patent drawing 1US11237744B2 — patent drawing 2
View full filing
Most-cited records in this corpus
#Publication no.Patent titleCitations
1US20180349095A1Log-structured merge tree based data storage architecture80
2CN106708427A一种适用于键值对数据的存储方法49
3US20170249257A1Solid-state storage device flash translation layer42
4US20220156231A1Scalable I/O operations on a log-structured merge (LSM) tree38
5US20200201821A1Synchronization of index copies in an LSM tree file system33
6US10394822B2Systems and methods for data conversion and comparison33
7US10430433B2Systems and methods for data conversion and comparison33
8US10423626B2Systems and methods for data conversion and comparison32
9CN112395212A减少键值分离存储系统的垃圾回收和写放大的方法及系统31
10US20200201822A1Lockless synchronization of LSM tree metadata in a distributed system26

Citation counts reflect influence within this searched corpus and skew toward older filings; treat them as a signal of historical impact rather than current relevance.

Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.

Source: Patsnap Eureka. Citation counts and representative records. Derived from a Patsnap search on Storage Engines and Data Layout covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Run it yourself

Put your own technology through the same analysis


  
Where to run it
Fastest

Eureka on the web

When you want the answer in the next five minutes.

The agent works the prompt against patents and technical literature, citing every source.

Run your analysis now →
For builders

MCP server & REST API

When it has to run inside your own pipeline.

Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.

Browse MCP servers →
Insights

What the numbers signal

Three read-throughs from the filing trend, jurisdiction split and citation data that matter for anyone deciding where to file or which prior art to clear first.

Filing Trajectory
50 in 2023
peak year filings

Growth already happened before the peak

The midpoint year of 2022 already sat at 42 filings, just eight below the eventual 2023 peak. That compressed run-up, followed by a flattening rather than a continued climb, points to a technology space where the core architectural claims - LSM compaction strategies, columnar layout variants - have largely been staked out rather than one still being actively opened up.

Based on the 2017-2026 filing trend
Jurisdiction Mix
161 vs 86
China vs US receiving-office records

China-origin filing now outpaces the US

China accounts for 161 records against 86 from the US, with Europe, WIPO/PCT, Hong Kong and South Korea trailing well behind. For teams benchmarking freedom-to-operate, a China-first prior art search is no longer optional - it is where the bulk of recent claim language actually sits.

Based on receiving-office counts
Citation Concentration
80 citations
top-cited record

Influence sits with early LSM tree filings

The most-cited record in the corpus is a log-structured merge tree data storage architecture filing, with the next tier covering flash translation layers and scalable LSM I/O. Citation leadership clusters around foundational LSM tree mechanics rather than newer columnar or compression work, which is consistent with older filings simply having had more time to accumulate citations.

Based on the most-cited records table
Collaboration Pattern
4 co-assignee pairs
joint-filing relationships

Co-filing is rare and mostly corporate-to-subsidiary

Only four co-assignee pairs appear across 280 families, and the strongest link is a corporate parent-subsidiary pairing rather than an arm's-length collaboration. Joint industry-academia filing exists but is thin, suggesting most storage-engine IP here is developed and held internally rather than through formal partnership.

Based on co-assignee pair analysis
Eureka AI Agent
Looking for what nobody has claimed yet?

Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to storage engines and data layout, with the prior art for and against each one.

Find the white space →
Source: Patsnap Eureka. Co-assignee relationships and derived observations. Derived from a Patsnap search on Storage Engines and Data Layout covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Players

Who is filing, and who has slowed down

The assignee base spans large cloud and device vendors, Chinese enterprise software firms, and a handful of universities filing jointly with industry. Recent-year momentum, however, has cooled broadly rather than concentrating around a new leader.

Momentum Signal
0 in latest year
across multiple top filers

Several major assignees show no latest-year filings

Alibaba, Huawei, Huazhong University of Science and Technology, VMware and the Douyin/CapCut Cayman entity all register zero families in the most recent year, including one filer at -100% YoY. That is broad-based cooling rather than a single company pulling back, and it lines up with the flat 2022-2023-2026 trend.

Based on recent-year momentum by assignee
Cross-border Structure
6 shared families
strongest co-assignee pair

Samsung's Korea-China filing link is the tightest pairing

The strongest co-assignee relationship in the dataset pairs Samsung Electronics (Korea) with Samsung's China semiconductor subsidiary across 6 shared families, reflecting a single global R&D program filed through two corporate entities rather than an external collaboration.

Based on co-assignee pair analysis
Filing Geography
161 China-origin
receiving-office records

Chinese enterprise and cloud vendors dominate volume

Beyond the largest names, the assignee base includes numerous mid-size Chinese enterprise storage and cloud-infrastructure firms filing singly rather than in groups, consistent with the China receiving-office count leading the US by a wide margin.

Based on receiving-office counts
🔍
Under-claimed sub-areas worth a closer look
Branches with thin filing density relative to the core LSM-tree and columnar-storage claims
Adaptive compaction scheduling under mixed read/write loadTiered compression ratio selection by hot/cold dataPoint-lookup index synchronization across storage tiersScan-performance optimization for hybrid row-columnar layoutsFlash translation layer coordination with LSM compaction
Rank all filers by momentum →
Recent-year filing momentum by assignee
AssigneeRecent yearYoY
Alipay (Hangzhou) Digital Service Technology Co., Ltd.0-100%
VMware, Inc.0
Faceu Cayman Islands Co., Ltd. (CapCut/ByteDance)0
Huawei Technologies Co., Ltd.0
Huazhong University of Science and Technology0
Alibaba Group Holding Limited0
Inspur Cloud Information Technology Co., Ltd.0
Huaqiao University0-100%
Source: Patsnap Eureka. Assignee-level momentum. Derived from a Patsnap search on Storage Engines and Data Layout covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
What's Next

Where to take this analysis

The filing and citation patterns above raise questions that a static report can't answer on its own - they depend on what you're trying to file, clear, or license next.

Check freedom-to-operate against the China filing base

With 161 of 280 records originating in China, a claim-by-claim review of the leading Chinese storage-engine filers is likely to surface blocking art that a US-only search would miss.

Run a targeted FTO search in Eureka

Map the under-claimed sub-areas to your roadmap

Adaptive compaction scheduling and hybrid row-columnar scan optimization show thin filing density relative to core LSM and columnar claims - worth checking against your own development priorities before you file.

Explore white space in Eureka

Track assignees whose momentum has cooled

Several top filers show zero latest-year activity; confirming whether that reflects a genuine pullback or a publication-lag artifact changes how much weight to put on their existing portfolios.

Monitor assignee activity in Eureka
Source: Patsnap Eureka. Forward-looking reading of the same dataset. Derived from a Patsnap search on Storage Engines and Data Layout covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
FAQ

Common questions on storage engine patents

Answers are grounded in the same dataset. Derived from a Patsnap search on Storage Engines and Data Layout covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Research Storage Engines and Data Layout in depth with Eureka

Go past this page: query the whole storage engines and data layout corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.

Try Eureka

Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.

Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.

Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.

Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company's registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.