Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →This landscape tracks patent families addressing storage engine architecture and data layout - log-structured merge (LSM) trees, columnar storage, and the compaction, compression and lookup mechanics that determine how a storage engine performs under write and scan load. The search combines title/abstract terms for storage engine, LSM tree and columnar storage with claim-level language on write amplification, compaction, compression ratio, point lookup and scan performance, restricted to the G06F16/G06F12/G06F3 data-processing and storage IPC classes.
Coverage runs from 2015 through the July 2026 data cut-off, with 280 published families forming the ranking base. Because publication trails filing by roughly eighteen months, the 2025 and 2026 counts in any trend line understate real filing activity for those years.
Pick a task. Every answer cites the patents behind it.
Two views of the same 280-family dataset: how filing volume has moved year over year, and how those families sort across IPC subclasses.
Filings rose from 7 in 2017 to a peak of 50 in 2023, with the 2022 midpoint already at 42 - meaning most of the growth had already happened well before the peak, and the curve has since leveled rather than continued climbing. Read the final one to two years as undercounts given publication lag.
All 280 records carry a G06F classification, confirming the search is anchored correctly in core data-processing and storage IPC codes. Secondary classes are thin - 9 records touch G06N (AI-based computing), and single digits appear in coding/transmission classes H03M and H04L - indicating storage-engine work here is rarely filed jointly with AI-model or networking claims.
Shares are the percentage of the 280 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
This page is one run against one query. Ask Eureka your own question about storage engines and data layout and every answer comes back with the patent numbers behind it.
Try EurekaThe patent discloses a method for tying a storage engine's write amplification factor directly to a compaction value - the minimum number of valid records a data file must hold before a clean-up process triggers. A counter tracks records per file; once it satisfies a threshold, each record is evaluated against a second criterion to decide whether it gets rewritten to another file. The write amplification factor is then derived from the compaction value itself, linking a tunable configuration parameter to a measurable performance outcome.Filed by Yahoo Assets LLC, granted 2022.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20180349095A1 | Log-structured merge tree based data storage architecture | 80 |
| 2 | CN106708427A | 一种适用于键值对数据的存储方法 | 49 |
| 3 | US20170249257A1 | Solid-state storage device flash translation layer | 42 |
| 4 | US20220156231A1 | Scalable I/O operations on a log-structured merge (LSM) tree | 38 |
| 5 | US20200201821A1 | Synchronization of index copies in an LSM tree file system | 33 |
| 6 | US10394822B2 | Systems and methods for data conversion and comparison | 33 |
| 7 | US10430433B2 | Systems and methods for data conversion and comparison | 33 |
| 8 | US10423626B2 | Systems and methods for data conversion and comparison | 32 |
| 9 | CN112395212A | 减少键值分离存储系统的垃圾回收和写放大的方法及系统 | 31 |
| 10 | US20200201822A1 | Lockless synchronization of LSM tree metadata in a distributed system | 26 |
Citation counts reflect influence within this searched corpus and skew toward older filings; treat them as a signal of historical impact rather than current relevance.
Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →Three read-throughs from the filing trend, jurisdiction split and citation data that matter for anyone deciding where to file or which prior art to clear first.
The midpoint year of 2022 already sat at 42 filings, just eight below the eventual 2023 peak. That compressed run-up, followed by a flattening rather than a continued climb, points to a technology space where the core architectural claims - LSM compaction strategies, columnar layout variants - have largely been staked out rather than one still being actively opened up.
China accounts for 161 records against 86 from the US, with Europe, WIPO/PCT, Hong Kong and South Korea trailing well behind. For teams benchmarking freedom-to-operate, a China-first prior art search is no longer optional - it is where the bulk of recent claim language actually sits.
The most-cited record in the corpus is a log-structured merge tree data storage architecture filing, with the next tier covering flash translation layers and scalable LSM I/O. Citation leadership clusters around foundational LSM tree mechanics rather than newer columnar or compression work, which is consistent with older filings simply having had more time to accumulate citations.
Only four co-assignee pairs appear across 280 families, and the strongest link is a corporate parent-subsidiary pairing rather than an arm's-length collaboration. Joint industry-academia filing exists but is thin, suggesting most storage-engine IP here is developed and held internally rather than through formal partnership.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to storage engines and data layout, with the prior art for and against each one.
The assignee base spans large cloud and device vendors, Chinese enterprise software firms, and a handful of universities filing jointly with industry. Recent-year momentum, however, has cooled broadly rather than concentrating around a new leader.
Alibaba, Huawei, Huazhong University of Science and Technology, VMware and the Douyin/CapCut Cayman entity all register zero families in the most recent year, including one filer at -100% YoY. That is broad-based cooling rather than a single company pulling back, and it lines up with the flat 2022-2023-2026 trend.
The strongest co-assignee relationship in the dataset pairs Samsung Electronics (Korea) with Samsung's China semiconductor subsidiary across 6 shared families, reflecting a single global R&D program filed through two corporate entities rather than an external collaboration.
Beyond the largest names, the assignee base includes numerous mid-size Chinese enterprise storage and cloud-infrastructure firms filing singly rather than in groups, consistent with the China receiving-office count leading the US by a wide margin.
| Assignee | Recent year | YoY |
|---|---|---|
| Alipay (Hangzhou) Digital Service Technology Co., Ltd. | 0 | -100% |
| VMware, Inc. | 0 | — |
| Faceu Cayman Islands Co., Ltd. (CapCut/ByteDance) | 0 | — |
| Huawei Technologies Co., Ltd. | 0 | — |
| Huazhong University of Science and Technology | 0 | — |
| Alibaba Group Holding Limited | 0 | — |
| Inspur Cloud Information Technology Co., Ltd. | 0 | — |
| Huaqiao University | 0 | -100% |
The filing and citation patterns above raise questions that a static report can't answer on its own - they depend on what you're trying to file, clear, or license next.
With 161 of 280 records originating in China, a claim-by-claim review of the leading Chinese storage-engine filers is likely to surface blocking art that a US-only search would miss.
Run a targeted FTO search in EurekaAdaptive compaction scheduling and hybrid row-columnar scan optimization show thin filing density relative to core LSM and columnar claims - worth checking against your own development priorities before you file.
Explore white space in EurekaSeveral top filers show zero latest-year activity; confirming whether that reflects a genuine pullback or a publication-lag artifact changes how much weight to put on their existing portfolios.
Monitor assignee activity in EurekaThis landscape covers patent families whose title or abstract references a storage engine, LSM tree or columnar storage, combined with claim-level language on write amplification, compaction, compression ratio, point lookup or scan performance. It is restricted to IPC classes G06F16, G06F12 and G06F3, which cover data retrieval, memory architecture and input/output arrangements respectively. That combination excludes general database-query patents that don't touch the underlying storage layer, and excludes storage-hardware patents that don't address data layout or compaction logic.
No - filings rose steadily from 7 in 2017 to a peak of 50 in 2023, but the midpoint year of 2022 was already at 42, meaning most of the growth had happened before the peak. The trend has since flattened rather than continued climbing. Because publication lags filing by roughly 18 months, the 2025 and 2026 figures will revise upward somewhat, but the broader pattern is a plateau, not sustained growth.
China leads by receiving-office count, with 161 records against 86 from the United States in this dataset. Europe (EPO), WIPO/PCT filings, Hong Kong and South Korea make up the remainder in much smaller numbers. For anyone assessing where competitive claim density is highest, a China-first search of assignees and claim language is essential rather than optional.
US11237744B2, assigned to Yahoo Assets LLC and granted in 2022, claims a method for deriving a storage engine's write amplification factor from a compaction value - the minimum count of valid records in a data file needed to trigger clean-up. It ties a specific configuration parameter (compaction value) to a specific performance metric (write amplification) through a counter-and-threshold mechanism. Designs that determine write amplification independently of a stored compaction value, or that trigger compaction through different criteria such as fixed time intervals or size thresholds unrelated to record counts, sit outside its literal claim scope, but any implementation using this exact linkage should be reviewed against it directly.
The thinnest filing density relative to the core LSM-tree and columnar-storage claims appears in adaptive compaction scheduling that responds to mixed read/write load, tiered compression ratio selection based on hot/cold data classification, and scan-performance optimization for hybrid row-columnar layouts. These are adjacent to heavily-claimed core mechanics but are not yet occupied by dense prior art in this corpus. That doesn't guarantee novelty for any specific claim, but it does mark areas where a full clearance search is more likely to come back clean than in core compaction or LSM indexing claims.
Go past this page: query the whole storage engines and data layout corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company's registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.