Vector Database Patents: Top Companies & Filing Trends 2026
- Zero to 38 in eight years. Filings sat at zero as recently as 2022 and reached a peak of 38 in 2025, meaning almost the entire patent record for this field was created in the last few years.
- China leads the receiving offices. China accounts for 32 of the tracked filings against 26 for the United States, with the most-cited Chinese filings tied to large-model retrieval-augmented generation rather than standalone indexing.
- No single company holds the field. Several previously active assignees — including major technology and hardware firms — show a full drop to zero filings in the latest year, while smaller entrants post sharp year-over-year gains from a low base.
A young, fast-accelerating field of claims
Patent filings covering vector databases and approximate nearest neighbor search were essentially nonexistent before the last few years: the dataset shows zero filings as recently as 2022 and a peak of 38 in 2025. Most of what exists today sits inside the G06F electric-digital-data-processing classification, with only a minority also tagged to AI-model classifications, which tells you the field is still framed and claimed primarily as an indexing and data-structure problem rather than as an applied-AI technique.
China and the United States account for the bulk of receiving-office activity, with the most-cited individual records split between a foundational graph-index method and newer large-model-driven retrieval-augmented generation filings. Because publication lags filing by roughly 18 months, the most recent year in any trend understates real filing activity, and 2026 figures in particular should be read as a floor, not a ceiling.
Filing trends and technology composition
Filings in this dataset start from zero in 2017 and climb sharply toward a 2025 peak, with 2026 numbers still incomplete because publication trails filing by roughly a year and a half. The technology composition confirms this is overwhelmingly a data-processing story, with artificial-intelligence classifications a distant second.
A late, steep filing curve
With 2022 still at zero and 2025 reaching 38, essentially all activity in this dataset is concentrated in the last few years, and the curve has not yet shown signs of leveling off.
G06F dominates the classification mix
70 of 72 records touch G06F (electric digital data processing), while G06N (AI models) appears in 20 and every other subclass — image recognition, business processing, healthcare informatics — sits in single digits, showing the field is still framed as a data-structure problem more than an applied-AI one.
Shares are the percentage of the 72 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Vector Databases and Similarity Search with Eureka
This page is one run against one query. Ask Eureka your own question about vector databases and similarity search and every answer comes back with the patent numbers behind it.
Try EurekaThe most-cited and most recent filings
Approximate nearest neighbor search method and approximate nearest neighbor search system
According to one embodiment, an approximate nearest neighbor search method manages graph-based index information for defining an inter-cluster graph. The approximate nearest neighbor search method searches for a vector closest to a query vector from vectors belonging to a search start cluster that is closest to the query vector among a plurality of clusters. The approximate nearest neighbor search method selects one or more search target clusters close to the search start cluster while traversing the inter-cluster graph, and searches for a vector closest to the query vector from vectors belonging to each of the one or more search target clusters.Published 2026-06-04 by Kioxia Corporation; combines cluster routing with inter-cluster graph traversal, sitting downstream of earlier hierarchical navigable small world claims.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | CN119988588A | 一种基于大模型的多模态文档检索增强生成方法 | 54 |
| 2 | US20020178158A1 | Vector index preparing method, similar vector searching method, and apparatuses for the methods | 34 |
| 3 | CN110008256A | 一种基于分层可导航小世界图的近似最近邻搜索方法 | 29 |
| 4 | CN118093962A | 数据检索方法、装置、系统、电子设备及可读存储介质 | 23 |
| 5 | CN118069812A | 一种基于大模型的导览方法 | 21 |
| 6 | US20140086492A1 | Storage method and storage device for database for approximate nearest neighbor search | 16 |
| 7 | US20250094400A1 | Deploying a vector index on multiple nodes of a cluster | 9 |
| 8 | CN117033394A | 一种大语言模型驱动的向量数据库构建方法及系统 | 8 |
| 9 | CN110866024A | 一种矢量数据库增量更新方法及系统 | 8 |
| 10 | US20250217428A1 | Web Browser with Integrated Vector Database | 6 |
Citation counts are drawn from the searched corpus only and favour older filings; treat them as a measure of influence within this dataset, not as a ranking of current commercial importance.
Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the citation and filing patterns show
Citation counts and filing timing point in different directions here: the most-cited prior art is several years old, while the filing volume is almost entirely recent, which is typical of a field where foundational methods are settled but applied claims are still being staked out.
A multimodal retrieval method leads citations
The most-cited record in this dataset is a large-model-based multimodal document retrieval and generation method, ahead of an older vector-index-preparation patent at 34 citations and a hierarchical navigable small world method at 29. High citation counts here favour earlier filings simply because they have had more time to be cited, not because they are more relevant today.
Growth is still accelerating, not leveling
The filing trend shows zero activity at the 2022 midpoint and a peak of 38 filings in 2025, with 2026 numbers still partial because publication typically lags filing by around 18 months. There is no sign yet of the curve flattening.
Data-structure framing dominates over AI framing
Only 20 of 72 records also carry a G06N (AI model) classification, and every other subclass — image recognition, commerce, healthcare — sits in single digits. The field is still being claimed primarily as an indexing and data-processing problem, even though most commercial use cases sit downstream of large AI models.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to vector databases and similarity search, with the prior art for and against each one.
Who is filing, and where the momentum sits
No single assignee dominates this dataset by volume, and the momentum data shows real turnover: firms with prior filing history have gone quiet in the latest year while newer, smaller filers post sharp percentage gains from a low base. Co-assignee filing is nearly absent, with only one weak cross-institutional pair recorded.
Academic filers gaining share from a low base
Vellore Institute of Technology doubled its filings year-over-year, a pattern typical of academic and research-institute entrants stepping into a field where large-company filing has recently slowed.
Several established filers dropped to zero
Assignees with prior filing history — including large technology and memory-hardware companies — recorded zero filings in the latest year, a swing large enough to reopen claim space that looked settled a year earlier.
Cross-institutional filing is nearly nonexistent
Only one co-assignee pairing appears in the dataset, at the minimum observed strength, indicating that most filings here are still produced by single organisations rather than joint ventures or university-industry partnerships.
| Assignee | Recent year | YoY |
|---|---|---|
| VELLORE INSITUTE OF TECH | 2 | +100% |
| Oracle International Corporation | 0 | -100% |
| D NOTITIA INC | 0 | -100% |
| Kioxia Corporation | 0 | -100% |
| COACTIVE SYSTEMS INC | 0 | -100% |
| Google LLC | 0 | — |
| Suzhou Yuannao Intelligent Technology Co., Ltd. | 0 | -100% |
| Shenzhen UGREEN Technology Co., Ltd. | 0 | — |
Where to take this analysis
The filing curve is still climbing and the leaderboard is still turning over, so a single snapshot understates how fast this space is moving. The next steps depend on whether you are trying to file, license, or simply track competitors.
Check freedom to operate before filing
Run the specific cluster-and-graph-traversal sequence, or any filtered-search claim, against the most-cited records in this dataset before drafting, since the densest claim territory sits around graph-based indexing.
Check claims in EurekaTrack assignee momentum, not just totals
Several long-standing filers dropped to zero in the latest year while smaller entrants gained; a static ranking will miss this turnover, so momentum should be revisited at least twice a year.
Set up assignee tracking in EurekaWatch the white space in filtered and incremental search
Filtered search and incremental index updates are named in the underlying search criteria but thin in the claim record, making them a lower-risk area for a first claim relative to core graph-index territory.
Explore white space in EurekaCommon questions about vector database patents
Filing in this dataset is recent and fragmented rather than concentrated in one dominant company. Several assignees that appear in the historical record — including large technology firms and specialist hardware makers — show zero filings in the most recent year, while smaller or newer entrants show sharp year-over-year gains from a low base. This pattern suggests the field has not yet settled around a small number of clear leaders, and rankings from a year ago may not hold going forward.
Graph-based methods, exemplified by the hierarchical navigable small world approach that is the most-cited prior art in this dataset, build a navigable graph over the vector space and traverse it to find near neighbors. Cluster-based methods, such as the newest representative filing in this dataset, first route a query to a nearby cluster and then search within it, sometimes using a graph between clusters rather than between individual vectors. Cluster-first approaches tend to target lower memory footprint and faster search-start latency, while pure graph methods are typically claimed for higher recall at larger scale.
China is the single largest receiving office in this dataset, ahead of the United States, and several of the most-cited records — including a multimodal retrieval-augmented generation method and a data retrieval method and system — originate from Chinese filers. This concentration lines up with heavy commercial deployment of large-model retrieval pipelines in China, where vector search is being claimed as an infrastructure layer underneath generative AI products rather than as a standalone database technology.
US20260154293A1 claims a specific sequence — identifying a search-start cluster, traversing an inter-cluster graph to select further target clusters, then comparing vectors within each selected cluster. It does not automatically block flat graph-only search or cluster-based search that skips the inter-cluster graph step, since those sit closer to earlier, already-cited prior art. Any team building a clustered index should check specifically whether cluster selection in their system uses a graph traversal step, since that is the feature this filing turns on.
Both terms appear directly in the search criteria used to build this dataset, but neither shows up as a large, distinct claim cluster compared to core index-structure filings, which sit overwhelmingly in the G06F16 classification. That gap suggests filtered search and incremental (online) index updates are still relatively open claim territory, even though they are operationally critical for any production vector database that needs to delete, update, or apply metadata filters without full reindexing.
This dataset covers 72 published records, all of which are also counted as 72 distinct patent families in the assignee ranking, meaning there is little duplicate or continuation filing within the set as gathered. Filing activity is heavily weighted toward the most recent years, with 2025 as the peak year so far at 38 filings, so the total family count is likely to keep growing as more recent applications publish.
Research Vector Databases and Similarity Search in depth with Eureka
Go past this page: query the whole vector databases and similarity search corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.