Book a demo

Vector Database Patents: Top Companies & Filing Trends 2026

Vector Database Patents: Top Companies & Filing Trends 2026
https://www.patsnap.com/resources/blog/rd-blog/vector-databases-and-similarity-search-patent-landscape/ · Patsnap · data cut-off 2026-07-31 · downloaded from the live page
Patent Landscape · Database Systems
Vector Database and Similarity Search Patents: Who Is Filing, and Where the Claims Sit
  • Zero to 38 in eight years. Filings sat at zero as recently as 2022 and reached a peak of 38 in 2025, meaning almost the entire patent record for this field was created in the last few years.
  • China leads the receiving offices. China accounts for 32 of the tracked filings against 26 for the United States, with the most-cited Chinese filings tied to large-model retrieval-augmented generation rather than standalone indexing.
  • No single company holds the field. Several previously active assignees — including major technology and hardware firms — show a full drop to zero filings in the latest year, while smaller entrants post sharp year-over-year gains from a low base.
Get a prior-art report on your approach
72
Published Records
25%
Top-5 Share of All Records
CN
Leading Jurisdiction
61
Active Filers Ranked
Published byPatsnap Research··8 min readSourced from Patsnap Eureka
Overview

A young, fast-accelerating field of claims

Patent filings covering vector databases and approximate nearest neighbor search were essentially nonexistent before the last few years: the dataset shows zero filings as recently as 2022 and a peak of 38 in 2025. Most of what exists today sits inside the G06F electric-digital-data-processing classification, with only a minority also tagged to AI-model classifications, which tells you the field is still framed and claimed primarily as an indexing and data-structure problem rather than as an applied-AI technique.

China and the United States account for the bulk of receiving-office activity, with the most-cited individual records split between a foundational graph-index method and newer large-model-driven retrieval-augmented generation filings. Because publication lags filing by roughly 18 months, the most recent year in any trend understates real filing activity, and 2026 figures in particular should be read as a floor, not a ceiling.

Filing volume by year, 2017–2026
  1. 1ORACLE INT CORP5
  2. 2D NOTITIA INC4
  3. 3KIOXIA CORP3
  4. 4COACTIVE SYSTEMS INC3
  5. 5VELLORE INSITUTE OF TECH3
  6. 6INSPUR SUZHOU INTELLIGENT TECH CO LTD2
  7. 7SHENZHEN GREEN CONNECTION TECH CO LTD2
  8. 8BEIJING ELECTRONICS SCI & TECH INST2
  9. 9GOOGLE LLC2
  10. 10HANGZHOU DIANZI UNIV2
Source: Patsnap Eureka. Assignee ranking and totals. Derived from a Patsnap search on Vector Databases and Similarity Search covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Filing Data

Filing trends and technology composition

Filings in this dataset start from zero in 2017 and climb sharply toward a 2025 peak, with 2026 numbers still incomplete because publication trails filing by roughly a year and a half. The technology composition confirms this is overwhelmingly a data-processing story, with artificial-intelligence classifications a distant second.

A late, steep filing curve

With 2022 still at zero and 2025 reaching 38, essentially all activity in this dataset is concentrated in the last few years, and the curve has not yet shown signs of leveling off.

A late, steep filing curve01020304002017201820192020202120222023202438202552026Most recent year is partial — publication lag means later filings are not yet visible.

G06F dominates the classification mix

70 of 72 records touch G06F (electric digital data processing), while G06N (AI models) appears in 20 and every other subclass — image recognition, business processing, healthcare informatics — sits in single digits, showing the field is still framed as a data-structure problem more than an applied-AI one.

G06F dominates the classification mixG06F · Electric digital data processi…7097.2%G06N · Computing based on AI models2027.8%G06V · Image/video recognition56.9%G06Q · Business, commerce & admin dat…45.6%G06K · Data recognition & presentation11.4%G06T · Image data processing & genera…11.4%G16H · Healthcare informatics11.4%H04L · Digital information transmissi…11.4%

Shares are the percentage of the 72 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.

Source: Patsnap Eureka. Filing trend and technology composition. Derived from a Patsnap search on Vector Databases and Similarity Search covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.

Go deeper on Vector Databases and Similarity Search with Eureka

This page is one run against one query. Ask Eureka your own question about vector databases and similarity search and every answer comes back with the patent numbers behind it.

Try Eureka
Key Patents

The most-cited and most recent filings

Representative Filing
US20260154293A12026-06-04

Approximate nearest neighbor search method and approximate nearest neighbor search system

KIOXIA CORPORATION

According to one embodiment, an approximate nearest neighbor search method manages graph-based index information for defining an inter-cluster graph. The approximate nearest neighbor search method searches for a vector closest to a query vector from vectors belonging to a search start cluster that is closest to the query vector among a plurality of clusters. The approximate nearest neighbor search method selects one or more search target clusters close to the search start cluster while traversing the inter-cluster graph, and searches for a vector closest to the query vector from vectors belonging to each of the one or more search target clusters.Published 2026-06-04 by Kioxia Corporation; combines cluster routing with inter-cluster graph traversal, sitting downstream of earlier hierarchical navigable small world claims.

US20260154293A1 — patent drawing 1US20260154293A1 — patent drawing 2
View full filing
Most-cited records in this dataset
#Publication no.Patent titleCitations
1CN119988588A一种基于大模型的多模态文档检索增强生成方法54
2US20020178158A1Vector index preparing method, similar vector searching method, and apparatuses for the methods34
3CN110008256A一种基于分层可导航小世界图的近似最近邻搜索方法29
4CN118093962A数据检索方法、装置、系统、电子设备及可读存储介质23
5CN118069812A一种基于大模型的导览方法21
6US20140086492A1Storage method and storage device for database for approximate nearest neighbor search16
7US20250094400A1Deploying a vector index on multiple nodes of a cluster9
8CN117033394A一种大语言模型驱动的向量数据库构建方法及系统8
9CN110866024A一种矢量数据库增量更新方法及系统8
10US20250217428A1Web Browser with Integrated Vector Database6

Citation counts are drawn from the searched corpus only and favour older filings; treat them as a measure of influence within this dataset, not as a ranking of current commercial importance.

Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.

Source: Patsnap Eureka. Citation counts and representative records. Derived from a Patsnap search on Vector Databases and Similarity Search covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Run it yourself

Put your own technology through the same analysis

 
Where to run it
Fastest

Eureka on the web

When you want the answer in the next five minutes.

The agent works the prompt against patents and technical literature, citing every source.

Run your analysis now →
For builders

MCP server & REST API

When it has to run inside your own pipeline.

Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.

Browse MCP servers →
Signals

What the citation and filing patterns show

Citation counts and filing timing point in different directions here: the most-cited prior art is several years old, while the filing volume is almost entirely recent, which is typical of a field where foundational methods are settled but applied claims are still being staked out.

Citation concentration
54 citations
top-cited record

A multimodal retrieval method leads citations

The most-cited record in this dataset is a large-model-based multimodal document retrieval and generation method, ahead of an older vector-index-preparation patent at 34 citations and a hierarchical navigable small world method at 29. High citation counts here favour earlier filings simply because they have had more time to be cited, not because they are more relevant today.

Read as influence, not current activity.
Filing acceleration
0 → 38
2022 to 2025 filings

Growth is still accelerating, not leveling

The filing trend shows zero activity at the 2022 midpoint and a peak of 38 filings in 2025, with 2026 numbers still partial because publication typically lags filing by around 18 months. There is no sign yet of the curve flattening.

Expect the true 2026 total to be revised upward.
Classification concentration
70 of 72
records tagged G06F

Data-structure framing dominates over AI framing

Only 20 of 72 records also carry a G06N (AI model) classification, and every other subclass — image recognition, commerce, healthcare — sits in single digits. The field is still being claimed primarily as an indexing and data-processing problem, even though most commercial use cases sit downstream of large AI models.

Applied-AI framing of these claims remains comparatively rare.
Eureka AI Agent
Looking for what nobody has claimed yet?

Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to vector databases and similarity search, with the prior art for and against each one.

Find the white space →
Source: Patsnap Eureka. Co-assignee relationships and derived observations. Derived from a Patsnap search on Vector Databases and Similarity Search covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Players

Who is filing, and where the momentum sits

No single assignee dominates this dataset by volume, and the momentum data shows real turnover: firms with prior filing history have gone quiet in the latest year while newer, smaller filers post sharp percentage gains from a low base. Co-assignee filing is nearly absent, with only one weak cross-institutional pair recorded.

Momentum shift
+100% YoY
Vellore Institute of Technology

Academic filers gaining share from a low base

Vellore Institute of Technology doubled its filings year-over-year, a pattern typical of academic and research-institute entrants stepping into a field where large-company filing has recently slowed.

Small absolute numbers, but directionally notable.
Momentum shift
-100% YoY
multiple prior filers

Several established filers dropped to zero

Assignees with prior filing history — including large technology and memory-hardware companies — recorded zero filings in the latest year, a swing large enough to reopen claim space that looked settled a year earlier.

Zero-filing years do not necessarily signal exit; check for stealth continuations.
Collaboration density
1 pair
strongest co-assignee link

Cross-institutional filing is nearly nonexistent

Only one co-assignee pairing appears in the dataset, at the minimum observed strength, indicating that most filings here are still produced by single organisations rather than joint ventures or university-industry partnerships.

Joint-development filing strategies remain largely untested in this space.
🔍
Under-claimed sub-areas worth a first-mover claim
These branches are named in the underlying search criteria but are thin relative to core index-structure filings.
Filtered search with live metadata constraintsIncremental index updates without full reindexMemory-footprint-bounded graph indexesRecall-latency tradeoff tuning at query timeDomain-specific filter-and-index fusion (healthcare, finance)
Rank all filers by momentum →
Recent-year filing momentum by assignee
AssigneeRecent yearYoY
VELLORE INSITUTE OF TECH2+100%
Oracle International Corporation0-100%
D NOTITIA INC0-100%
Kioxia Corporation0-100%
COACTIVE SYSTEMS INC0-100%
Google LLC0
Suzhou Yuannao Intelligent Technology Co., Ltd.0-100%
Shenzhen UGREEN Technology Co., Ltd.0
Source: Patsnap Eureka. Assignee-level momentum. Derived from a Patsnap search on Vector Databases and Similarity Search covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
What's Next

Where to take this analysis

The filing curve is still climbing and the leaderboard is still turning over, so a single snapshot understates how fast this space is moving. The next steps depend on whether you are trying to file, license, or simply track competitors.

Check freedom to operate before filing

Run the specific cluster-and-graph-traversal sequence, or any filtered-search claim, against the most-cited records in this dataset before drafting, since the densest claim territory sits around graph-based indexing.

Check claims in Eureka

Track assignee momentum, not just totals

Several long-standing filers dropped to zero in the latest year while smaller entrants gained; a static ranking will miss this turnover, so momentum should be revisited at least twice a year.

Set up assignee tracking in Eureka

Watch the white space in filtered and incremental search

Filtered search and incremental index updates are named in the underlying search criteria but thin in the claim record, making them a lower-risk area for a first claim relative to core graph-index territory.

Explore white space in Eureka
Source: Patsnap Eureka. Forward-looking reading of the same dataset. Derived from a Patsnap search on Vector Databases and Similarity Search covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
FAQ

Common questions about vector database patents

Answers are grounded in the same dataset. Derived from a Patsnap search on Vector Databases and Similarity Search covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Research Vector Databases and Similarity Search in depth with Eureka

Go past this page: query the whole vector databases and similarity search corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.

Try Eureka

Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.

Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.

Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.

Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.

Help us improve this page

Found incorrect or outdated information? Let us know and we'll get it fixed.