Book a demo

Hybrid Retrieval Patents: Who Leads, Where the Gaps Are 2026

Hybrid Retrieval Patents: Who Leads, Where the Gaps Are 2026
https://www.patsnap.com/resources/blog/rd-blog/vector-databases-and-retrieval-hybrid-sparse-dense-retrieval-patent-landscape-patent-landscape/ · Patsnap · data cut-off 2026-07-31 · downloaded from the live page
Patent Landscape · Hybrid Sparse-Dense Retrieval
Hybrid sparse-dense retrieval patents: mapping the vector search and retrieval-augmented generation claim space
  • No dominant filer. the leading assignee holds 25 of 184 records and the top 5 combined account for just 25.0% of all records in scope — this field has not consolidated.
  • Filings are accelerating late. activity climbed from zero in 2017 to a peak of 99 records in 2025, with 2026 already at 60 before publication lag is even accounted for.
  • Claims cluster in general-purpose computing. 87.0% of records sit in G06F and 44.6% in G06N, while healthcare, speech and image-specific applications remain thinly claimed.
Get a prior-art report on your approach
184
Published Records
25%
Top-5 Share of All Records
US
Leading Jurisdiction
100
Active Filers Ranked

Top-5 share is the combined record count of the five largest assignees divided by all 184 records in scope (CR5), not by the ranked leaders only.

Published byPatsnap Research··7 min readSourced from Patsnap Eureka
Overview

What this landscape covers

This landscape tracks 184 published patent families matching hybrid sparse-dense retrieval and hybrid vector search claims layered against vector database, vector search or similarity search terminology. The scope captures systems that combine token-based (sparse) ranking with embedding-based (dense) similarity search, including retrieval-augmented generation pipelines that route queries through a vector database. Coverage runs from 2015-01-01 through the 2026-07-31 data cut-off.

Because publication lags filing by roughly 18 months, the most recent filing year is always undercounted relative to where activity actually stands. Family-level counting is used throughout so that continuation filings and multi-jurisdiction copies of the same invention are not double-counted in the assignee ranking or the trend line.

Filing activity, 2017-2026
  1. 1MADISETTI VIJAY25
  2. 2VELLORE INSITUTE OF TECH7
  3. 3AIRA TECHNOLOGIES INC7
  4. 4ORACLE INT CORP4
  5. 5SYMBIOSIS INTERNATIONAL UNIVERSITY3
  6. 6DR NILESH SABLE2
  7. 7MANIPAL UNIVERSITY JAIPUR2
  8. 8SYMBIOSIS INT DEEMED UNIV2
  9. 9ZSCALER INC2
  10. 10RMINT INC2
Source: Patsnap Eureka. Assignee ranking and totals. Derived from a Patsnap search on Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Let an AI agent run this analysis on your own technology

Pick a task. Every answer cites the patents behind it.

10,000 free credits to start
The data

Filing trend and technology composition

Two views of the same 184 records: how filing activity has moved year over year, and which IPC subclasses carry the claim density.

Filing trend

Filings sat near zero through the late 2010s, then rose sharply, reaching a peak of 99 records in 2025. The 2026 figure of 60 is already substantial despite the year being incomplete and subject to publication lag, suggesting the true 2026 total will exceed 2025 once later filings surface.

Filing trend0255075100020172018201920202021202220232024992025602026Most recent year is partial — publication lag means later filings are not yet visible.

IPC composition

G06F (electric digital data processing) covers 87.0% of the 184 records and G06N (AI-based computing) covers 44.6%, confirming that most claims are written as general computing or machine-learning methods rather than domain-specific applications. G06Q (business/commerce data processing) reaches 18.5%, while healthcare informatics (G16H, 4.9%), image/video recognition (G06V, 3.8%) and speech analysis (G10L, 3.8%) remain minor branches. Because a record can carry multiple IPC classes, these shares sum to more than 100% of the record total.

IPC compositionG06F · Electric digital data processi…16087.0%G06N · Computing based on AI models8244.6%G06Q · Business, commerce & admin dat…3418.5%G16H · Healthcare informatics94.9%G06V · Image/video recognition73.8%G10L · Speech & audio analysis/synthe…73.8%G01N · Material analysis & testing42.2%G06T · Image data processing & genera…42.2%Other3619.6%

Shares are the percentage of the 184 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.

Source: Patsnap Eureka. Filing trend and technology composition. Derived from a Patsnap search on Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.

Go deeper on Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape with Eureka

This page is one run against one query. Ask Eureka your own question about vector databases & retrieval — hybrid sparse–dense retrieval patent landscape and every answer comes back with the patent numbers behind it.

Try Eureka
Key patents

Representative filing and most-cited records

Representative filing
US20260228288A12026-08-06

Hybrid retrieval augmented generation for rich document queries using a large language model

ZOOM COMMUNICATIONS, INC.

The filing describes a computing system that receives a query about documents held in a vector database, generates tokens from the query terms, and produces a probabilistic ranking of documents from those tokens. In parallel, it issues a vector similarity search against an embedded representation of the same query to the vector database and receives a second document ranking back. The system then combines the two rankings to determine a final top result set for the query.Filed by Zoom Communications, Inc., published 2026-08-06.

US20260228288A1 — patent drawing 1US20260228288A1 — patent drawing 2
View full record
Most-cited records in scope
#Publication no.Patent titleCitations
1US20250190460A1Method and System for Optimizing Use of Retrieval Augmented Generation Pipelines in Generative Artificial Int…45
2US20250139140A1Method and system for context-aware telecommunications, cellular, and radio based generative pre-trained tran…20
3US20260050860A1System and method for comprehensive ESG performance management with multi-dimensional business value quantifi…18
4US20260123620A1Laser – based targeting and object detection system15
5CN120448502A基于知识图谱自适应混合检索增强的农业病虫害问答方法15
6US20250094398A1Unified rdbms framework for hybrid vector search on different data types via SQL and nosql14
7US12204524B1Method and system for electronic processing of user queries maintaining factual consistency during processing14
8US20250328560A1Security Methods and Systems for Multi-Agent Generative AI Applications12
9US12405977B1Method and system for optimizing use of retrieval augmented generation pipelines in generative artificial int…12
10US12259913B1Caching large language model (LLM) responses using hybrid retrieval and reciprocal rank fusion12

Citation counts are drawn from within this searched corpus and favour older filings that have had more time to accumulate citations; treat them as a signal of influence rather than current importance.

Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.

Source: Patsnap Eureka. Citation counts and representative records. Derived from a Patsnap search on Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Run it yourself

Put your own technology through the same analysis

 
Where to run it
Fastest

Eureka on the web

When you want the answer in the next five minutes.

The agent works the prompt against patents and technical literature, citing every source.

Run your analysis now →
For builders

MCP server & REST API

When it has to run inside your own pipeline.

Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.

Browse MCP servers →
Insights

What the numbers say about the field

Three read-throughs from the assignee, trend and citation data that matter for a filing or freedom-to-operate decision.

Concentration
25.0%
top 5 share of 184 records

A modest leader, not a gatekeeper

The leading assignee holds 25 of the 184 records in scope, and the top 5 combined reach only 25.0% of all records. The top 10 combined reach 30.4%. No single filer has cornered core hybrid retrieval claims, which leaves room for a well-drafted application to stake out unclaimed combinations rather than having to design around one blocking portfolio.

Based on the 184-record assignee ranking.
Momentum
99 in 2025
peak filing year

Activity is recent and still rising

The trend moved from 0 filings in 2017 to a peak of 99 in 2025, with 2026 already at 60 before the year closes. Most of the field's claim history is less than three years old, so freedom-to-operate searches need to weight recent filings heavily rather than relying on older, more-cited records alone.

Trend covers 2017 through the 2026-07-31 cut-off.
Geography
74 US filings
of 184 records

US and India lead filing offices

The United States (74) and India (63) are the two largest receiving offices, followed by China (24) and WIPO/PCT filings (14). Europe and Australia each carry a small single-digit share, which suggests hybrid retrieval claims filed through European or Australian offices remain comparatively rare relative to demand in this space.

Counts are by receiving office across the 184 records.
Eureka AI Agent
Looking for what nobody has claimed yet?

Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to vector databases & retrieval — hybrid sparse–dense retrieval patent landscape, with the prior art for and against each one.

Find the white space →
Source: Patsnap Eureka. Co-assignee relationships and derived observations. Derived from a Patsnap search on Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Players

Who is filing, and where the field is still open

The ranking returned by the data endpoint covers 100 companies, counted in records — not a top-50 or top-100 cut, but the entirety of what the ranking returns for this search.

Leader
25 records
leading assignee

One filer ahead of a long tail

The top-ranked assignee holds 25 of the 184 records, well ahead of fifth place at 3 and tenth place at 2. That gap between first and the rest of the ranked leaders is the clearest sign of an unsettled field: most participants have filed only a handful of times.

Ranking covers 100 companies by record count.
Momentum
+33% YoY
fastest-growing ranked filer

Academic filers are the ones still accelerating

Among tracked assignees, the fastest year-over-year growth in the latest year belongs to an academic filer, up 33%, while several of the historically larger filers show flat or sharply declining latest-year counts. That pattern points to universities and smaller labs as the ones currently pushing new claims into this space.

Momentum measured year over year on latest complete filing year.
Long tail
30.4%
top 10 share of 184 records

Most of the field files once or twice

With the top 10 assignees combined accounting for only 30.4% of all 184 records, the remaining 90 ranked filers and the unranked entrants collectively hold the majority of the field. This is a landscape shaped by many single- or double-filing entrants rather than a handful of repeat players.

Based on the 184-record assignee ranking.
🔍
Under-claimed sub-areas worth checking before filing
These branches show low record counts relative to the core G06F/G06N claim density, based on the IPC composition above.
Healthcare-informatics-specific hybrid retrievalSpeech and audio query embedding fusionImage/video similarity re-ranking with sparse fallbackMaterial-analysis retrieval pipelinesBusiness-process-specific RAG scoring
Rank all filers by momentum →
Recent-year filing momentum by assignee
AssigneeRecent yearYoY
VELLORE INSITUTE OF TECH4+33%
MADISETTI VIJAY1-95%
AIRA TECHNOLOGIES INC1-50%
Symbiosis International University1-50%
Oracle International Corporation0
INVENTUS HOLDINGS LLC0-100%
Accenture Global Solutions Limited0-100%
International Business Machines Corporation (IBM)0-100%
Source: Patsnap Eureka. Assignee-level momentum. Derived from a Patsnap search on Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
What's next

Where to take this analysis

The dataset points to a field still forming its claim boundaries. Two practical next steps follow from that.

Run a freedom-to-operate check on the current filing

Before drafting a hybrid retrieval or RAG pipeline claim, check it against the recent US and Indian filings driving 2025-2026 volume, not just the older, more-cited records.

Explore with Patsnap Eureka

Draft into the under-claimed branches

Healthcare, speech and image-specific hybrid retrieval remain thin relative to the general-purpose G06F/G06N cluster, which is where a narrowly scoped first claim is more likely to clear prior art.

Explore with Patsnap Eureka
Source: Patsnap Eureka. Forward-looking reading of the same dataset. Derived from a Patsnap search on Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
FAQ

Common questions about hybrid retrieval patents

Answers are grounded in the same dataset. Derived from a Patsnap search on Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Research Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape in depth with Eureka

Go past this page: query the whole vector databases & retrieval — hybrid sparse–dense retrieval patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.

Try Eureka

Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.

Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.

Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.

Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.

Help us improve this page

Found incorrect or outdated information? Let us know and we'll get it fixed.