Hybrid Retrieval Patents: Who Leads, Where the Gaps Are 2026
- No dominant filer. the leading assignee holds 25 of 184 records and the top 5 combined account for just 25.0% of all records in scope — this field has not consolidated.
- Filings are accelerating late. activity climbed from zero in 2017 to a peak of 99 records in 2025, with 2026 already at 60 before publication lag is even accounted for.
- Claims cluster in general-purpose computing. 87.0% of records sit in G06F and 44.6% in G06N, while healthcare, speech and image-specific applications remain thinly claimed.
Top-5 share is the combined record count of the five largest assignees divided by all 184 records in scope (CR5), not by the ranked leaders only.
What this landscape covers
This landscape tracks 184 published patent families matching hybrid sparse-dense retrieval and hybrid vector search claims layered against vector database, vector search or similarity search terminology. The scope captures systems that combine token-based (sparse) ranking with embedding-based (dense) similarity search, including retrieval-augmented generation pipelines that route queries through a vector database. Coverage runs from 2015-01-01 through the 2026-07-31 data cut-off.
Because publication lags filing by roughly 18 months, the most recent filing year is always undercounted relative to where activity actually stands. Family-level counting is used throughout so that continuation filings and multi-jurisdiction copies of the same invention are not double-counted in the assignee ranking or the trend line.
Let an AI agent run this analysis on your own technology
Pick a task. Every answer cites the patents behind it.
Filing trend and technology composition
Two views of the same 184 records: how filing activity has moved year over year, and which IPC subclasses carry the claim density.
Filing trend
Filings sat near zero through the late 2010s, then rose sharply, reaching a peak of 99 records in 2025. The 2026 figure of 60 is already substantial despite the year being incomplete and subject to publication lag, suggesting the true 2026 total will exceed 2025 once later filings surface.
IPC composition
G06F (electric digital data processing) covers 87.0% of the 184 records and G06N (AI-based computing) covers 44.6%, confirming that most claims are written as general computing or machine-learning methods rather than domain-specific applications. G06Q (business/commerce data processing) reaches 18.5%, while healthcare informatics (G16H, 4.9%), image/video recognition (G06V, 3.8%) and speech analysis (G10L, 3.8%) remain minor branches. Because a record can carry multiple IPC classes, these shares sum to more than 100% of the record total.
Shares are the percentage of the 184 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape with Eureka
This page is one run against one query. Ask Eureka your own question about vector databases & retrieval — hybrid sparse–dense retrieval patent landscape and every answer comes back with the patent numbers behind it.
Try EurekaRepresentative filing and most-cited records
Hybrid retrieval augmented generation for rich document queries using a large language model
The filing describes a computing system that receives a query about documents held in a vector database, generates tokens from the query terms, and produces a probabilistic ranking of documents from those tokens. In parallel, it issues a vector similarity search against an embedded representation of the same query to the vector database and receives a second document ranking back. The system then combines the two rankings to determine a final top result set for the query.Filed by Zoom Communications, Inc., published 2026-08-06.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20250190460A1 | Method and System for Optimizing Use of Retrieval Augmented Generation Pipelines in Generative Artificial Int… | 45 |
| 2 | US20250139140A1 | Method and system for context-aware telecommunications, cellular, and radio based generative pre-trained tran… | 20 |
| 3 | US20260050860A1 | System and method for comprehensive ESG performance management with multi-dimensional business value quantifi… | 18 |
| 4 | US20260123620A1 | Laser – based targeting and object detection system | 15 |
| 5 | CN120448502A | 基于知识图谱自适应混合检索增强的农业病虫害问答方法 | 15 |
| 6 | US20250094398A1 | Unified rdbms framework for hybrid vector search on different data types via SQL and nosql | 14 |
| 7 | US12204524B1 | Method and system for electronic processing of user queries maintaining factual consistency during processing | 14 |
| 8 | US20250328560A1 | Security Methods and Systems for Multi-Agent Generative AI Applications | 12 |
| 9 | US12405977B1 | Method and system for optimizing use of retrieval augmented generation pipelines in generative artificial int… | 12 |
| 10 | US12259913B1 | Caching large language model (LLM) responses using hybrid retrieval and reciprocal rank fusion | 12 |
Citation counts are drawn from within this searched corpus and favour older filings that have had more time to accumulate citations; treat them as a signal of influence rather than current importance.
Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the numbers say about the field
Three read-throughs from the assignee, trend and citation data that matter for a filing or freedom-to-operate decision.
A modest leader, not a gatekeeper
The leading assignee holds 25 of the 184 records in scope, and the top 5 combined reach only 25.0% of all records. The top 10 combined reach 30.4%. No single filer has cornered core hybrid retrieval claims, which leaves room for a well-drafted application to stake out unclaimed combinations rather than having to design around one blocking portfolio.
Activity is recent and still rising
The trend moved from 0 filings in 2017 to a peak of 99 in 2025, with 2026 already at 60 before the year closes. Most of the field's claim history is less than three years old, so freedom-to-operate searches need to weight recent filings heavily rather than relying on older, more-cited records alone.
US and India lead filing offices
The United States (74) and India (63) are the two largest receiving offices, followed by China (24) and WIPO/PCT filings (14). Europe and Australia each carry a small single-digit share, which suggests hybrid retrieval claims filed through European or Australian offices remain comparatively rare relative to demand in this space.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to vector databases & retrieval — hybrid sparse–dense retrieval patent landscape, with the prior art for and against each one.
Who is filing, and where the field is still open
The ranking returned by the data endpoint covers 100 companies, counted in records — not a top-50 or top-100 cut, but the entirety of what the ranking returns for this search.
One filer ahead of a long tail
The top-ranked assignee holds 25 of the 184 records, well ahead of fifth place at 3 and tenth place at 2. That gap between first and the rest of the ranked leaders is the clearest sign of an unsettled field: most participants have filed only a handful of times.
Academic filers are the ones still accelerating
Among tracked assignees, the fastest year-over-year growth in the latest year belongs to an academic filer, up 33%, while several of the historically larger filers show flat or sharply declining latest-year counts. That pattern points to universities and smaller labs as the ones currently pushing new claims into this space.
Most of the field files once or twice
With the top 10 assignees combined accounting for only 30.4% of all 184 records, the remaining 90 ranked filers and the unranked entrants collectively hold the majority of the field. This is a landscape shaped by many single- or double-filing entrants rather than a handful of repeat players.
| Assignee | Recent year | YoY |
|---|---|---|
| VELLORE INSITUTE OF TECH | 4 | +33% |
| MADISETTI VIJAY | 1 | -95% |
| AIRA TECHNOLOGIES INC | 1 | -50% |
| Symbiosis International University | 1 | -50% |
| Oracle International Corporation | 0 | — |
| INVENTUS HOLDINGS LLC | 0 | -100% |
| Accenture Global Solutions Limited | 0 | -100% |
| International Business Machines Corporation (IBM) | 0 | -100% |
Where to take this analysis
The dataset points to a field still forming its claim boundaries. Two practical next steps follow from that.
Run a freedom-to-operate check on the current filing
Before drafting a hybrid retrieval or RAG pipeline claim, check it against the recent US and Indian filings driving 2025-2026 volume, not just the older, more-cited records.
Explore with Patsnap EurekaDraft into the under-claimed branches
Healthcare, speech and image-specific hybrid retrieval remain thin relative to the general-purpose G06F/G06N cluster, which is where a narrowly scoped first claim is more likely to clear prior art.
Explore with Patsnap EurekaCommon questions about hybrid retrieval patents
This landscape identifies 184 published patent families matching hybrid sparse-dense retrieval, hybrid vector search or dense-sparse fusion claims combined with vector database or similarity search terminology, covering filings from 2015 through the 2026-07-31 data cut-off. That count reflects one specific search string and scope, so a broader query on adjacent embedding or ranking terminology would return a different total. Because publication lags filing by roughly 18 months, the true count for the most recent year is understated at any point in time.
The field is not consolidated: the leading assignee holds 25 of the 184 records, and the top 5 assignees combined account for only 25.0% of all records in scope. The top 10 combined reach 30.4%. That leaves the large majority of filings spread across a long tail of companies and universities filing once or twice, rather than concentrated in a handful of dominant portfolios.
The United States receives the most filings at 74 of the 184 records, followed by India at 63 and China at 24. WIPO PCT filings account for 14 records, with Europe and Australia each in the low single digits. The heavy US and India presence suggests the commercial and research pressure behind these filings is currently strongest in those two jurisdictions.
Most records sit in G06F, electric digital data processing, at 87.0% of the 184 records, and G06N, AI-based computing, at 44.6%. G06Q, covering business and commerce data processing, appears in 18.5% of records. Healthcare informatics, image/video recognition and speech analysis each account for under 5% of records, indicating the technology is claimed mostly as general-purpose computing methods rather than tied to a specific application domain.
Filing activity has risen sharply from zero records in 2017 to a peak of 99 in 2025, and 2026 already shows 60 records before the year is complete. A precise growth rate cannot be calculated because fewer than four complete years of data remain once publication lag is accounted for, but the trajectory of the raw counts points to a field still in an active filing phase rather than a mature or declining one.
Research Vector Databases & Retrieval — Hybrid Sparse–Dense Retrieval Patent Landscape in depth with Eureka
Go past this page: query the whole vector databases & retrieval — hybrid sparse–dense retrieval patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.