RAG Foundation Model Patents: Who Leads, Where Gaps Are 2026
- Filings surged 281% from 2021 to 2024 (851 to 3,240), the clearest sign that RAG moved from research technique to a defended commercial architecture inside three years.
- The ranked leaders are diffuse: the top 5 assignees hold just 9.9% of all 838,520 records in scope, and the top 10 only 14.3% — no single filer controls the core architecture.
- G06F carries the claim density, at 1.8% of all records, while adjacent classes like G10L (speech) and G06T (image generation) sit near 0.1% — largely open ground for retrieval fused with other modalities.
Filing growth compares 2021 (851 records) with 2024 (3,240) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 838,520 records in scope (CR5), not by the ranked leaders only.
What the RAG patent record actually shows
Retrieval-Augmented Generation started as a way to ground large language model outputs in external documents, and the patent record now spans 838,520 records filed between 2015 and mid-2026. The search string captures both explicit RAG terminology and the broader pattern of retrieval combined with generation over context, knowledge or documents, so the corpus includes both narrowly-labelled RAG filings and the wider retrieval-and-generation infrastructure that underpins it. Publication lags filing by roughly 18 months, so filing activity in 2025 and 2026 will keep revising upward as those applications publish.
Filing activity accelerated sharply once RAG moved from an academic technique to a production requirement for enterprise LLM deployments: 2021 filings sat at 851, rising to 3,240 by 2024, a complete year in this dataset. The assignee base is wide rather than concentrated, and the IPC composition shows most claim density sitting in general data-processing classes rather than in AI-model-specific ones, which points to RAG being claimed as infrastructure and system architecture as much as as a modelling technique.
Filing trends and technology composition
Two views of the same corpus: how filing volume has moved year over year, and which IPC subclasses carry the claim density once a record can sit in more than one class.
A three-year filing surge, now cresting
Filings rose from 508 in 2017 to a peak of 3,300 in 2025, with the sharpest acceleration between 2021 and 2024 — an increase of 281% across just three years. 2026 figures (408 so far) are partial and will rise materially as later-filed applications publish; they should not be read as a decline.
Claim density concentrates in general data processing, not AI-specific classes
G06F (electric digital data processing) covers 1.8% of all 838,520 records, more than three times the next-largest class, G06N (AI-model computing) at 0.7%. Classes tied to specific modalities — G06V image recognition, G10L speech, G06T image generation — each sit at roughly 0.1% of records, suggesting RAG's system and pipeline claims are far more crowded than its multimodal extensions.
Shares are the percentage of the 838,520 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Foundation Models: Retrieval-Augmented Generation Patent Landscape with Eureka
This page is one run against one query. Ask Eureka your own question about foundation models: retrieval-augmented generation patent landscape and every answer comes back with the patent numbers behind it.
Try EurekaRepresentative filing and most-cited prior art
Method and system for improving retrieval accuracy in retrieval augmented generation (RAG) framework
Filed by LTI Mindtree and published 2026-07-02, this application describes a pipeline that analyses input documents with a small language model to classify document type before extracting content with type-specific methods, then runs adaptive chunking driven by use case, latency, cost and the target LLM's context window size. Tokenization strategy and embedding model selection are both made configurable against accuracy and vocabulary requirements, with quantization applied to the resulting vector representations.The claim scope centres on adaptive, use-case-driven configuration of the retrieval pipeline rather than a single retrieval algorithm — worth reading closely against any product that auto-tunes chunking or embedding choice.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20080120129A1 | Consistent set of interfaces derived from a business object model | 2,470 |
| 2 | US20120069131A1 | Reality alternate | 2,111 |
| 3 | US7181438B1 | Database access system | 1,902 |
| 4 | US20050132070A1 | Data security system and method with editor | 1,899 |
| 5 | US20100070448A1 | System and method for knowledge retrieval, management, delivery and presentation | 1,798 |
| 6 | US5321816A | Local-remote apparatus with specialized image storage modules | 1,789 |
| 7 | US8275836B2 | System and method for supporting collaborative activity | 1,703 |
| 8 | US6026388A | User interface and other enhancements for natural language information retrieval system and method | 1,703 |
| 9 | US20030126136A1 | System and method for knowledge retrieval, management, delivery and presentation | 1,532 |
| 10 | US6675159B1 | Concept-based search and retrieval system | 1,510 |
Citation counts favour older filings simply because they have had more time to accumulate citations within the searched corpus — treat this table as a map of foundational influence, not of current commercial relevance.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the numbers mean for a filing decision
Three read-throughs from the filing trend, the assignee spread and the IPC composition, each with a direct implication for where to file or watch.
The surge is real and recent
Filings nearly quadrupled from 851 in 2021 to 3,240 in 2024, tracking the enterprise adoption curve for LLM-grounding techniques. Because 2025 and 2026 are still publishing, the true current filing rate is higher than the raw counts show.
No single architecture owner
The top 5 assignees hold 9.9% of all 838,520 records and the top 10 just 14.3%, well short of the concentration seen in more mature hardware fields. That leaves substantial room for new entrants to build defensible positions in specific implementation layers.
Infrastructure claims dominate over modality claims
G06F filings outnumber the combined weight of image, speech and video-specific classes by a wide margin, indicating most RAG patent activity targets pipeline and system architecture rather than any single input modality.
Read momentum drops with the publication lag in mind
Several of the most active historical filers show sharp year-over-year drops in the latest year, but this mirrors the roughly 18-month publication lag rather than a real pullback — recent filings from these assignees simply have not published yet.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to foundation models: retrieval-augmented generation patent landscape, with the prior art for and against each one.
Who is filing, and where the gate sits for new entrants
The ranked leaders span telecoms infrastructure, enterprise software and semiconductor firms rather than a cluster of pure-play AI labs — a sign that RAG claims are being built into existing product portfolios more than staked out by specialists.
Breadth over a narrow moat
The leading assignee's count sits well above fifth place (10,890) and tenth (6,618), but even that lead only translates into a fraction of the total corpus, reflecting a portfolio built across many product lines rather than one core RAG patent family.
A real second tier, not a cliff
The drop from leader to fifth place is steep but not a cliff to zero — fifth place still holds over 10,000 records, indicating several large filers rather than one dominant player and a void beneath.
Filing is mostly solo, with a few dense internal pairs
The strongest co-assignee links are internal — a parent company filing jointly with its own regional IP or subsidiary entities — rather than cross-company collaboration, suggesting most RAG IP strategy is managed within single corporate groups.
| Assignee | Recent year | YoY |
|---|---|---|
| Microsoft Technology Licensing, LLC | 8 | -92% |
| Google LLC | 4 | -93% |
| Oracle International Corporation | 4 | -95% |
| NVIDIA Corporation | 4 | -95% |
| Shuo Power Corporation | 2 | -98% |
| SAP SE | 1 | -95% |
| International Business Machines Corporation (IBM) | 0 | -100% |
| Microsoft Corporation | 0 | — |
Where to take this analysis
The dataset points to open claim space in modality fusion and pipeline configuration, and to a filing surge that has not yet fully published. Both are reasons to look closer rather than treat this as a settled field.
Map a specific claim against this corpus
A single competitor filing or draft claim can be run against this same search scope to see how crowded its exact technical branch is, rather than relying on the field-wide averages shown here.
Explore this landscape in Eureka →Track the assignees actually moving right now
Because publication lag understates 2025 and 2026, a live monitoring view catches new filings from key assignees as they publish rather than waiting for the next annual dataset refresh.
Set up assignee tracking in Eureka →Common questions on RAG patent strategy
The tracked corpus for this landscape contains 838,520 published records matching RAG-related search terms between 2015 and mid-2026. This figure covers both filings that explicitly name Retrieval-Augmented Generation and the broader pattern of retrieval combined with generation over context, knowledge or documents, so it is intentionally wider than a search on the exact acronym alone. Because publication lags filing by roughly 18 months, the most recent one to two years of this total will keep rising as later applications publish.
The ranked leaders are led by a single assignee with 20,923 records, with fifth place at 10,890 and tenth place at 6,618, spanning telecoms, enterprise software and semiconductor firms rather than a small set of AI-only specialists. The top 5 assignees together hold only 9.9% of all 838,520 records in scope, and the top 10 hold 14.3%, so no single company controls the core architecture. This spread suggests RAG capability is being built into broader existing product portfolios rather than staked out as a standalone patent moat.
Yes, in specific branches. Filing density concentrates heavily in general data-processing classes like G06F (1.8% of all records) and AI-model classes like G06N (0.7%), while modality-specific areas such as speech (G10L) and image generation (G06T) each sit near 0.1% of records. That gap indicates the core retrieval pipeline is more crowded than multimodal or configuration-specific extensions of it, which is where a first mover currently faces a lighter prior-art field.
Filings rose from 851 in 2021 to 3,240 in 2024, a 281% increase over three complete years, tracking the shift of RAG from an academic grounding technique into a required component of production LLM deployments. Enterprises adopting large language models needed a defensible way to connect models to proprietary or current data, and patent filing followed that commercial need. The 2025 peak of 3,300 and lower 2026 count should be read alongside the roughly 18-month publication lag rather than as a slowdown.
This filing, from LTI Mindtree and published 2026-07-02, claims an adaptive pipeline that classifies input documents with a small language model, chunks content based on use case, latency, cost and the target LLM's context window, and selects tokenization strategy and embedding model according to accuracy and vocabulary needs before applying quantization. Companies with products that automatically tune chunking granularity or embedding choice based on document type or downstream model constraints should review this filing's specific claim language closely. It does not appear to claim a single retrieval algorithm, so the risk is concentrated in configurable, use-case-driven pipeline orchestration rather than retrieval itself.
Research Foundation Models: Retrieval-Augmented Generation Patent Landscape in depth with Eureka
Go past this page: query the whole foundation models: retrieval-augmented generation patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.