Multimodal Model Patents: Top Companies & Filing Trends 2026
- Filings up 914% in three years. Annual filings went from 51 in 2021 to 517 in 2024, the last year the dataset treats as complete.
- The top 5 hold 28.8% of the field. 459 of 1,592 records in scope sit with the leading five assignees, with a long tail behind them.
- G06F and G06N dominate, but overlap heavily. 42.6% and 40.9% of records respectively carry these classes, meaning most filings sit at the intersection of general computing and AI-model claims rather than in a single lane.
Filing growth compares 2021 (51 records) with 2024 (517) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 1,592 records in scope (CR5), not by the ranked leaders only.
What the dataset covers
This landscape covers 1,592 published records matching multimodal model and cross-modal learning terms combined with core machine-learning claim language such as training dataset, model parameter, feature vector, model inference, prediction model and loss function. The window runs from 2015 through the data cut-off of 31 August 2026, though publication lag means the most recent one to two years are undercounted. Records are drawn primarily through the United States, WIPO’s PCT route and the European Patent Office, with smaller volumes at the Indian, Canadian and Australian offices.
Read the filing trend and the assignee ranking together rather than in isolation. A company can file heavily without holding a proportionate share of the most-cited records, and a technology class can carry a high record count without any single filer controlling it. The sections below separate those two questions: who is filing, and what they are actually claiming.
Filing trend and technology composition
Two views of the same 1,592 records: how filing volume moved year over year, and which IPC subclasses the claims fall into.
From 8 filings in 2017 to a 2025 peak
Annual filings rose from 8 in 2017 to a peak of 664 in 2025, with 2026 (partial at 102) still filling in. The complete-year comparison the dataset supports is 2021 to 2024, where filings went from 51 to 517 — a 914% increase. Treat 2025 and 2026 figures as a floor, not a ceiling, given the roughly 18-month lag between filing and publication.
G06F and G06N carry the field, G06V and G06T follow
G06F (electric digital data processing) appears on 42.6% of records and G06N (AI-model computing) on 40.9%, confirming that most multimodal filings are framed as general computing architecture claims layered with AI-specific model claims rather than one or the other. G06V (image/video recognition, 16.1%) and G06T (image data processing and generation, 11.4%) show where the modality work concentrates, while G16H (healthcare informatics, 9.2%), G06Q (business/commerce, 9.0%), G10L (speech/audio, 6.3%) and H04L (digital transmission, 5.0%) mark smaller but distinct application fronts. Because a single record can carry several classes, these shares sum to well over 100% of the 1,592 records and should not be added together.
Shares are the percentage of the 1,592 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Multimodal Models Patent Landscape with Eureka
This page is one run against one query. Ask Eureka your own question about multimodal models patent landscape and every answer comes back with the patent numbers behind it.
Try EurekaMost-cited records in the corpus
Method and electronic device for learning multimodal model
A method of training a multimodal model, performed by at least one processor, includes acquiring an existing model, which is a pretrained multimodal model, obtaining a training dataset for training a multimodal model, and generating a new model by training the existing model based on the training dataset, wherein the training dataset includes true response data and false response data.Filed by Sionic AI, this June 2026 application illustrates a common current pattern: fine-tuning a pretrained multimodal model against a training dataset that explicitly labels true and false response pairs, rather than training a model from scratch.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20240386015A1 | Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search | 374 |
| 2 | US20240412720A1 | Real-time contextually aware artificial intelligence (AI) assistant system and a method for providing a conte… | 233 |
| 3 | US20240046318A1 | Social network with network-based rewards | 157 |
| 4 | US20170098153A1 | Intelligent image captioning | 124 |
| 5 | US20180189572A1 | Method and System for Multi-Modal Fusion Model | 115 |
| 6 | US20170147910A1 | Systems and methods for fast novel visual concept learning from sentence descriptions of images | 80 |
| 7 | US20170193545A1 | Filtering machine for sponsored content | 66 |
| 8 | US20200027557A1 | Multimodal modeling systems and methods for predicting and managing dementia risk for individuals | 57 |
| 9 | WO2018124309A1 | Method and system for multi-modal fusion model | 53 |
| 10 | US20250259075A1 | Advanced model management platform for optimizing and securing ai systems including large language models | 51 |
Citation counts inside a searched corpus favour older filings that have had more time to accumulate references — read them as a signal of influence on the field, not of current commercial importance.
Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the concentration and citation data actually show
Three findings that shape where a freedom-to-operate search or a new filing strategy should start.
The top of the field is concentrated, the rest is a long tail
459 of 1,592 records sit with the five leading assignees, rising to 590 (37.1%) across the top ten. Below that the ranking spreads across 100 companies, each holding a modest slice — the tenth-placed assignee holds only 25 records against the leader's 177.
Filing volume accelerated sharply through 2024
Annual filings moved from 51 in 2021 to 517 in 2024, the last year the dataset can treat as complete. That trajectory, not the partial 2025-2026 counts, is the reliable growth signal in this dataset.
Leading filers show a pullback in the most recent year, not a reversal
Every top assignee with recent-year data shows a year-on-year decline in the latest year, from -40% to -100%. Given the roughly 18-month publication lag, this reads as an artefact of incomplete recent-year data rather than a genuine slowdown in filing activity.
The most-cited records skew toward earlier, broader claims
The highest-cited records in the corpus include filings on composite symbolic/non-symbolic reasoning, contextual AI assistants, and earlier work on image captioning and multi-modal fusion dating back to 2017-2018. Their citation counts reflect years of accumulated references more than current relevance.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to multimodal models patent landscape, with the prior art for and against each one.
Who is filing, and where the claim space still has room
The ranked leaders span large platform companies and specialised AI labs; the co-assignee pairs suggest close internal collaboration rather than cross-company partnerships.
One filer sits well ahead of the field
The leading assignee holds 177 records against 32 at fifth place and 25 at tenth — a steep drop-off that marks this as a leader-plus-long-tail field rather than an evenly matched contest among a handful of firms.
Co-filing is limited and mostly internal
Only 10 co-assignee pairs appear in the dataset, and the strongest link — 15 shared records — sits between two entities under common corporate ownership rather than between competitors.
Recent-year counts are still filling in for every major filer
Even the most active recent filer shows only 19 records in the latest year, well down on prior-year pace. Given publication lag, this understates true recent filing activity across the ranked leaders.
| Assignee | Recent year | YoY |
|---|---|---|
| GDM Holding LLC | 19 | -74% |
| Google LLC | 8 | -87% |
| ANTHROPIC PBC | 3 | -40% |
| Oracle International Corp | 2 | -94% |
| Microsoft Technology Licensing, LLC | 0 | -100% |
| DeepMind Technologies Limited | 0 | -100% |
| Samsung Electronics Co., Ltd. | 0 | -100% |
| Qualcomm Incorporated | 0 | -100% |
Where to take this analysis
The record-level data behind this landscape supports closer work than a summary page can show.
Run a freedom-to-operate check against the leading assignees
With 28.8% of records held by five companies, a targeted search against their specific claim language is more efficient than a broad keyword sweep of the full corpus.
Open EurekaTrack the under-claimed branches before they fill in
Sub-areas like healthcare-specific multimodal inference and speech-to-vision alignment show lower record density than the core computing classes — a signal worth monitoring as filing volume continues to grow.
Open EurekaCommon questions on multimodal model patents
This dataset identifies 1,592 published records matching multimodal model and cross-modal learning terms combined with core machine-learning claim language, covering filings from 2015 through the August 2026 data cut-off. The true current total is higher than this because publication typically lags filing by around 18 months, so recent filings are still working their way into the public record. Family-level counts, used here, are a fairer measure than raw document counts because they neutralise continuation filings and multi-jurisdiction duplicates.
Filing is concentrated at the top: the leading assignee holds 177 records, and the top five combined hold 459 records, or 28.8% of the 1,592 records in scope. Beyond the top ten, which together hold 37.1% of the field, the ranking spreads across 100 companies each with a modest share. This pattern — a clear leader, a handful of strong followers, then a long tail — is typical of a fast-growing but not yet consolidated technology area.
Yes, sharply, through the last complete year the dataset supports: filings rose from 51 in 2021 to 517 in 2024, a 914% increase. Figures for 2025 and 2026 appear lower, but that reflects publication lag rather than an actual drop in filing activity, since it takes roughly 18 months for a filed application to publish. Any claim that the field is slowing should be checked against the 2021-2024 trend rather than the partial recent years.
The two largest IPC classes are G06F (electric digital data processing, 42.6% of records) and G06N (computing based on AI models, 40.9%), reflecting that most filings combine general computing architecture claims with AI-model-specific claims. Behind those, G06V (image/video recognition, 16.1%) and G06T (image data processing and generation, 11.4%) capture the visual-modality work, while G16H (healthcare informatics), G06Q (business/commerce) and G10L (speech/audio) mark smaller application-specific fronts. Because records often carry multiple classes, these percentages overlap rather than summing to a whole.
The technology composition data points to several sub-areas with comparatively low record density relative to the dominant G06F and G06N classes, including cross-modal loss function design, speech-to-vision feature alignment, and business-process-specific multimodal fusion. Healthcare informatics applications of multimodal inference also show a smaller footprint (9.2% of records) than the core computing classes despite active interest in the space. These are not guarantees of allowability, but they are areas where the existing claim density is lower and a well-drafted first claim has more room to stand apart from prior art.
Research Multimodal Models Patent Landscape in depth with Eureka
Go past this page: query the whole multimodal models patent landscape corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.