Model Compression Patents: Who Leads, Where the Gaps Are 2026
- Filing is still accelerating, with the count climbing from zero in 2017 to a 2025 peak of 19 filings and no sign of plateauing through the 2022 midpoint.
- India and China lead filing volume, with 29 and 24 receiving-office filings respectively, ahead of the United States at 18 — a filing geography that does not match where the most-cited patents were filed.
- Co-assignment is almost nonexistent, with only two multi-assignee pairs across 91 families, pointing to a field still dominated by single-entity filings rather than joint ventures or research consortia.
What this landscape covers
Model compression and knowledge distillation patents cover the techniques used to shrink large neural networks into forms that run on constrained hardware without giving up too much accuracy: teacher-student training, structured sparsity, quantization-aware training, and pruning for edge deployment. The search underlying this page pulls 91 patent families filed between 2015 and mid-2026, filtered to records that explicitly claim accuracy retention, structured sparsity, quantization-aware training, teacher-student architectures, or edge deployment, and classified under the AI-model and general computing IPC groups.
Because publication typically lags filing by around 18 months, the 2025 and 2026 figures in this dataset are undercounts of the filings that have actually happened — the true recent-year volume is higher than what has published so far.
Filing trend and technology composition
The filing curve and the IPC mix together show a field concentrated in AI-model classes but reaching into image recognition and general data processing as compression techniques get applied to specific deployment targets.
Filings, 2017–2026
Filings moved from zero in 2017 to a midpoint of 6 in 2022 and a peak of 19 in 2025; the 2026 figure of 9 is a partial year and will revise upward as later filings publish.
IPC subclass composition
Every record in this set sits in G06N (AI-model computing), with substantial overlap into G06V (image/video recognition, 21 records), G06F (general digital data processing, 19), and G06K (data recognition, 15) — evidence that most compression claims are written for a specific vision or recognition workload rather than as generic architecture claims.
Shares are the percentage of the 91 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Model Compression and Knowledge Distillation with Eureka
This page is one run against one query. Ask Eureka your own question about model compression and knowledge distillation and every answer comes back with the patent numbers behind it.
Try EurekaMost-cited records and a representative filing
Method for object detection using knowledge distillation (US20200293903A1, Cortica Ltd., 2020-09-17)
The patent describes training a student object-detection neural network (ODNN) to mimic a teacher ODNN by calculating a teacher-student detection loss based on the teacher's pre-bounding-box output, itself a function of several component ODNNs inside the teacher. The trained student network then detects objects in an image directly, producing its own pre-bounding-box output and bounding boxes without needing the full teacher network at inference time.The claims sit specifically on pre-bounding-box loss calculation for object detection, not on teacher-student distillation generally.


| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20200302295A1 | System and method for knowledge distillation between neural networks | 139 |
| 2 | US20200387782A1 | Sparsity constraints and knowledge distillation based learning of sparser and compressed neural networks | 61 |
| 3 | CN114037844A | 基于滤波器特征图的全局秩感知神经网络模型压缩方法 | 37 |
| 4 | CA3076424A1 | System and method for knowledge distillation between neural networks | 28 |
| 5 | CN112183670A | 一种基于知识蒸馏的少样本虚假新闻检测方法 | 26 |
| 6 | US20210073643A1 | Neural network pruning | 25 |
| 7 | US20220004803A1 | Semantic relation preserving knowledge distillation for image-to-image translation | 23 |
| 8 | CA3056098A1 | Sparsity constraints and knowledge distillation based learning of sparser and compressed neural networks | 20 |
| 9 | CN113705317A | 图像处理模型训练方法、图像处理方法及相关设备 | 19 |
| 10 | US20210256383A1 | Computer-implemented methods and systems for privacy-preserving deep neural network model compression | 18 |
Citation counts favour older records simply because they have had more time to accumulate citations within the searched corpus — treat this as a signal of influence, not of current commercial importance.
Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the numbers mean for filing strategy
Three patterns stand out once the raw counts are read against each other: where the volume sits, where the influence sits, and how thin collaboration is.
Volume leadership sits outside the US
India and China together account for more receiving-office filings than the United States, but the most-cited records in the set were filed through the US and Canadian offices — volume and influence are not concentrated in the same place.
The curve is still climbing, not flattening
Filings grew from zero in 2017 through a 2022 midpoint of 6 to a 2025 peak of 19. That trajectory, plus the understatement built into any 2025-2026 count from publication lag, means the technology is in an active build-out phase rather than a mature one.
Almost no joint filing
Only two co-assignee pairs appear across the full set of 91 families, both involving the same individual inventor pairing. Most filings here come from a single assignee acting alone, which is unusual for a technique this widely applied across industry and academia.
Claims cluster around specific workloads
Every record classifies under the core AI-model subclass, but a substantial share also carries image/video recognition or general data-processing classification, showing that compression claims are usually anchored to a target application rather than filed as pure architecture patents.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to model compression and knowledge distillation, with the prior art for and against each one.
Who is filing, and where the field is open
The assignee list is a long tail: momentum data shows most tracked assignees at zero or negative year-on-year change in the latest year, with no single filer showing sustained multi-year acceleration in this dataset.
Vellore Institute of Technology
One of the few assignees with any latest-year activity in this set, holding flat year-on-year rather than growing — consistent with steady academic output rather than a commercial filing push.
Tata Consultancy Services
Shows a sharp year-on-year drop to zero latest-year filings after prior activity, a pattern shared by several other named corporate assignees in this set rather than a company-specific anomaly.
Cortica Ltd.
Holds the representative filing in this set on teacher-student object detection loss, and its related US filing on knowledge distillation between neural networks is the most-cited record in the whole corpus.
| Assignee | Recent year | YoY |
|---|---|---|
| VELLORE INSITUTE OF TECH | 1 | 0% |
| Tata Consultancy Services | 0 | -100% |
| L'Oréal | 0 | — |
| Electronics and Telecommunications Research Institute (ETRI) | 0 | — |
| Royal Bank of Canada | 0 | — |
| Cortica Ltd. | 0 | — |
| Nankai University | 0 | — |
| East China Jiaotong University | 0 | — |
Where to take this analysis
The dataset points to a field with room to move, both in claim scope and in filing geography.
Map the white space in quantization claims
Quantization-aware training appears in the search criteria but has a thinner claim footprint than pruning or distillation in this set. A focused search on that sub-branch would clarify how open it still is.
Explore in EurekaTrack the India and China filing surge
Receiving-office volume in India and China now exceeds the United States, but citation influence still sits with US and Canadian filings. Watching how that gap closes is worth a standing search alert.
Set up a watch in EurekaCommon questions about this landscape
This landscape tracks 91 patent families filed between 2015 and mid-2026 that explicitly claim model compression, knowledge distillation, or neural network pruning alongside terms like accuracy retention, structured sparsity, or edge deployment. Filing activity has grown from zero in 2017 to a peak of 19 in 2025, with the 2026 figure still partial because publication typically lags filing by around 18 months. The real current total is therefore somewhat higher than the published count once later filings publish.
No single assignee dominates this set — the momentum data shows most tracked filers at zero or negative year-on-year change in the latest year, and co-assignment is rare, with only two co-assignee pairs across all 91 families. Cortica Ltd. holds one of the most-cited records in the corpus through its knowledge distillation filings, and academic filers such as Vellore Institute of Technology continue to file steadily. The overall picture is a long tail of single-assignee filings rather than a small group of dominant players.
India leads receiving-office filings at 29, followed by China at 24 and the United States at 18, with smaller volumes through the EPO, WIPO/PCT, and Canada. That ordering does not match citation influence, however: the most-cited records in this set were filed through the US and Canadian offices. A company deciding where to file defensively should weigh both figures rather than volume alone.
In this dataset the three terms are treated as related but distinct claim strategies: pruning claims typically target removing redundant weights or structures (often framed as structured sparsity), knowledge distillation claims target training a smaller student network to mimic a larger teacher network's outputs, and model compression is used as an umbrella term covering both plus quantization. Filings often combine more than one of these techniques in a single claim set, which is why the search terms are joined rather than treated as separate categories. IPC classification alone does not distinguish them — all three sit predominantly under G06N.
Based on IPC composition and the abstracts reviewed, quantization-aware training for transformer attention layers, structured sparsity for multi-modal fusion networks, and on-device distillation without a stored teacher model all show up in claim language but carry thinner filing density than core teacher-student distillation and pruning. These are the branches most likely to still have room for a first-mover claim. Any freedom-to-operate check should still be run against the specific claim language of the most-cited records before filing.
Research Model Compression and Knowledge Distillation in depth with Eureka
Go past this page: query the whole model compression and knowledge distillation corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.