Low-Bit Quantization Patents: Top Companies & Filing Trends 2026
- Filings jumped 800% from 2021 to 2024, rising from 1 to 9 records a year as post-training quantization moved from research curiosity to production requirement.
- Half of all activity sits with five assignees, 16 of 32 records, while the next five names add only 6 more — a leader group with a long single-filer tail behind it.
- China and the US anchor filing, 10 and 9 records respectively, with WIPO PCT filings (5) signalling that some applicants are already planning multi-jurisdiction coverage.
Filing growth compares 2021 (1 records) with 2024 (9) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 32 records in scope (CR5), not by the ranked leaders only.
What this landscape covers
This dataset tracks 32 published patent records matching low-bit quantization techniques for language model inference — post-training quantization, four-bit inference formats, outlier channel handling, kernel support and memory bandwidth reduction claimed together with quantization. The search spans filings from 2015 through the 2026 cut-off, though publication lag means the most recent one to two years are still filling in.
The field is small enough that individual assignees move the concentration numbers meaningfully, and young enough that no single architecture has settled into dominant prior art. That combination — low volume, high recent growth — is exactly where freedom-to-operate work pays off before the claim space fills in.
Filing trend and technology composition
Two views of the same 32-record dataset: filings by year, and the IPC subclasses those records fall under. Because a single record can carry more than one classification, the subclass shares add up to more than the record total.
A three-year run from near-zero to a peak of 9
Annual filings held near zero through the late 2010s, then rose from 1 in 2021 to 9 in 2024 — the peak year so far, and the last year the trend can be read as complete given typical 18-month publication lag.
G06N dominates; supporting classes point to hardware-adjacent claims
G06N (AI model computing) covers 75.0% of the 32 records, by far the largest single class. G06F (digital data processing) at 21.9% and H04B/G06V each at 9.4% show a meaningful minority of filings anchoring quantization claims to transmission hardware or vision pipelines rather than pure model architecture.
Shares are the percentage of the 32 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.
Go deeper on Low-Bit Quantization for Language Model Inference with Eureka
This page is one run against one query. Ask Eureka your own question about low-bit quantization for language model inference and every answer comes back with the patent numbers behind it.
Try EurekaThe most-cited records in this dataset
| # | Publication no. | Patent title | Citations |
|---|---|---|---|
| 1 | US20210224658A1 | Parametric Power-Of-2 Clipping Activations for Quantization for Convolutional Neural Networks | 37 |
| 2 | US20220416937A1 | Autoencoder-based error correction coding for low-resolution communication | 11 |
| 3 | WO2021041551A2 | Autoencoder-based error correction coding for low-resolution communication | 8 |
| 4 | CN116483774A | 一种兼容脉动阵列加速器的矢量处理器及处理方法 | 5 |
| 5 | CN118350420A | 一种针对大语言模型的内存高效与量化感知微调方法及装置 | 2 |
| 6 | US20250045572A1 | Quantization for neural networks | 2 |
| 7 | KR102689249B1 | Method and apparatus for implementing quantization technique and scaling technique for light-weighting of dif… | 2 |
| 8 | CN110378466A | 基于神经网络差分的量化方法及系统 | 2 |
| 9 | CN120220659A | 一种大型语音模型超低位训练后量化方法及系统 | 1 |
| 10 | US20240362470A1 | Panoptic perception system, method thereof and non-transitory computer-readable media | 1 |
Citation counts reflect activity inside this searched corpus and favour older filings; treat them as a signal of influence rather than current commercial weight.
Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.
Put your own technology through the same analysis
Eureka on the web
When you want the answer in the next five minutes.
The agent works the prompt against patents and technical literature, citing every source.
Run your analysis now →MCP server & REST API
When it has to run inside your own pipeline.
Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.
Browse MCP servers →What the data means for filing strategy
Three patterns stand out once the raw counts are read against each other: where growth is concentrated, how classification splits between software and hardware framing, and where receiving-office choice signals strategic intent.
Growth is real but the base is small
Nine filings in the peak year is a meaningful jump from one, but it is still a small absolute number. A handful of new entrants in any single year can shift the ranking, so treat leadership in this field as provisional rather than settled.
Half the field, five names
The top five assignees combined account for 16 of the 32 records in scope. The next five add only 6 more, which means the practical competitive set to watch is small even though 20 companies appear somewhere in the ranking.
Filing is split, not one-sided
China leads receiving offices with 10 records and the United States follows closely with 9. Five WIPO PCT filings suggest a subset of applicants are already building multi-jurisdiction coverage rather than filing domestically only.
Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to low-bit quantization for language model inference, with the prior art for and against each one.
Leaders and the long tail
The ranking covers 20 companies total, counted by patent family. Activity concentrates early and thins quickly — a pattern typical of a technology still moving from lab demonstration to shipped hardware.
A single leader, not yet a dominant one
The top-ranked assignee holds 5 of the 32 records in scope — a meaningful lead in a field this size, but not enough to foreclose the space. Fifth place already drops to 2 records, showing how quickly the ranking flattens.
A thin middle tier
Between the leader and the long tail sits a mid-tier of assignees filing two records apiece. This is the group most likely to move up or drop out as the field matures over the next few publication cycles.
A long single-filing tail
By tenth place, filing volume is down to a single record. That long tail includes university research groups and regional players testing the space rather than building a defensive portfolio around it.
| Assignee | Recent year | YoY |
|---|---|---|
| Shenbi Maliang Artificial Intelligence (Hangzhou) Co., Ltd. | 1 | — |
| DEEPLITE INC | 0 | — |
| Board of Regents, The University of Texas System | 0 | — |
| Texas Instruments Inc. | 0 | — |
| Electronics and Telecommunications Research Institute (ETRI) | 0 | — |
| Macronix International Co., Ltd. | 0 | — |
| Nanjing University | 0 | — |
| Peking University | 0 | -100% |
Where to take this analysis
The numbers on this page describe what has been filed; deciding what to do next means running the same lens against a specific product roadmap or claim draft.
Check freedom-to-operate against the leader group
With half the field held by five assignees, a targeted claim-chart review of their published families is the fastest way to see whether a planned quantization pipeline sits close to existing coverage.
Explore assignee portfolios in Eureka →Watch the under-claimed branches before they fill in
Outlier handling and four-bit kernel support show up in the search terms but carry thin dedicated claim density today. That gap narrows fast once filing volume compounds at the current growth rate.
Run a white space search in Eureka →Common questions on this landscape
Within this 32-record dataset, one assignee leads with 5 records, followed by a mid-tier of companies filing 2 apiece by fifth place. The top five assignees combined account for 50.0% of all 32 records, and the top ten reach 68.8%. That is meaningful concentration for a field this size, but the leader's margin is thin enough that new entrants could shift the ranking within a year or two of publication catching up.
Filings grew sharply from 1 record in 2021 to a peak of 9 in 2024, an increase of 800% over that span. 2025 and 2026 show fewer published records, but that reflects publication lag of roughly 18 months rather than a real slowdown — filings made in those years are still working through the pipeline. 2024 is the most recent year that can be read as a complete picture.
China leads receiving offices with 10 records in this dataset, followed closely by the United States with 9. South Korea, Taiwan and Canada each show smaller counts, and 5 filings went through the WIPO PCT route, indicating some applicants are pursuing multi-jurisdiction protection rather than filing in a single home market.
Three-quarters of the 32 records (75.0%) carry a G06N classification, covering AI-model-specific computing methods. G06F, general digital data processing, appears in 21.9% of records, and smaller shares fall under image recognition (G06V), transmission (H04B) and digital information transmission (H04L). Because records can carry multiple classes, these shares add up to more than 100%, reflecting how quantization claims often bridge model architecture and supporting hardware.
The search terms surface several sub-areas with limited dedicated claim density in this dataset: outlier channel handling for activation quantization, kernel-level support specifically for four-bit inference formats, and accuracy-recovery methods after aggressive weight quantization. Speech and audio quantization pipelines, overlapping with G10L, also show only a single record. These branches are worth checking early, since filing density in the core G06N space is already substantial relative to the field's total size.
Research Low-Bit Quantization for Language Model Inference in depth with Eureka
Go past this page: query the whole low-bit quantization for language model inference corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.
Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.
Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.
Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.