Book a demo

Low-Bit Quantization Patents: Top Companies & Filing Trends 2026

Low-Bit Quantization Patents: Top Companies & Filing Trends 2026
https://www.patsnap.com/resources/blog/rd-blog/low-bit-quantization-for-language-model-inference-patent-landscape/ · Patsnap · data cut-off 2026-07-31 · downloaded from the live page
Patent Landscape · AI & Computer Vision
Low-Bit Quantization for Language Model Inference: Patent Trends and Who Holds the Claims
  • Filings jumped 800% from 2021 to 2024, rising from 1 to 9 records a year as post-training quantization moved from research curiosity to production requirement.
  • Half of all activity sits with five assignees, 16 of 32 records, while the next five names add only 6 more — a leader group with a long single-filer tail behind it.
  • China and the US anchor filing, 10 and 9 records respectively, with WIPO PCT filings (5) signalling that some applicants are already planning multi-jurisdiction coverage.
Get a prior-art report on your approach
32
Published Records
50%
Top-5 Share of All Records
+800%
Filing Growth 2021→2024
CN
Leading Jurisdiction

Filing growth compares 2021 (1 records) with 2024 (9) — a three-year span. 2024 is the most recent year we treat as complete: publication lags filing by roughly 18 months, so 2025 onwards are still filling in and any growth rate that ends there would understate the field. Top-5 share is the combined record count of the five largest assignees divided by all 32 records in scope (CR5), not by the ranked leaders only.

Published byPatsnap Research··6 min readSourced from Patsnap Eureka
Overview

What this landscape covers

This dataset tracks 32 published patent records matching low-bit quantization techniques for language model inference — post-training quantization, four-bit inference formats, outlier channel handling, kernel support and memory bandwidth reduction claimed together with quantization. The search spans filings from 2015 through the 2026 cut-off, though publication lag means the most recent one to two years are still filling in.

The field is small enough that individual assignees move the concentration numbers meaningfully, and young enough that no single architecture has settled into dominant prior art. That combination — low volume, high recent growth — is exactly where freedom-to-operate work pays off before the claim space fills in.

Filing activity and technology composition, 2015-2026
  1. 1DEEPLITE INC5
  2. 2BOARD OF RGT THE UNIV OF TEXAS SYST4
  3. 3TEXAS INSTRUMENTS INC3
  4. 4NANJING UNIV2
  5. 5MACRONIX INTERNATIONAL CO LTD2
  6. 6ELECTRONICS & TELECOMM RES INST2
  7. 7SHANGHAI JIAOTONG UNIV1
  8. 8SRM UNIV AP1
  9. 9Hefei Jun Zheng Technology Co., Ltd.1
  10. 10KOREA ADVANCED INST OF SCI & TECH1
Source: Patsnap Eureka. Assignee ranking and totals. Derived from a Patsnap search on Low-Bit Quantization for Language Model Inference covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
The Numbers

Filing trend and technology composition

Two views of the same 32-record dataset: filings by year, and the IPC subclasses those records fall under. Because a single record can carry more than one classification, the subclass shares add up to more than the record total.

A three-year run from near-zero to a peak of 9

Annual filings held near zero through the late 2010s, then rose from 1 in 2021 to 9 in 2024 — the peak year so far, and the last year the trend can be read as complete given typical 18-month publication lag.

A three-year run from near-zero to a peak of 90358100201720182019202020212022202392024202522026Most recent year is partial — publication lag means later filings are not yet visible.

G06N dominates; supporting classes point to hardware-adjacent claims

G06N (AI model computing) covers 75.0% of the 32 records, by far the largest single class. G06F (digital data processing) at 21.9% and H04B/G06V each at 9.4% show a meaningful minority of filings anchoring quantization claims to transmission hardware or vision pipelines rather than pure model architecture.

G06N dominates; supporting classes point to hardware-adjacent claimsG06N · Computing based on AI models2475.0%G06F · Electric digital data processi…721.9%G06V · Image/video recognition39.4%H04B · Transmission (general)39.4%H04L · Digital information transmissi…26.3%G06K · Data recognition & presentation13.1%G06T · Image data processing & genera…13.1%G10L · Speech & audio analysis/synthe…13.1%

Shares are the percentage of the 32 records in scope. A patent can carry several IPC classes, so the shares add up to more than 100%.

Source: Patsnap Eureka. Filing trend and technology composition. Derived from a Patsnap search on Low-Bit Quantization for Language Model Inference covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.

Go deeper on Low-Bit Quantization for Language Model Inference with Eureka

This page is one run against one query. Ask Eureka your own question about low-bit quantization for language model inference and every answer comes back with the patent numbers behind it.

Try Eureka
Key Patents

The most-cited records in this dataset

Highest-citation records in scope
#Publication no.Patent titleCitations
1US20210224658A1Parametric Power-Of-2 Clipping Activations for Quantization for Convolutional Neural Networks37
2US20220416937A1Autoencoder-based error correction coding for low-resolution communication11
3WO2021041551A2Autoencoder-based error correction coding for low-resolution communication8
4CN116483774A一种兼容脉动阵列加速器的矢量处理器及处理方法5
5CN118350420A一种针对大语言模型的内存高效与量化感知微调方法及装置2
6US20250045572A1Quantization for neural networks2
7KR102689249B1Method and apparatus for implementing quantization technique and scaling technique for light-weighting of dif…2
8CN110378466A基于神经网络差分的量化方法及系统2
9CN120220659A一种大型语音模型超低位训练后量化方法及系统1
10US20240362470A1Panoptic perception system, method thereof and non-transitory computer-readable media1

Citation counts reflect activity inside this searched corpus and favour older filings; treat them as a signal of influence rather than current commercial weight.

Patent titles are shown in the language they were filed in, not translated, so that each record stays verifiable against the original filing — a translated title will not match in Eureka or in any national register. Each row carries its publication number; clicking a row searches Eureka by that number.

Source: Patsnap Eureka. Citation counts and representative records. Derived from a Patsnap search on Low-Bit Quantization for Language Model Inference covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Run it yourself

Put your own technology through the same analysis

 
Where to run it
Fastest

Eureka on the web

When you want the answer in the next five minutes.

The agent works the prompt against patents and technical literature, citing every source.

Run your analysis now →
For builders

MCP server & REST API

When it has to run inside your own pipeline.

Patent search, landscape analysis and assignee resolution as MCP tools. Drop them into any agent framework, or call REST directly.

Browse MCP servers →
Insights

What the data means for filing strategy

Three patterns stand out once the raw counts are read against each other: where growth is concentrated, how classification splits between software and hardware framing, and where receiving-office choice signals strategic intent.

Growth
+800%
2021 → 2024 filings

Growth is real but the base is small

Nine filings in the peak year is a meaningful jump from one, but it is still a small absolute number. A handful of new entrants in any single year can shift the ranking, so treat leadership in this field as provisional rather than settled.

2021-2024, complete years only
Concentration
50.0%
of 32 records held by top 5

Half the field, five names

The top five assignees combined account for 16 of the 32 records in scope. The next five add only 6 more, which means the practical competitive set to watch is small even though 20 companies appear somewhere in the ranking.

Top 5 vs top 10: 50.0% vs 68.8%
Geography
10 vs 9
China vs US receiving offices

Filing is split, not one-sided

China leads receiving offices with 10 records and the United States follows closely with 9. Five WIPO PCT filings suggest a subset of applicants are already building multi-jurisdiction coverage rather than filing domestically only.

Receiving offices: CN 10, US 9, WIPO 5, KR 3, TW 2, CA 1
Eureka AI Agent
Looking for what nobody has claimed yet?

Eureka can read the same corpus for gaps instead of for coverage: under-claimed branches adjacent to low-bit quantization for language model inference, with the prior art for and against each one.

Find the white space →
Source: Patsnap Eureka. Co-assignee relationships and derived observations. Derived from a Patsnap search on Low-Bit Quantization for Language Model Inference covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
Who's Filing

Leaders and the long tail

The ranking covers 20 companies total, counted by patent family. Activity concentrates early and thins quickly — a pattern typical of a technology still moving from lab demonstration to shipped hardware.

Leader
5 records
top-ranked assignee

A single leader, not yet a dominant one

The top-ranked assignee holds 5 of the 32 records in scope — a meaningful lead in a field this size, but not enough to foreclose the space. Fifth place already drops to 2 records, showing how quickly the ranking flattens.

Leader: 5 records; 5th place: 2 records
Mid-field
2 records
fifth-ranked assignee

A thin middle tier

Between the leader and the long tail sits a mid-tier of assignees filing two records apiece. This is the group most likely to move up or drop out as the field matures over the next few publication cycles.

Ranked from 20 companies total
Tail
1 record
tenth-ranked assignee

A long single-filing tail

By tenth place, filing volume is down to a single record. That long tail includes university research groups and regional players testing the space rather than building a defensive portfolio around it.

Top 10 combined: 68.8% of 32 records
🔍
Under-claimed branches worth watching
Sub-areas that show up in the search terms but carry little dedicated claim density in this dataset — early filing here faces thinner prior art.
Outlier channel handling for activation quantizationKernel-level support for four-bit inference formatsAccuracy recovery after aggressive weight quantizationMemory bandwidth reduction techniques tied to quantized inferenceSpeech/audio quantization pipelines (G10L overlap)
Rank all filers by momentum →
Recent-year filing momentum
AssigneeRecent yearYoY
Shenbi Maliang Artificial Intelligence (Hangzhou) Co., Ltd.1
DEEPLITE INC0
Board of Regents, The University of Texas System0
Texas Instruments Inc.0
Electronics and Telecommunications Research Institute (ETRI)0
Macronix International Co., Ltd.0
Nanjing University0
Peking University0-100%
Source: Patsnap Eureka. Assignee-level momentum. Derived from a Patsnap search on Low-Bit Quantization for Language Model Inference covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
What's Next

Where to take this analysis

The numbers on this page describe what has been filed; deciding what to do next means running the same lens against a specific product roadmap or claim draft.

Check freedom-to-operate against the leader group

With half the field held by five assignees, a targeted claim-chart review of their published families is the fastest way to see whether a planned quantization pipeline sits close to existing coverage.

Explore assignee portfolios in Eureka →

Watch the under-claimed branches before they fill in

Outlier handling and four-bit kernel support show up in the search terms but carry thin dedicated claim density today. That gap narrows fast once filing volume compounds at the current growth rate.

Run a white space search in Eureka →
Source: Patsnap Eureka. Forward-looking reading of the same dataset. Derived from a Patsnap search on Low-Bit Quantization for Language Model Inference covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP
FAQ

Common questions on this landscape

Answers are grounded in the same dataset. Derived from a Patsnap search on Low-Bit Quantization for Language Model Inference covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish.Run this in Eureka MCP

Research Low-Bit Quantization for Language Model Inference in depth with Eureka

Go past this page: query the whole low-bit quantization for language model inference corpus yourself, in your own scope.
Every answer comes back with patent numbers you can open.

Try Eureka

Disclaimer. This page is generated from Patsnap Eureka data drawn from a limited snapshot of global patent and scientific-literature records, and is provided for general information and reference only.

Patent data carries inherent limitations: recent filings (typically the most recent 18–24 months) are under-counted due to standard publication lag; counts may be reported at either a patent-family or a patent-record basis and are not always directly comparable; classification, applicant-name, and citation data may contain errors, duplicates, or omissions; and the underlying search query defines and constrains the scope shown. As a result, the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.

Nothing on this page constitutes an exhaustive prior-art, novelty, freedom-to-operate, or validity search, nor does it constitute legal, financial, investment, or professional advice, and it should not be relied upon as such. Any patent, commercial, or strategic decision should be verified independently and reviewed with qualified patent, legal, and domain professionals. Patsnap makes no warranties, express or implied, as to the accuracy, completeness, or fitness for any particular purpose of the information presented.

Machine translation. Assignee and organisation names originally recorded in Chinese, Japanese or Korean have been rendered into English by an AI translation step so that the tables stay readable. These renderings are best-effort and may not match a company’s registered English name; the original name is what the underlying patent record carries, and it is what any Eureka query launched from this page uses.

Help us improve this page

Found incorrect or outdated information? Let us know and we'll get it fixed.