Protein Language Model Patents: Mapping Who Files, What They Claim, and What's Still Open
43.2% of records sit with the top five filers, yet 71 companies appear in the ranking — concentration at the top with a long tail below it. Filings rose from zero in 2017 to a peak of 54 in 2025, with 7 recorded so far in 2026. Because publication typically lags filing by roughly 18 months, the 2026 figure understates actual filing activity for that year and should not be read as a slowdown.
- 1X DEVELOPMENT LLC26
- 2NE47 BIO INC5
- 3GLAXOSMITHKLINE BIOLOGICALS SA4
- 4THE GOVERNING COUNCIL OF THE UNIV OF TORONTO3
- 5MICROSOFT TECHNOLOGY LICENSING LLC3
See the full protein language models analysis in Eureka
- The complete ranking, not just the top five
- Every IPC branch with its share of the corpus
- The most-cited records, and where claim space is still thin
Common questions about protein language model patents
How many patent families exist for protein language models?
The dataset in scope contains 95 published patent records, treated as 95 families for ranking purposes, filed between 2015 and the 2026 data cut-off. Filing was negligible before 2017 and rose to a peak of 54 records in 2025. Because publication typically lags filing by around 18 months, the true 2025 and 2026 totals are likely higher than currently visible.
Who are the leading patent filers in protein language models?
The ranking covers 71 companies, with the leading assignee holding 26 of the 95 records in scope — well ahead of the fifth-ranked filer at 3 and the tenth-ranked filer at 2. The top five assignees combined account for 43.2% of all records, and the top ten reach 54.7%, meaning the field has a clear leader but roughly half the filing activity remains spread across a long tail of smaller filers. Several early leaders show no filings in the most recent year, which is worth checking against pipeline lag before assuming exit.
What patent classes cover protein language model technology?
G16B (bioinformatics) is the dominant class, appearing on 71.6% of the 95 records, followed by G06N (AI-based computing) at 42.1% and G06F (digital data processing) at 24.2%. Smaller classes including G01N, G16H, C12M, C12N and G16C each cover under 10% of records, indicating narrower and potentially less contested claim territory. Because records can carry multiple IPC codes, these shares add up to more than 100% and are not directly comparable to a single-class total.
Disclaimer. This analysis is based on Patsnap Eureka data drawn from a limited snapshot of global patent records and is provided for general information and reference only. Patent data carries inherent limitations — recent filings are under-counted because of publication lag, counts may be on a record or family basis, classification and applicant-name data may contain errors or duplicates, and the underlying search query defines the scope shown — so the analysis may be incomplete or inaccurate and may not reflect the full technology landscape.
Nothing here is an exhaustive prior-art, novelty, freedom-to-operate or validity search, nor does it constitute legal, financial or professional advice, and it should not be relied upon as such. Verify independently and review with qualified patent and legal professionals before acting on it.
Method: Filing trend and technology composition. Derived from a Patsnap search on Protein Language Models covering 2015–2026, data cut-off 2026-07-31. Counts reflect published records only and shift as new filings publish. Every share divides by all records in scope. Data: Patsnap Eureka. See the full landscape report.