Book a demo

Non-Patent Literature Search: The Overlooked Critical Step

Introduction

The vast majority of patentability searchers habitually search only patent databases, but non-patent literature search is often just as important. Yet the reality is: a significant amount of destructive prior art exists precisely outside of patent databases. Academic papers, technical standards, conference proceedings, open-source code, product manuals — these are all “Non-Patent Literature” (NPL), frequently underestimated or overlooked in patentability searches.

According to WIPO statistics, approximately 30%–40% of prior art exists solely in non-patent literature. In certain technology sectors (such as artificial intelligence, biotechnology, and software), this proportion is even higher. This article methodically presents strategies and tools for non-patent literature search, helping you address the weakest link in patentability searching.

Why Non-Patent Literature Search Is Critical in Patentability Searching

Data That Will Change How You View NPL

  • Over 3 million scholarly research papers are published globally each year
  • Over 200 million public repositories exist on GitHub
  • Over 200,000 preprints in AI/ML alone are available on arXiv
  • Technical standard documents are updated far more frequently than the patent prosecution cycle

Non-patent literature is typically published far faster than patents — a paper may go from submission to online publication in just a few weeks, whereas a patent takes 18 months from filing to publication. This means: by the time you can search the latest patent literature, the truly cutting-edge technical information may already have been publicly disclosed through NPL 6–12 months earlier.

The Special Value of NPL

  1. Bridging the time gap: during the 18-month “blackout window” between patent filing and publication, NPL is the only channel for discovering the latest technical developments of competitors
  2. Cross-domain discovery: patent classification tools have inherent limitations — a cross-disciplinary paper may not appear in your pre-defined classification search results
  3. Academic prior art: innovations from universities and research institutions are often published as papers first, with patent applications considered only later
  4. The open-source community: in AI/software domains, a great deal of innovation is released in open-source form, which may not necessarily be converted into patent applications
  5. “Hidden” prior art from industry: while product white papers and technical blogs do not qualify as strict “academic publications,” they nonetheless constitute publication disclosure in the legal sense

Major Types of Non-Patent Literature and Search Strategies

Type 1: Academic Papers and Journals

This is the largest and most structured category of NPL.

Major NPL source categories Platforms:

PlatformCoverageFeaturesFree/Paid
CNKI (中国知网)Most comprehensive Chinese-language academic literatureCovers journals, theses, conference papers, standardsPaid (institutional subscription)
万方 / 维普Chinese-language academic literatureSupplementary to CNKI; cross-validationPaid
Google ScholarGlobal academic literatureBroad coverage, free, citation relationship trackingFree
IEEE XploreElectrical/electronic, computer, and communications engineeringAuthoritative source for electrical and electronics fieldsPaid
PubMedBiomedicine / life sciencesMaintained by NIH; core medical and biological literatureFree
Scopus / Web of ScienceGlobal academic literatureHigh-quality indexing, professional analytical toolsPaid
arXivPhysics, mathematics, CS, AIPreprints, extremely fast updatesFree
Semantic ScholarAI-powered academic searchFree, semantic search + citation analysisFree

Recommended Search Strategy:

  • Begin with a broad initial search on Google Scholar (broadest coverage, no paywall barrier)
  • Once high-quality relevant papers are identified, leverage their citation relationships and references to “snowball”
  • If the target technical field is clearly defined, supplement with targeted searches on specialized platforms (e.g., biology → PubMed, electronics → IEEE)

Type 2: Technical Standards

Technical standards often contain extensive descriptions of technical solutions, yet they are not found in patent databases.

Major Sources:

  • National standards: China National Standards Full-Text Disclosure portal at openstd.samr.gov.cn, free
  • International standards: official websites of ISO, IEC, and IEEE standards organizations
  • Industry standards: telecommunications (3GPP, ETSI), automotive (SAE), medical (ISO 13485)
  • Industry consortium standards: e.g., Wi-Fi Alliance, Bluetooth SIG

Search tip: standards typically follow numbering schemes. Begin with keyword searches on the standards organization’s official website to identify relevant standard numbers, then review the full text.

Type 3: Theses and Dissertations

Master’s and doctoral theses are important sources of prior art. In China, in particular, theses publicly disclosed through CNKI legally constitute publication disclosure.

  • Domestic (China): CNKI doctoral/master’s thesis database, 万方 theses
  • International: ProQuest Dissertations & Theses

Type 4: Conference Proceedings and Presentation Materials

Presentation slides, posters, and proceedings from academic conferences and technical forums.

  • IEEE conference papers (via IEEE Xplore)
  • ACM conference papers (via ACM Digital Library)
  • Public presentation platforms such as Slideshare, Speaker Deck
  • Publicly available technical materials from industry summits (e.g., CES, MWC, NeurIPS)

Type 5: Open-Source Code and Software Repositories

In the software and AI domains, open-source code repositories are an extremely important source of prior art.

  • GitHub: the world’s largest code hosting platform. Code and documentation in public repositories constitute publication disclosure
  • GitLab / Bitbucket: other mainstream code hosting platforms
  • PyPI / npm / Maven: publicly released libraries in programming-language package managers
  • Papers with Code: a platform linking papers with code, particularly useful in AI patentability searches

Key judgment: the contents of public code repositories (including code comments, README documentation, issue discussions, etc.) can all constitute prior art. In software/AI patentability searches, GitHub searching should be given due attention.

Type 6: Corporate Technical Materials and Product Documentation

  • Technical white papers and technical blogs on corporate websites
  • Product datasheets, user manuals
  • Product crowdfunding pages predating patent applications (Kickstarter, Indiegogo, etc.)
  • Public demonstrations on technical video platforms such as YouTube

Practical Workflow for NPL Searching

Not every invention requires a full-scale NPL search. Determine based on the invention’s characteristics:

  • Software/AI-related inventions → high NPL search priority (open-source code + arXiv/conference papers are especially important)
  • Chemistry/pharmaceuticals → CAS/SciFinder specialized chemistry databases + PubMed
  • Mechanical/structural → relatively lower NPL priority, but supplementary standards searching is still recommended
  • Biotechnology → PubMed + BLAST sequence searching

It is recommended to proceed in the following order, from highest to lowest efficiency:

  1. Google Scholar (broadest coverage, free, first-choice starting point)
  2. Specialized platform search (select CNKI, PubMed, IEEE, arXiv, etc. based on the field)
  3. GitHub/code repositories (mandatory for software-related inventions)
  4. Standards organization websites (when the technology involves standards)
  5. General search engine supplementation (to discover overlooked unstructured public information)

Step 3: Screening and Assessment of NPL Documents

The assessment criteria for NPL documents are consistent with those for patent documents, focusing on:

  • Was the publication date earlier than the filing date / priority date? (exercise careful judgment where uncertain)
  • Relevance of technical content: does it address the same or similar technical problem and technical solution?
  • Reliability of the source: formal publications > personal blogs; timestamped materials > materials without date information

Common Challenges in NPL Searching and How to Address Them

Challenge 1: Paywalls

Many high-quality academic papers require payment or institutional subscriptions to access the full text.

Mitigation strategies:

  • Google Scholar can often locate preprint versions or author self-archived free copies
  • Free full texts may be available on arXiv, ResearchGate, or the author’s personal homepage
  • If the full text cannot be obtained, make a preliminary assessment based on at least the abstract and keywords, and annotate: “Full-text acquisition required for confirmation”

Challenge 2: Language Barriers

Non-English NPL (Chinese CNKI, Japanese J-STAGE, etc.) may contain unique technical information.

Mitigation strategies:

  • If the target market includes the relevant country, NPL in that country’s language must be searched
  • AI translation tools (such as DeepL) can assist in reading non-native-language abstracts
  • When uncertain, engage a domain expert fluent in that language to assist in assessment

Challenge 3: Uncertain Publication Dates

Blog posts, forum threads, and GitHub repositories on the internet may lack explicit creation dates.

Mitigation strategies:

  • Use the Wayback Machine (archive.org) to verify the archival date of web pages
  • GitHub’s commit history can prove the public disclosure timeline of code
  • Where the publication date cannot be verified, note “date uncertain” in the patentability search report
  • As a matter of principle, information whose publication cannot be confirmed as predating the filing date cannot serve as reliable prior art basis

Challenge 4: Overwhelming Volume of NPL

Sometimes the volume of NPL search hits far exceeds that of patent searches, making screening difficult.

Mitigation strategies:

  • Sort by “cited by” count in Google Scholar, prioritizing highly-cited classic literature
  • Filter by “time range,” prioritizing papers that first proposed similar solutions
  • Apply a classification mindset: identify which journals/conferences the technology is primarily published in, then narrow the scope

Real-World Case Study: How NPL Reversed a Patentability Assessment

Scenario: An AI startup developed a “rapid method for predicting three-dimensional protein structures” based on the Transformer architecture. The patent search only uncovered some traditional physics-simulation-based structure prediction methods, and the initial assessment suggested good novelty and inventive step.

NPL supplementary search: The patentability searcher discovered a preprint paper published 6 months earlier on arXiv, titled “Large-scale Protein Structure Prediction via Enhanced Attention Mechanisms,” which proposed a highly similar approach. Further investigation revealed that the author team had already open-sourced the complete code on GitHub.

Outcome: This combination of the arXiv preprint + GitHub open-source code posed a serious threat to the novelty of the invention. The company decided to abandon the original approach and re-focus its innovation on a specific post-processing optimization module not covered in the paper.

Lesson: In the AI domain, NPL (arXiv, GitHub) often publicly discloses innovations earlier than patent databases. If you search only patent databases, you may believe you are “first,” when in reality a publicly available solution already exists.

NPL Search Checklist

  • [ ] Has Google Scholar been searched for core keywords? (in both Chinese and English)
  • [ ] Have supplementary searches been conducted on field-specific platforms? (CNKI / PubMed / IEEE / arXiv as appropriate)
  • [ ] For software/AI inventions: has GitHub and relevant open-source platforms been searched?
  • [ ] For inventions involving industry standards: have the public documents of relevant standards organizations been searched?
  • [ ] Have general search engines been used to supplementarily discover corporate white papers, technical blogs, etc.?
  • [ ] Do all cited NPL references have confirmable publication dates?
  • [ ] Have NPL references whose full text could not be obtained been annotated with an explanation of the uncertainty?

Key Takeaway: Non-Patent Literature (NPL) is an easily overlooked yet critically important source of prior art in patentability searching. In fields such as AI, biotechnology, and software, the prior-art value of NPL may even exceed that of patent literature. Non-patent literature search must cover academic papers (Google Scholar + specialized platforms), technical standards, open-source code (GitHub), corporate technical materials, and more. Fully documenting the non-patent literature search process and any uncertainties is a fundamental requirement of professional patentability searching. PatSnap Analytics can also support patent and non-patent evidence review when teams need a repeatable workflow.

Your Agentic AI Partner
for Smarter Innovation

Patsnap fuses the world’s largest proprietary innovation dataset with cutting-edge AI to
supercharge R&D, IP strategy, materials science, and drug discovery.

Book a demo