Non-Patent Literature Search: The Overlooked Critical Step
Introduction
The vast majority of patentability searchers habitually search only patent databases, but non-patent literature search is often just as important. Yet the reality is: a significant amount of destructive prior art exists precisely outside of patent databases. Academic papers, technical standards, conference proceedings, open-source code, product manuals — these are all “Non-Patent Literature” (NPL), frequently underestimated or overlooked in patentability searches.
According to WIPO statistics, approximately 30%–40% of prior art exists solely in non-patent literature. In certain technology sectors (such as artificial intelligence, biotechnology, and software), this proportion is even higher. This article methodically presents strategies and tools for non-patent literature search, helping you address the weakest link in patentability searching.
Why Non-Patent Literature Search Is Critical in Patentability Searching
Data That Will Change How You View NPL
- Over 3 million scholarly research papers are published globally each year
- Over 200 million public repositories exist on GitHub
- Over 200,000 preprints in AI/ML alone are available on arXiv
- Technical standard documents are updated far more frequently than the patent prosecution cycle
Non-patent literature is typically published far faster than patents — a paper may go from submission to online publication in just a few weeks, whereas a patent takes 18 months from filing to publication. This means: by the time you can search the latest patent literature, the truly cutting-edge technical information may already have been publicly disclosed through NPL 6–12 months earlier.
The Special Value of NPL
- Bridging the time gap: during the 18-month “blackout window” between patent filing and publication, NPL is the only channel for discovering the latest technical developments of competitors
- Cross-domain discovery: patent classification tools have inherent limitations — a cross-disciplinary paper may not appear in your pre-defined classification search results
- Academic prior art: innovations from universities and research institutions are often published as papers first, with patent applications considered only later
- The open-source community: in AI/software domains, a great deal of innovation is released in open-source form, which may not necessarily be converted into patent applications
- “Hidden” prior art from industry: while product white papers and technical blogs do not qualify as strict “academic publications,” they nonetheless constitute publication disclosure in the legal sense
Major Types of Non-Patent Literature and Search Strategies
Type 1: Academic Papers and Journals
This is the largest and most structured category of NPL.
Major NPL source categories Platforms:
| Platform | Coverage | Features | Free/Paid |
|---|---|---|---|
| CNKI (中国知网) | Most comprehensive Chinese-language academic literature | Covers journals, theses, conference papers, standards | Paid (institutional subscription) |
| 万方 / 维普 | Chinese-language academic literature | Supplementary to CNKI; cross-validation | Paid |
| Google Scholar | Global academic literature | Broad coverage, free, citation relationship tracking | Free |
| IEEE Xplore | Electrical/electronic, computer, and communications engineering | Authoritative source for electrical and electronics fields | Paid |
| PubMed | Biomedicine / life sciences | Maintained by NIH; core medical and biological literature | Free |
| Scopus / Web of Science | Global academic literature | High-quality indexing, professional analytical tools | Paid |
| arXiv | Physics, mathematics, CS, AI | Preprints, extremely fast updates | Free |
| Semantic Scholar | AI-powered academic search | Free, semantic search + citation analysis | Free |
Recommended Search Strategy:
- Begin with a broad initial search on Google Scholar (broadest coverage, no paywall barrier)
- Once high-quality relevant papers are identified, leverage their citation relationships and references to “snowball”
- If the target technical field is clearly defined, supplement with targeted searches on specialized platforms (e.g., biology → PubMed, electronics → IEEE)
Type 2: Technical Standards
Technical standards often contain extensive descriptions of technical solutions, yet they are not found in patent databases.
Major Sources:
- National standards: China National Standards Full-Text Disclosure portal at openstd.samr.gov.cn, free
- International standards: official websites of ISO, IEC, and IEEE standards organizations
- Industry standards: telecommunications (3GPP, ETSI), automotive (SAE), medical (ISO 13485)
- Industry consortium standards: e.g., Wi-Fi Alliance, Bluetooth SIG
Search tip: standards typically follow numbering schemes. Begin with keyword searches on the standards organization’s official website to identify relevant standard numbers, then review the full text.
Type 3: Theses and Dissertations
Master’s and doctoral theses are important sources of prior art. In China, in particular, theses publicly disclosed through CNKI legally constitute publication disclosure.
- Domestic (China): CNKI doctoral/master’s thesis database, 万方 theses
- International: ProQuest Dissertations & Theses
Type 4: Conference Proceedings and Presentation Materials
Presentation slides, posters, and proceedings from academic conferences and technical forums.
- IEEE conference papers (via IEEE Xplore)
- ACM conference papers (via ACM Digital Library)
- Public presentation platforms such as Slideshare, Speaker Deck
- Publicly available technical materials from industry summits (e.g., CES, MWC, NeurIPS)
Type 5: Open-Source Code and Software Repositories
In the software and AI domains, open-source code repositories are an extremely important source of prior art.
- GitHub: the world’s largest code hosting platform. Code and documentation in public repositories constitute publication disclosure
- GitLab / Bitbucket: other mainstream code hosting platforms
- PyPI / npm / Maven: publicly released libraries in programming-language package managers
- Papers with Code: a platform linking papers with code, particularly useful in AI patentability searches
Key judgment: the contents of public code repositories (including code comments, README documentation, issue discussions, etc.) can all constitute prior art. In software/AI patentability searches, GitHub searching should be given due attention.
Type 6: Corporate Technical Materials and Product Documentation
- Technical white papers and technical blogs on corporate websites
- Product datasheets, user manuals
- Product crowdfunding pages predating patent applications (Kickstarter, Indiegogo, etc.)
- Public demonstrations on technical video platforms such as YouTube
Practical Workflow for NPL Searching
Step 1: Determine the Scope of NPL Search
Not every invention requires a full-scale NPL search. Determine based on the invention’s characteristics:
- Software/AI-related inventions → high NPL search priority (open-source code + arXiv/conference papers are especially important)
- Chemistry/pharmaceuticals → CAS/SciFinder specialized chemistry databases + PubMed
- Mechanical/structural → relatively lower NPL priority, but supplementary standards searching is still recommended
- Biotechnology → PubMed + BLAST sequence searching
Step 2: Execute the NPL Search
It is recommended to proceed in the following order, from highest to lowest efficiency:
- Google Scholar (broadest coverage, free, first-choice starting point)
- Specialized platform search (select CNKI, PubMed, IEEE, arXiv, etc. based on the field)
- GitHub/code repositories (mandatory for software-related inventions)
- Standards organization websites (when the technology involves standards)
- General search engine supplementation (to discover overlooked unstructured public information)
Step 3: Screening and Assessment of NPL Documents
The assessment criteria for NPL documents are consistent with those for patent documents, focusing on:
- Was the publication date earlier than the filing date / priority date? (exercise careful judgment where uncertain)
- Relevance of technical content: does it address the same or similar technical problem and technical solution?
- Reliability of the source: formal publications > personal blogs; timestamped materials > materials without date information
Common Challenges in NPL Searching and How to Address Them
Challenge 1: Paywalls
Many high-quality academic papers require payment or institutional subscriptions to access the full text.
Mitigation strategies:
- Google Scholar can often locate preprint versions or author self-archived free copies
- Free full texts may be available on arXiv, ResearchGate, or the author’s personal homepage
- If the full text cannot be obtained, make a preliminary assessment based on at least the abstract and keywords, and annotate: “Full-text acquisition required for confirmation”
Challenge 2: Language Barriers
Non-English NPL (Chinese CNKI, Japanese J-STAGE, etc.) may contain unique technical information.
Mitigation strategies:
- If the target market includes the relevant country, NPL in that country’s language must be searched
- AI translation tools (such as DeepL) can assist in reading non-native-language abstracts
- When uncertain, engage a domain expert fluent in that language to assist in assessment
Challenge 3: Uncertain Publication Dates
Blog posts, forum threads, and GitHub repositories on the internet may lack explicit creation dates.
Mitigation strategies:
- Use the Wayback Machine (archive.org) to verify the archival date of web pages
- GitHub’s commit history can prove the public disclosure timeline of code
- Where the publication date cannot be verified, note “date uncertain” in the patentability search report
- As a matter of principle, information whose publication cannot be confirmed as predating the filing date cannot serve as reliable prior art basis
Challenge 4: Overwhelming Volume of NPL
Sometimes the volume of NPL search hits far exceeds that of patent searches, making screening difficult.
Mitigation strategies:
- Sort by “cited by” count in Google Scholar, prioritizing highly-cited classic literature
- Filter by “time range,” prioritizing papers that first proposed similar solutions
- Apply a classification mindset: identify which journals/conferences the technology is primarily published in, then narrow the scope
Real-World Case Study: How NPL Reversed a Patentability Assessment
Scenario: An AI startup developed a “rapid method for predicting three-dimensional protein structures” based on the Transformer architecture. The patent search only uncovered some traditional physics-simulation-based structure prediction methods, and the initial assessment suggested good novelty and inventive step.
NPL supplementary search: The patentability searcher discovered a preprint paper published 6 months earlier on arXiv, titled “Large-scale Protein Structure Prediction via Enhanced Attention Mechanisms,” which proposed a highly similar approach. Further investigation revealed that the author team had already open-sourced the complete code on GitHub.
Outcome: This combination of the arXiv preprint + GitHub open-source code posed a serious threat to the novelty of the invention. The company decided to abandon the original approach and re-focus its innovation on a specific post-processing optimization module not covered in the paper.
Lesson: In the AI domain, NPL (arXiv, GitHub) often publicly discloses innovations earlier than patent databases. If you search only patent databases, you may believe you are “first,” when in reality a publicly available solution already exists.
NPL Search Checklist
- [ ] Has Google Scholar been searched for core keywords? (in both Chinese and English)
- [ ] Have supplementary searches been conducted on field-specific platforms? (CNKI / PubMed / IEEE / arXiv as appropriate)
- [ ] For software/AI inventions: has GitHub and relevant open-source platforms been searched?
- [ ] For inventions involving industry standards: have the public documents of relevant standards organizations been searched?
- [ ] Have general search engines been used to supplementarily discover corporate white papers, technical blogs, etc.?
- [ ] Do all cited NPL references have confirmable publication dates?
- [ ] Have NPL references whose full text could not be obtained been annotated with an explanation of the uncertainty?
Key Takeaway: Non-Patent Literature (NPL) is an easily overlooked yet critically important source of prior art in patentability searching. In fields such as AI, biotechnology, and software, the prior-art value of NPL may even exceed that of patent literature. Non-patent literature search must cover academic papers (Google Scholar + specialized platforms), technical standards, open-source code (GitHub), corporate technical materials, and more. Fully documenting the non-patent literature search process and any uncertainties is a fundamental requirement of professional patentability searching. PatSnap Analytics can also support patent and non-patent evidence review when teams need a repeatable workflow.