Book a demo

How Accurate Is AI Prior Art Search? What the Benchmarks Measure

Prior art search · Accuracy

What the published benchmarks actually measure, what the human baseline looks like, and how to test a search tool on your own art instead of taking a percentage on trust.

Most claims about AI accuracy are unfalsifiable, because there is no answer key. Prior art search is the rare exception. Every granted patent carries a list of the documents an examiner cited against it, categorised by how damaging each one is, and that list is an objective target that no vendor controls.

That answer key is what turns the question from a marketing argument into a measurement. It is also why the serious numbers in this field are built the same way: take a set of applications, hide the search report, run the tool on the disclosure, and check how much of the examiner-cited art comes back and how near the top it lands. Patsnap publishes its Novelty Search agent against exactly that target in PatentBench, with the dataset composition, the metric definitions and the comparison systems all stated on the page.1

A single percentage, though, is close to meaningless on its own. Three things have to travel with it: the sample it was measured on, the definition of a correct answer, and which metric produced the number. This article walks through all three, then sets them against what human search actually achieves, and ends with a procedure for running the test on your own portfolio.

In short
  • Patent search is a recall problem, not a precision problem. One missed document can invalidate a patent or stop a launch, while one extra irrelevant hit costs a minute of reading. A dedicated recall-oriented metric, PRES, was proposed for this exact task in 2010 because precision-weighted scores were hiding the failure that matters.
  • The answer key is examiner citations. X and Y references, deduplicated and normalised by patent family, give a target defined by patent offices rather than by the tool being tested.
  • The human baseline is not perfection. WIPO states that no single office can search the whole of the PCT minimum documentation in its original language, and published research puts a US examiner at roughly nineteen hours per application in total.
  • Cross-field art is where every system degrades. A 2025 family-level benchmark measured out-of-domain retrieval roughly five times worse than in-domain across every configuration it tested.
  • Finding art and judging it are different tasks. On a hands-on patent attorney examination the strongest model reached 0.82 accuracy, and the study reports that no model exceeded the 0.90 threshold it identifies for professional-level standards.

What accuracy means here, and why the metric decides the answer

Three different numbers get called accuracy in patent search, and they answer different questions.

Hit rate asks whether at least one correct answer appeared near the top. It is reported at a cut-off, so a hit rate at the top 5 means the proportion of test cases where a correct document was among the first five results. It is the number closest to the experience of opening a result list and finding something usable immediately.

Recall asks how much of the known-correct set was found at all, within some larger window such as the top 100. It is the number closest to the question a searcher is actually being paid to answer, which is whether anything relevant was left behind.

Precision asks what proportion of returned documents were relevant. It governs how much reading a result list costs, and it is the metric most information retrieval research has historically optimised for.

The mismatch between that last fact and the nature of patent work is not a new observation. In 2010 Magdy and Jones set out the problem in a SIGIR paper, noting that while precision and recall “are often discussed with equal importance, in practice most attention has been given to precision focused metrics,” and that “even for recall-oriented IR tasks of growing importance, such as patent retrieval, these precision based scores remain the primary evaluation measures.” They proposed PRES, a metric that accounts for “recall and the user’s search effort,” and demonstrated it on 48 runs from the CLEF-IP 2009 patent retrieval track.7

The practical consequence is simple. A tool can look excellent on a precision-weighted score and still be the wrong instrument for a novelty or freedom-to-operate search, because the document that kills the claim is the one it did not return. When a vendor quotes an accuracy figure, the first question is which of these three it is.

Prior art search · Eureka

Run a case where you already know the answer

Feed the agent a disclosure from a family you have already prosecuted, then compare what comes back against the examiner citations on the file. It is the fastest honest test of any search tool, and a single case runs in about 30 minutes.

Try Eureka free

10,000 free credits to get started

The answer key: what counts as a correct result

Search reports classify citations by how they bear on patentability, using categories defined in WIPO Standard ST.14. A category X document means “the claimed invention cannot be considered novel or cannot be considered to involve an inventive step when the document is taken alone.” A category Y document means the invention “cannot be considered to involve an inventive step when the document is combined with one or more other such documents, such combination being obvious to a person skilled in the art.” A category A document merely defines the general state of the art and is “not considered to be of particular relevance.”6

That gives a graded target rather than a binary one, and the X and Y references are the part worth measuring against. They are the documents that changed the outcome.

PatentBench builds its ground truth this way: X and Y references cited by examiners at different receiving offices are “collected, deduplicated, and normalized by patent family.”1 The family normalisation matters more than it sounds. The same disclosure is typically published several times across offices, so a system that returns three members of one family has found one piece of art, not three. Counting publications instead of families is one of the easiest ways to make a search score look better than it is.

Two properties make this target useful. It is defined by patent offices rather than by whoever is running the test, and it is independent of the retrieval method, so a keyword search, an embedding model and an agent are all scored against the same thing. Its limitation is that examiner-cited art is what one examiner found in the time available, which makes recall measured against it a floor rather than a ceiling.

What the published numbers actually say

Novelty search against examiner-cited art

PatentBench, dated July 2026, runs on 340 cross-jurisdiction patent family samples, 68% English and 32% Chinese, with US and Chinese applications at roughly 32% each and EP and WIPO at roughly 18% each, distributed across IPC classifications. Two metrics are reported, and the page glosses both in plain terms. X Hit Rate is the share of test cases in which the agent “successfully identified at least one relevant X document”; X Recall Rate is the share of “all relevant X documents” it retrieved.1

The page reports that Patsnap’s Novelty Search AI Agent “achieved an 85% X Hit Rate and a 37% X Recall Rate within the top 100 results,” and that “Claude Opus 4.8 ranked second, with a 52.37% X Hit Rate and an 11.68% X Recall Rate.” The remaining systems evaluated are also general-purpose models with web search enabled: Perplexity Pro, ChatGPT 5.4, Gemini 3.1 Pro and DeepSeek 3.2.1

The system behind those runs is an agent workflow rather than a query box, described on the product page as turning an invention disclosure into ranked prior art and feature-level evidence. Coverage is 200M+ patents across 174 jurisdictions, with SOC 2 and ISO 27001 certification, GDPR and CCPA compliance and no AI training on user data, which is the part that matters when the input is an unfiled invention.3

Read the two numbers together rather than picking the flattering one, because they count different things. A hit rate counts samples: did at least one correct answer come back for this application. A recall rate counts documents: how much of the whole known-correct set came back. A single application often carries several X and Y references, which is why the page can report finding at least one in 85% of cases while retrieving 37% of all of them.1 The distance between the two numbers is exactly that difference, and it is the honest picture of where retrieval currently sits: dependable at surfacing a relevant document, some way from exhaustively recovering every one.

Design freedom to operate, where the target is an image

A separate benchmark covers design rights, where the query is a product appearance rather than a text disclosure and the ground truth is a confirmed infringing registration. It runs on 261 samples spread roughly evenly across China, the European Union and the United States, covering 26 Locarno primary classifications, split between e-commerce infringement images at 64.8% and invalidation scenarios at 35.2%, with the infringement relationships “rigorously verified by a human annotation team.” The metric is a High-Risk Patent Hit Rate, defined as the “proportion of test samples where the confirmed infringing patent appears in the top K results,” reported within the top 200. The Patsnap Design FTO agent records 77%, against 3.71% for Gemini 3.1 Pro and 0.93% for ChatGPT 5.4. Across offices the page reports performance strongest at the EU office, where the agent reached a 92% hit rate.2

That page also reports a PRES score of 0.7 for what it describes as end-to-end infringement determination quality, evaluating not only whether the infringing registration was found but how highly it was ranked, and it credits the recall-oriented PRES metric Magdy and Jones published in 2010.2, 7

What the academic literature reports

Independent benchmarks give the wider picture, and it is a sobering one. DAPFAM, a 2025 family-level dataset built specifically to test cross-domain retrieval, ran 249 controlled experiments across lexical, dense and rank-fusion configurations over 1,247 query families and 45,336 target families, and found that “OUT-domain performance remains roughly five times lower than IN-domain across all configurations.”12 PatenTEB, a companion benchmark of 15 patent embedding tasks and 2.06 million examples from the same group, reports its strongest model reaching 0.377 nDCG@100 on DAPFAM.13 Those are the numbers for off-the-shelf retrieval, and they are a long way from solved.

The baseline nobody quotes: how accurate is human search?

Asking whether AI search is accurate enough is only answerable against something. The something is usually assumed to be an exhaustive expert search, and the published evidence does not support that assumption.

No office searches everything. WIPO’s own explanation of why supplementary international search exists is direct: “no single Office is capable of searching the whole of the PCT minimum documentation in its original language,” and a supplementary search “should further reduce the risk of new prior art being cited in the national phase.”8 The international system is designed around the expectation that one authoritative search will miss things.

The time available is finite. Frakes and Wasserman, studying USPTO examination, report that “on average, an examiner spends only nineteen hours reviewing an application, including reading the patent application, searching for prior art, comparing the prior art with the patent application, writing a rejection, responding to the patent applicant’s arguments, and often conducting an interview with the applicant’s attorney.” Searching is one line item inside that budget, not the whole of it. Their central finding is that when promotion reduces the time allotted, grant rates rise substantially and the share of citations originating from examiner search falls, with obviousness rejections, the most search-intensive kind, dropping first.9

Applicant-supplied art rarely does the work. Cotropia, Lemley and Sampat found that examiners “focus almost exclusively on art they find themselves,” with roughly 87% of references used in rejections coming from examiner searches, only about 2% of applicant-submitted references ever cited in a rejection, and over 90% of the US patents used in rejections emanating from examiner search.10 Whatever the search process is, it is one person’s search under time pressure.

The offices are automating too. The USPTO has deployed an AI-based Similarity Search inside its internal PE2E examiner system, which uses trained AI models to output a list of documents similar to the application being searched. A memorandum to the Patent Examining Corps dated 24 October 2025 states that “examiners are required to use Similarity Search during examination of a plant or utility application and record the conducted Similarity Search in the application file wrapper.”23, 24 The question of whether machine retrieval belongs in prior art search has effectively been answered by the office that sets the standard.

The misses surface later. In USPTO PTAB trials for fiscal year 2025, 22% of challenged claims and 47% of instituted claims were found unpatentable by a preponderance of the evidence.11 Those outcomes rest largely on art that a motivated challenger found and the original search did not.

None of this is an argument that automation is therefore good enough. It is an argument that the comparison should be honest. The useful question is not whether a tool is perfect, but whether it recovers more of the known-correct art per hour than the process it is replacing, and whether a professional can check its work document by document.

From a percentage to your own evidence

Source-linked results you can actually check

Every finding maps back to the passage it came from, so verification is reading a citation rather than trusting a score. Coverage is 200M+ patents across 174 jurisdictions, with SOC 2 and ISO 27001 and no AI training on your data.

Try Eureka free

SOC 2 and ISO 27001, no training on your data

Where AI prior art search still fails

Art from another field. This is the documented hard case and it is also the commercially expensive one, because the reference that invalidates a mechanical claim is often sitting in a different technology area entirely. DAPFAM measured the gap at roughly five times across every configuration tested, and concluded that neither dense retrieval nor lexical search closes it.12

Coverage nobody can search around. Machine-readable full text exists for a minority of offices. The EPO’s worldwide bibliographic dataset contains abstracts, citations and families but “no full text or images.”18 Google Patents indexes “over 120 million patent publications from 100+ patent offices” but lists full text for 22 of them.19 A search over abstracts and a search over claims are different instruments, and no model compensates for text that was never indexed. The free official collections describe their own boundaries plainly for the same reason: Espacenet publishes more than 150 million documents from more than 100 patent authorities, PATENTSCOPE publishes a per-collection coverage table listing over 128 million documents across its national and regional collections, and USPTO Patent Public Search states that it does not use semantic searching and offers no access to foreign patent databases.20, 21, 22

Finding art is not judging it. A 2025 study built the first dataset for novelty evaluation from real examination cases and found that generative models “make predictions with a reasonable level of accuracy” with explanations good enough to follow the relationship between claim and reference, while classification models “struggle to effectively assess novelty.”14 Reasonable is not the same as decisive. On a hands-on patent attorney examination, the strongest model scored 0.82 accuracy and 0.81 F1 while smaller open models landed between 0.50 and 0.55, and the study reports that accuracy “never exceeded the average threshold of 0.90 required for professional-level standards.” Its human patent experts, it notes, valued clarity and legal rationale over the raw correctness of the answers.15

Fabrication when nothing is retrieved. A model that generates a citation instead of returning one from an index will produce something plausible and wrong. A research team profiling legal hallucinations, measuring general-purpose models on verifiable questions about real court cases across more than 800,000 queries, found hallucination rates from 58% to 88% depending on the model.16 Retrieval augmentation reduces the problem without removing it: the same group found purpose-built commercial legal research tools still “hallucinate between 17% and 33% of the time.”17 For patent work the practical rule follows directly. A result that does not carry a publication number you can open in an official register is not a result.

How to measure accuracy on your own art

Every published benchmark is measured on someone else’s sample. Your technology area, your claim language and your jurisdictions are the only test that predicts anything about your work, and you already own the answer key, because your granted families come with search reports. Running the test takes about a day.

  1. Assemble 20 to 50 families you have already prosecuted. Prefer cases where the examiner cited X or Y art, and spread them across the technology areas you actually file in rather than the one you know best.
  2. Build the answer key before you run anything. Take the X and Y references from every office that examined the family, deduplicate them, and normalise to patent family so that one disclosure counts once.
  3. Use the disclosure, not the granted claims. Feed the tool the invention as it looked before drafting. Feeding it the granted claim set leaks the examiner’s hindsight into the query and inflates every score.
  4. Record two numbers, not one. Hit rate at the top 5 tells you what the first screen looks like. Recall at the top 100 tells you what a full review recovers. Reporting only the first is how a weak system looks strong.
  5. Include cases you expect it to fail. Deliberately load in families whose killer reference came from another technology field. That subset is where the differences between tools are real.
  6. Score the analysis separately from the retrieval. Whether the decisive passage was correctly mapped to the claim feature is a different question from whether the document was returned, and a tool can be strong on one and weak on the other.
  7. Re-run it when anything changes. A model update, a re-chunk or a coverage change can move results in either direction, and the only way to know is to keep the set and repeat.

If you want this inside your own harness rather than a browser, the Patsnap Open Platform exposes patent research, novelty and freedom-to-operate agents as 32 MCP servers callable from Claude, Cursor or a custom client, which makes a scripted evaluation loop over your own families straightforward to build.4, 5

What no accuracy number settles

  • A benchmark is a distribution, not a promise about your case. A hit rate measured across hundreds of samples says nothing about whether the decisive reference for the disclosure in front of you was among them.
  • The answer key is itself incomplete. Examiner citations are what one examiner found in the time available, so a score measured against them understates how much art exists and overstates how much has been ruled out.
  • Novelty is not patentability. Novelty is assessed one reference at a time. Inventive step is a separate analysis under a named jurisdictional framework, and no retrieval score speaks to it.
  • Ask what produced any figure you are quoted. Sample size, the definition of a correct answer, the cut-off and the metric. A percentage without those four is not a measurement.
  • The certification stays with the person who signs. The USPTO states that “simply relying on the accuracy of an AI tool is not a reasonable inquiry,” and under 37 CFR 11.18(b) whoever signs certifies that factual contentions have evidentiary support, to the best of their knowledge, information and belief “formed after an inquiry reasonable under the circumstances.”25, 26

Frequently asked questions

How accurate is AI prior art search?
Accurate against what, measured how, on which sample. Prior art search is one of the few AI applications with an objective answer key, because examiners publish the references they cited and grade them by category. On PatentBench, a 340-sample cross-jurisdiction dataset whose ground truth is X and Y references cited by examiners at different offices, deduplicated and normalised by patent family, the published result is that Patsnap’s Novelty Search AI Agent “achieved an 85% X Hit Rate and a 37% X Recall Rate within the top 100 results,” and that “Claude Opus 4.8 ranked second, with a 52.37% X Hit Rate and an 11.68% X Recall Rate.” The two metrics count different things: a hit rate counts samples where at least one correct answer came back, while a recall rate counts how much of the whole known-correct set came back, which is why the second number is much lower than the first. Any figure quoted without a sample size, a definition of a correct answer and a named metric is not a measurement.
What is the difference between hit rate and recall in patent search?
Hit rate counts samples: it asks whether at least one correct document came back for a given application, reported at a stated cut-off. Recall counts documents: it asks how much of the entire known-correct set came back within that cut-off. Because a single application often carries several relevant references, a system can succeed on nearly every sample and still recover only a minority of the full set, so the two numbers can differ sharply on the same run. Reporting only the flattering one is the most common way accuracy claims mislead.
Why is recall more important than precision for prior art search?
Because the costs are asymmetric. One missed document can invalidate a granted patent or stop a product launch. One extra irrelevant result costs a minute of reading. Information retrieval research has historically optimised for precision, and in 2010 Magdy and Jones published PRES, a metric accounting for recall and the user’s search effort, precisely because precision-weighted scores were the primary measures even for recall-oriented tasks such as patent retrieval.
What counts as a correct answer in a patent search benchmark?
Usually an examiner-cited X or Y reference. Under WIPO Standard ST.14, a category X document means the claimed invention cannot be considered novel or inventive when the document is taken alone, and a category Y document means it cannot be considered inventive when combined with one or more other such documents, that combination being obvious to a person skilled in the art. Category A documents merely define the general state of the art. Serious benchmarks also normalise by patent family, so that returning three publications of the same disclosure counts as finding one piece of art rather than three.
Are human patent searches more accurate than AI?
Human search is the reference standard, but published evidence shows it is not exhaustive. WIPO states that no single office is capable of searching the whole of the PCT minimum documentation in its original language, which is why supplementary international search exists at all. Research on USPTO examination reports that on average an examiner spends only nineteen hours reviewing an application, covering reading, prior art searching, comparison, writing a rejection, responding to the applicant’s arguments and often an interview, and finds that when time is compressed, examiner-sourced citations fall. In PTAB trials for fiscal year 2025, 22% of challenged claims and 47% of instituted claims were found unpatentable, largely on art a later challenger located. The realistic comparison is how much known-correct art each process recovers per hour, and whether the work can be checked.
Can ChatGPT or another general chatbot do a prior art search?
It can produce something that looks like one. Whether it retrieved the documents or generated them is the question. A research team profiling legal hallucinations measured general-purpose models on verifiable questions about real court cases across more than 800,000 queries and found hallucination rates from 58% to 88% depending on the model, and the same group found that even purpose-built legal research tools with retrieval augmentation hallucinate between 17% and 33% of the time. On PatentBench, where the comparison systems are general-purpose models with web search enabled, the page reports Claude Opus 4.8 ranking second with a 52.37% X Hit Rate and an 11.68% X Recall Rate, against 85% and 37% within the top 100 results for an agent searching a licensed patent corpus. The working rule is that a result without a publication number you can open in an official register is not a result.
Why do AI patent searches miss prior art from other technology fields?
Because both lexical and semantic retrieval are anchored to the vocabulary and classification of the field the query came from. DAPFAM, a 2025 family-level benchmark, ran 249 controlled experiments over 1,247 query families and 45,336 target families and found out-of-domain performance roughly five times lower than in-domain across every configuration, with neither dense transformer retrieval nor lexical search closing the gap. This is also the commercially expensive failure, because cross-field art is exactly what a later challenger goes looking for.
Does a good benchmark score mean my search is complete?
No. A benchmark reports a distribution across a sample, not a guarantee about the disclosure in front of you, and the answer key it is measured against is itself only what examiners found in the time they had. A search establishes what was found under a recorded method on a recorded date. The USPTO states that simply relying on the accuracy of an AI tool is not a reasonable inquiry, and under 37 CFR 11.18(b) whoever signs certifies that factual contentions have evidentiary support to the best of their knowledge, information and belief formed after an inquiry reasonable under the circumstances.
How do I test a patent search tool on my own portfolio?
Take 20 to 50 families you have already prosecuted, build the answer key from the X and Y references cited by every office that examined them, deduplicate and normalise by family, then run the tool on the original disclosure rather than the granted claims so the examiner’s hindsight does not leak into the query. Record hit rate at the top 5 and recall at the top 100 separately, include a subset of families whose decisive reference came from another technology field, and score the claim-to-passage analysis separately from the retrieval. Keep the set and re-run it whenever the tool changes.

Sources and verification

Disclosure & disclaimer

Who published this. This article is published by Patsnap, which develops and sells Patsnap Eureka, one of the systems whose published results are discussed below. It is an editorial overview written by a participant in this market and is intended as general information for professionals evaluating search tools.

How the information was gathered. Benchmark figures, research findings, statutory text and office guidance are taken from the published pages, papers and documents cited, as accessed on August 18, 2026. Performance figures attributed to a tool are that operator’s published results obtained under that operator’s own methodology. Published results, datasets and office practices change, so confirm anything material directly with the source before relying on it.

Scope and limitations. The benchmarks and studies cited measure specific tasks on specific datasets and do not generalise to every retrieval system, technology field or jurisdiction. Metric definitions differ between publications, so figures reported here should be compared only within the benchmark that produced them. Nothing here is a representation that any search is complete or that any particular result will be obtained.

Trademarks. All trademarks, service marks, product names and company names are the property of their respective owners and are used here solely for identification and descriptive purposes. Their use does not imply any affiliation with or endorsement by those owners.

Not professional advice. This article is general information about search methodology and evaluation. It is not legal advice, it does not create an attorney-client or any other professional relationship, and it should not be relied on in place of advice from a qualified patent attorney or agent admitted in the relevant jurisdiction.

  1. Patsnap, PatentBench for Novelty Search: metric definitions, the 340-sample cross-jurisdiction dataset, examiner-cited ground truth and the full result set.
  2. Patsnap, PatentBench for Design FTO Search: 261-sample dataset across CN, EU and US, 26 Locarno primary classifications, High-Risk Patent Hit Rate and PRES results.
  3. Patsnap Eureka, IP Search agents: Novelty Search, FTO Search and Design FTO Search agent workflows; 200M+ patents across 174 jurisdictions; SOC 2, ISO 27001, GDPR and CCPA; no AI training on user data.
  4. Patsnap Open Platform, MCP Servers marketplace: 32 servers and client setup instructions for Claude, Cursor and custom MCP clients.
  5. Patsnap Open Platform, pricing: Starter tier, 10,000 credits for 90 days.
  6. WIPO Standard ST.14 (PDF): definitions of citation categories X, Y and A used in search reports.
  7. Magdy and Jones, PRES: A Score Metric for Evaluating Recall-Oriented Information Retrieval Applications: SIGIR 2010, demonstrated on 48 runs from the CLEF-IP 2009 patent retrieval track.
  8. WIPO, PCT Newsletter 01/2012, Practical Advice: the benefits of requesting supplementary international search: PCT minimum documentation and language coverage.
  9. Frakes and Wasserman, Is the Time Allocated to Review Patent Applications Inducing Examiners to Grant Invalid Patents? (PDF): NBER Working Paper 20337; examiner time allocation and its effect on search and grant behaviour.
  10. Cotropia, Lemley and Sampat, Do Applicant Patent Citations Matter? (PDF): published in Research Policy 42(4), 2013; share of rejection references originating from examiner search.
  11. USPTO, PTAB Trial Statistics FY2025 end-of-year roundup (PDF): challenged and instituted claims found unpatentable.
  12. DAPFAM: A Domain-Aware Family-level Dataset to benchmark cross domain patent retrieval: Ayaou, Cavallucci and Chibane, arXiv:2506.22141, 2025; 249 controlled experiments over 1,247 query families and 45,336 target families.
  13. PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding: Ayaou and Cavallucci, arXiv:2510.22264, 2025; 15 tasks, 2.06 million examples, patembed-large result on DAPFAM.
  14. Can AI Examine Novelty of Patents? Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art: Ikoma and Mitamura, arXiv:2502.06316, 2025.
  15. Can Large Language Models Understand As Well As Apply Patent Regulations to Pass a Hands-On Patent Attorney Test?: Khera, Alamian, Scherz and Goetz, arXiv:2507.10576, 2025.
  16. Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models: Dahl, Magesh, Suzgun and Ho, arXiv:2401.01301; peer-reviewed in the Journal of Legal Analysis, vol. 16 (2024).
  17. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools: Magesh, Surani, Dahl, Suzgun, Manning and Ho, arXiv:2405.20362; peer-reviewed in the Journal of Empirical Legal Studies (2025).
  18. EPO, DOCDB bulk data: worldwide bibliographic data, abstracts, citations and simple families, with no full text or images.
  19. Google Patents, Coverage: document counts, the 22 full-text offices and the inclusion of Scholar, Books and the Prior Art Archive.
  20. EPO, Espacenet now offers more than 150 million freely accessible patent documents: 7 February 2024.
  21. WIPO PATENTSCOPE, data coverage: 128.8 million documents across 123 national and regional collections.
  22. USPTO, Patent Public Search FAQs: database coverage, the absence of semantic searching, and no access to foreign patent databases.
  23. USPTO, AI-based Similarity Search in PE2E Search (PDF): model description and training data.
  24. USPTO, Memorandum to the Patent Examining Corps, Advance notice of change to the MPEP regarding Similarity Search (PDF): 24 October 2025; requirement to use and record Similarity Search on plant and utility applications.
  25. USPTO, Guidance on Use of Artificial Intelligence-Based Tools in Practice Before the USPTO: 89 FR 25609, 11 April 2024.
  26. 37 CFR 11.18(b): eCFR: certifications made by the party presenting a paper to the USPTO.

Measure it, do not take it on trust

Run a family you have already prosecuted through the agent and compare the output against the examiner citations on the file. That single test tells you more than any published percentage.

Try Patsnap Eureka free

Your Agentic AI Partner
for Smarter Innovation

Patsnap fuses the world’s largest proprietary innovation dataset with cutting-edge AI to
supercharge R&D, IP strategy, materials science, and drug discovery.

Book a demo