Install
openclaw skills install @docsor1212/pubmed-verifierReference checker for AI-fabricated citations: batch-verify PMIDs against PubMed and catch the hallucination existence checks miss — a REAL PMID pointing to a DIFFERENT paper. Five-state citation verification (correct / mismatch / partial / invalid / unknown), citation-context parsing, dual fuzzy matching, Crossref DOI cross-check, retraction detection (capped at partial), correct-PMID suggestion, arXiv ID verification, SQLite cache, CSV/JSON claims, HTML/JSON/text reports. Dual data sources with automatic Europe PMC fallback, optional NCBI API key, Crossref polite pool, Retry-After backoff, UA rotation, host circuit breaker. Network failures are honestly reported as unverified, never as "not found". Zero dependencies, runs fully local. Triggers: verify PMIDs, check citations, validate references, citation audit, reference check, PMID check, audit references, batch verify references, AI hallucination detection, verify DOI, DOI check, validate citations, PubMed citation verifier, BibTeX audit.
openclaw skills install @docsor1212/pubmed-verifierBatch verification of PMID citations via the PubMed E-utilities API. Not just "does this PMID exist" — does this PMID point to the paper you claim? Zero dependencies, pure standard library, fully local.
Invoke it whenever citation truth matters:
--claims-file
when they supply the expected titles.Trigger priority & tool choice — explicit "verify / check / audit citations, references, PMIDs, DOIs" requests invoke this skill first. cite-holmes is for deep research with machine-verified citations; when a request mixes research and verification, run the research first, then this tool for the final reference audit.
| Verdict | Meaning |
|---|---|
| ✅ Correct | PMID exists AND matches the claimed paper |
| ⚠️ Mismatch | PMID exists but points to a different paper (the most common AI hallucination!) |
| 🔶 Partial | Some metadata matches (e.g. author+journal but title differs) |
| ❌ Invalid | PMID does not exist in PubMed |
| ❓ Unknown | Not enough claimed metadata to cross-check — or both data sources unreachable (never misreported as invalid) |
Why existence checks are not enough: a large share of fabricated citations use REAL PMIDs that point to a different paper from the same year/journal/field — in one of our own audits, 4 of 5 "valid" PMIDs were wrong this way. A binary exists/not-exists check misses them all.
# Scan a project directory for PMIDs (parses citation context automatically)
python3 scripts/verify_pmids.py --source /path/to/project --output report.html
# Verify specific PMIDs
python3 scripts/verify_pmids.py --pmids 31018962,22213727
# Mismatch demo: PMID 34078778 is actually a dental-materials paper, so the
# JIA claims below will NOT match it — expect ⚠️ mismatch verdicts
python3 scripts/verify_pmids.py --claims '[{"pmid":"34078778","title":"JIA pathogenesis","authors":["Zaripova"],"journal":"Pediatr Rheumatol Online J","year":"2021"}]' --output report.html
# Claims from a CSV file + suggest correct PMIDs for mismatches
python3 scripts/verify_pmids.py --claims-file claims.csv --suggest --output report.html
# Crossref DOI cross-verification + audit working-paper + BibTeX + full pipeline
python3 scripts/verify_pmids.py --source /path/to/files --verify-doi --suggest --output report.html --export-audit audit.json --export-bibtex refs.bib
# Verify DOIs directly (no PMIDs) + delta audit vs a previous run
python3 scripts/verify_pmids.py --dois "10.1038/nature12968,10.4012/dmj.2020-408" --workers 4 --export-audit audit.json --export-csv table.csv
python3 scripts/verify_pmids.py --source /path/to/project --diff audit.json --output report.html
# Verify arXiv IDs (preprints) — mixed audits supported
python3 scripts/verify_pmids.py --arxivs "2401.12345,cs/0211004" --no-cache
# Pre-flight: are the four data sources reachable right now?
python3 scripts/verify_pmids.py --check-net
# Audit a BibTeX or RIS bibliography file directly (PMID > DOI > arXiv routing)
python3 scripts/verify_pmids.py --bibliography refs.bib --no-cache
python3 scripts/verify_pmids.py --bibliography refs.ris --no-cache --export-ris verified.ris
# Institutional niceties (recommended): NCBI API key + contact email
python3 scripts/verify_pmids.py --source . --verify-doi --ncbi-api-key $NCBI_API_KEY --mailto you@lab.org
Top 10 things NOT to do (each is detailed below or in Anti-patterns):
| # | Don't | Do instead |
|---|---|---|
| 1 | Treat --pmids existence output as "verified" | Feed --claims-file with titles for real verification |
| 2 | Submit claims without title | Always include titles — the verdict caps at partial without one |
| 3 | Trust cached verdicts on publication day | Final check with --no-cache |
| 4 | Read "not found" as "fabricated" for auto-extracted DOIs | Check doi.org / arxiv.org by hand first |
| 5 | Treat the leading ' in CSV cells as corruption | It is the formula-injection guard — strip after import |
| 6 | Read the READY line as a quality score | It means "no problems among the checks that ran" |
| 7 | Pass --source together with --pmids | --source is ignored entirely when --pmids is given |
| 8 | Deep-verify (--verify-doi / --suggest) a thousand-entry sweep | Sweep first, deep-verify the flagged subset |
| 9 | Expect author matching across CJK↔Latin names | They are skipped honestly (author_check: skipped) |
| 10 | Ship a reference list without the audit trail | --export-audit writes a replayable working paper |
Large batch (hundreds of PMIDs) is slow — how to speed it up?
Metadata-only verification queries in batches of 50 with 0.4 s spacing
(0.12 s with --ncbi-api-key); cached re-runs are ~5 s. --verify-doi adds
one Crossref call per citation and --suggest adds one search per
mismatch — skip them for bulk sweeps, run them on the flagged subset.
When must I use --claims-file instead of scanning?
Context parsing is heuristic (v3.3.0 guards common abbreviations, exotic formatting can still mis-split).
For precise verification — or DOIs in claims (splice detection needs doi)
— feed structured JSON/CSV claims.
My citation text is in Chinese — the scan parses little?
GB/T 7714-style references (……标题[J]. 刊名, 年… PMID: xxx) parse since
v3.9.0 — title/journal comparisons against Latin registries are skipped
cross-language, so verdicts rest on year/DOI evidence. Free Chinese prose
without a PMID marker still yields little — feed --claims-file for full
verdicts regardless of language.
Slow or unstable network (China)?
Standard HTTPS_PROXY/HTTP_PROXY env vars are honored natively; raise
--timeout; --meta-source europepmc routes via Europe PMC when NCBI is
unreachable (per-entry meta_source shows which was used); cached results
are reused for 30 days.
❓ unknown vs ❌ invalid?
unknown (exit 2) = "could not verify, sources unreachable" — retry later;
invalid (exit 1) = "verified not-found". Network failures are never
reported as not-found and never cached.
What does RETRACTED mean in a report?
The registry itself lists the paper's publication type as "Retracted
Publication" (checked for every citation since v2.7.0 — no DOI or flags
needed), and/or Crossref records a retraction. The verdict is capped at
partial and a human review note is attached — citing it would propagate
withdrawn science. A retraction notice is never flagged; papers under
Expression of Concern (an editorial note, not a retraction) are not
flagged either. Retraction status reflects the registry at cache time — for
a final pre-submission check, run with --no-cache.
Mismatch reported but the title looks similar?
Check details for which field diverged; thresholds are strict on purpose.
Feed the full citation via --claims-file for a precise verdict.
Can I audit my .bib or .ris file directly?
Yes — --bibliography refs.bib (BibTeX) and --bibliography refs.ris
(RIS/Zotero/EndNote/Mendeley) route each entry by PMID > DOI > arXiv
and cross-check the claimed metadata. --lint-claims validates either
format offline first. The round trip works both ways: --export-bibtex
and --export-ris output can be fed back after edits.
My claims file seems to lose rows / verdicts look weaker than expected?
Lint it offline first: python3 scripts/verify_pmids.py --lint-claims claims.csv reports unusable rows, ID shape errors, unknown columns
(typo'd headers like "titel"), missing titles and DOI prefix problems —
no network, exit 1 on errors.
Can I verify a DOI with claimed metadata (full verdict)?
Yes — since v3.4.0 a claims row keyed by doi (with title, optionally
authors/journal/year) gets the same cross-check as PMID claims:
correct / mismatch / partial against the registered metadata. A DOI row
without claims stays unknown (existence only).
Why did my arXiv citation drop from correct to partial?
Your claims row paired an arxiv_id with a doi, and the DOI does not
match the version-of-record DOI registered on that arXiv entry — a typo,
or a DOI from a different paper. The registered DOI is in details;
fix the claim or drop the doi cell.
Each entry: the mistake → why it fails → the right way.
--pmids output as "fully verified" — existence-only.
→ Wrong: "all 5 PMIDs exist, so the citations are correct."
→ Right: existence-checked only; feed --claims-file with titles for
real verification (the READY line says so explicitly).title — author/journal/year alone can never reach
correct; the report caps at partial. → Always include titles.--no-cache.' from CSV cells — that apostrophe is the
formula-injection guard, not data corruption. → Strip it after import.What this tool can NOT do, consolidated in one place:
author_check: skipped); the verdict rests
on title/journal/year alone.--claims-file..bib files commonly hold DataCite/repository DOIs — check doi.org
by hand); only --dois/claims DOIs count a Crossref 404 as invalid.--no-cache.unknown; network failures are never reported as "not found" and
never cached.unknown,
not partial; an explicitly user-provided DOI (--dois or claims) that
is missing from Crossref counts as invalid.--claims-file accepts JSON (an array of objects) or CSV. Recognized
columns: pmid, title, authors (semicolon/pipe-separated),
journal, year, doi, arxiv_id. A row needs one of pmid,
arxiv_id or doi; a missing title caps the verdict at partial.
pmid,title,authors,journal,year,doi,arxiv_id
31018962,Candidate criteria for diagnosis of familial...,Gattorno,Ann Rheum Dis,2019,10.1136/annrheumdis-2019-215048,
,Attention Is All You Need,Vaswani,NeurIPS,2017,,1706.03762
,City size and the spreading of COVID-19 in Brazil,Silva Junior;Other,PLOS ONE,2020,10.1371/journal.pone.0239699,
Validate any file offline first: --lint-claims file.csv reports
unusable rows, ID shape errors, unknown columns, duplicates and missing
titles (no network). Lint wins when combined with verification flags —
only the lint runs.
--workers (default 4, cap 8) applies to
Crossref resolution and Europe PMC linking; arXiv stays serial by
official etiquette (≥3 s between calls).--verify-doi, no --suggest), then deep-verify only the flagged
subset; each deep flag adds one API call per citation.--export-bibtex / --export-ris
write verified entries; after edits, --bibliography refs.bib /
refs.ris re-audits the file (the PMID re-links the full record).--claims-file: it enables the full verdict ladder and the DOI /
arXiv pairing checks. Validate the file offline first:
python3 scripts/verify_pmids.py --lint-claims claims.csv reports ID
shape errors, missing titles, unknown columns and duplicates without any
network access.--timeout; HTTPS_PROXY/HTTP_PROXY are
honored natively; unreachable NCBI falls back to Europe PMC
automatically (meta_source shows which answered).--cache-days to tune; --no-cache for the final pre-submission pass.scripts/verify_pmids.py anywhere with
Python 3.8+ and it runs, no pip, no venv (that portability is why the
code is not split into modules).| Feature | Flag | Effect |
|---|---|---|
| NCBI API key | --ncbi-api-key / env NCBI_API_KEY | Rate ceiling 3→10 req/s, batch interval 0.4s→0.12s (~3x faster) |
| Europe PMC fallback | --meta-source auto|ncbi|europepmc | NCBI batch failure automatically retries via Europe PMC (free, no key); per-entry origin in JSON (meta_source) |
| Crossref polite pool | --mailto / env PUBMED_VERIFIER_MAILTO | ?mailto= on Crossref + tool/email params on NCBI — more generous limits |
| Retry-After backoff | automatic | 429 responses honored (clamped 1–5 s) instead of failing |
| UA rotation | automatic | 403/406 retried with a browser User-Agent |
| Host circuit breaker | automatic | After 2 call-level transport failures a host is skipped with an actionable message; success resets; HTTP errors never trip it |
| Honest unknown | automatic | Network failures report as ❓ unknown + exit code 2, never as "PMID not found", and are never cached |
Exit codes: 0 clean · 1 problems found (invalid / mismatch / retracted /
DOI-splice, incl. arXiv claimed-DOI pairing mismatches) · 2 could not verify
(data sources unreachable) — automation can tell "all good" from "no answer".
With --verify-doi, each cited DOI is also checked against Crossref's
withdrawal records (updated-by). A paper Crossref lists as RETRACTED is:
retracted: true + retraction_note) and in reports,Corrections and other update types do not trigger the cap. Crossref outages never flag anything (a missing check is not a retraction).
doi field to your claims (JSON or CSV).
The claimed DOI is compared with the DOI registered for that PMID — a
mismatch is a splice/fabrication signal (a real DOI attached to the wrong
paper): flagged in JSON (doi_splice_suspect), verdict capped at 🔶
partial, counted in exit 1. Uses the PubMed record only — no extra API call.Known limits: highly ambiguous abbreviations can over-match at the journal-only level ("J Immunol" ~ "Journal of Immunology Research") — the title remains the decisive field. A DOI-splice flag can also appear on an otherwise-unverifiable citation (the DOI mismatch is an independent fact).
author_check: skipped and the verdict falls back
to what was actually comparable (title/journal/year), or to ❓ unknown
when nothing else is checkable.eutils.ncbi.nlm.nih.gov, www.ebi.ac.uk (Europe PMC),
api.crossref.org, export.arxiv.org. No other hosts are contacted; no telemetry, no
analytics, no data collection — the only outbound payloads are the PMIDs,
DOIs and titles you asked to verify.~/.cache/pubmed-verifier/ (--no-cache to disable).NCBI_API_KEY / PUBMED_VERIFIER_MAILTO
authenticate or attribute your own API requests and are never sent
anywhere else.--export-audit audit.json — a self-contained JSON working-paper for
transparent review: tool identity and version, the exact (API-key-redacted)
invocation, per-citation evidence chains (claimed vs registered fields,
title match scores from both algorithms, author match with cross-language
skip records, DOI cross-check, retraction signals) and the
verdict-ladder trace for every citation. A reviewer can replay the entire
verification from this file alone.--verify-doi required, and the flag survives the cache
(schema v3). Crossref updated-by remains the detail source (the
retraction-notice DOI) when --verify-doi is on. A retraction notice
itself is never flagged.--export-bibtex refs.bib — export the verified bibliography: correct
entries as @article, partial entries commented out with their divergence
note, mismatched/invalid/unknown/retracted entries excluded and counted.SUBMISSION READY or NOT SUBMISSION-READY — <per-problem counts>.retracted column, auto-migrated).--dois "10.x/a, 10.y/b" — verify DOIs natively, no PMID required.
--source scans now also extract DOIs from your files automatically.
Each DOI is resolved via the Crossref works API: not-found on an explicitly
provided DOI = fabrication signal (invalid, exit 1); on one auto-extracted
from scanned text it stays a suspect (unknown) — scanned strings are never
user-endorsed, and DataCite/repository DOIs do not live in Crossref, so
always double-check at doi.org. Resolved = existence confirmed with the
registered metadata attached for manual comparison — existence is never
dressed up as a match.--diff previous-audit.json — delta audit against a previous working
paper: newly retracted (the safety signal — a paper retracted after
your last audit; act on it: swap or drop the citation, cite the retraction
notice instead, and re-check any conclusion that relied on it), degraded,
improved, new and dropped citations, with counts in every report format.
Built for periodic knowledge-base audits: "what changed since last time?"--workers N (default 4, max 8) resolves
DOI batches on a thread pool (roughly 3x faster on large lists), with
live progress output. For large --dois batches, set --mailto to stay
in Crossref's polite pool.--export-csv table.csv — spreadsheet-friendly audit table
(key/verdict/flags/fields/details; formula-injection hardened).Reference lists carry preprints. v3.0.0 verifies arXiv IDs alongside
PMIDs and DOIs: arXiv:2401.12345 and arxiv.org/abs/... patterns are
extracted from scans (or passed via --arxivs), checked against the
official arXiv API, and judged — nonexistent ID = fabrication signal
(invalid, exit 1); resolving ID = registered title/year attached, verdict
stays unknown. Malformed IDs (bad YYMM month) are flagged by shape.
Timely: arXiv penalizes submissions containing hallucinated or unverified
references (2026-05 policy) — audit before you submit.
arXiv entries carry the version-of-record DOI their authors registered at
publication (arxiv:doi). v3.3.0 puts it to work:
arxiv_id
and doi is cross-checked: agreement is reported as evidence
(fields.doi ✓); disagreement caps the verdict at partial — the DOI
belongs to a different paper (same failure class as PMID DOI-splice).Pairing example (match → correct with DOI evidence; wrong DOI → partial):
python3 scripts/verify_pmids.py --claims '[{"arxiv_id":"2005.13892",
"title":"City size and the spreading of COVID-19 in Brazil",
"doi":"10.1371/journal.pone.0239699"}]'
Claims rows could carry a PMID or an arXiv ID — a row keyed by DOI alone
was silently ignored, and DOI entries always stayed unknown ("no claimed
metadata to cross-verify"). v3.4.0 closes the matrix: all three citation
types now accept claimed metadata.
doi (no PMID, no arXiv ID) is cross-checked against
the registered metadata — the linked PubMed record when the DOI resolves
to one, Crossref otherwise — and gets the full verdict ladder:
correct / mismatch / partial.python3 scripts/verify_pmids.py --claims '[{"doi":"10.1371/journal.pone.0239699",
"title":"City size and the spreading of COVID-19 in Brazil",
"journal":"PLoS ONE","year":"2020"}]'
--lint-claims FILE — offline pre-flight for claims files
(JSON/CSV, zero network): ID shape errors, missing titles (the verdict
would cap at partial), unknown/typo'd columns, DOI prefix checks,
unusable rows — exit 1 on errors. Fix the format before the run
instead of guessing from weak verdicts.--bibliography refs.bib audits a .bib file directly: entries route by
PMID (the pmid field, or a "PMID: NNNN" in the note — including the
notes --export-bibtex itself writes) > DOI > arXiv (eprint, or an
"arXiv:XXXX.XXXXX" in the journal/note), each carrying its claimed
title/authors/journal/year for the full cross-check. Entries without any
routable ID are counted and skipped. This closes the loop with
--export-bibtex: a verified bibliography can be re-audited after edits.
--lint-claims refs.bib validates the file offline (unroutable entries,
missing titles, unclosed blocks).
--bibliography now accepts RIS files (refs.ris — Zotero/EndNote/
Mendeley exports) alongside BibTeX: records route by PMID (AN tag, or a
"PMID: NNNN" note) > DOI (DO) > arXiv (UR/eprint), each with claimed
metadata for the full cross-check. --export-ris completes the loop —
verified entries as TY JOUR, partials as TY DATA with a PARTIAL note.
--lint-claims refs.ris validates offline (unroutable records, missing
titles, duplicates, unterminated records).
An independent real-data audit (24 scenarios, registry-ground-truthed) found three judgment-ladder defects; all three are fixed and locked:
--lint-claims to catch duplicate claim rows offline); the
cross-language skip no longer coexists with contradictory details text;
empty --claims gets a one-line hint; docs state that Chinese-language
citation contexts parse weakly — use --claims-file.……标题[J]. 刊名, 年… PMID: xxx
contexts now yield the title (via the [J]/[M]/[R] type marker),
authors, journal, year and any DOI that follows the PMID. The title
comparison against a Latin registry title is skipped cross-language
(never a mismatch source); a DOI in the reference is cross-checked
against the registry like any claims DOI. Requires a PMID marker per
reference.--check-net — probes the four data sources (5 s each), shows your
proxy state and prints practical next steps when something is
unreachable. Run it when results come back unknown on a constrained
network.references/python_api.md documents the
stable importable surfaces (fetch_summaries, cross_check_citation,
verify_doi_entry, parse_citation_context, suggest_correct_pmid)
with copy-paste examples.PMID: 12345678 / PubMed URLs in
.html .md .txt .htm .json, and parses the surrounding reference into
claimed authors / title / journal / year.--cache-days), 3 retries with backoff.
Europe PMC steps in per failed batch when NCBI is unreachable.doi) — a claimed
DOI differing from the PMID's registered DOI is a splice/fabrication
signal (capped at partial).--verify-doi) — resolves each
cited DOI via Crossref, compares the registered title with the PubMed
record (doi_title_match in JSON), and detects RETRACTED papers (verdict
capped at partial). A doi_verified: false with note
"crossref unreachable" is a network fact, not a verdict.--suggest) — for mismatches,
searches PubMed with the claimed metadata and proposes top-3 candidates.
(Suggestion search always uses NCBI, even with --meta-source europepmc.)Context parsing is heuristic — since v3.3.0, common abbreviations
("U.S.", "e.g.", "vs.") no longer split a title, but exotic formatting
still can. For precise verification, feed structured claims via
--claims-file.
| Report | Flag | Use |
|---|---|---|
| HTML | --output report.html | Human review: claimed vs actual side by side |
| JSON | --output report.json | Programmatic processing (includes meta_source per entry) |
| Text | default | Quick terminal look |
Measured on a 225-PMID audit (5 esummary batches): metadata-only verification
runs in seconds; cached re-runs take ~5 s. Each batch waits 0.4 s between
calls (0.12 s with --ncbi-api-key). Optional extras are per-citation:
--verify-doi adds one Crossref call (~0.5–1 s) per cited DOI, and
--suggest adds one PubMed search per mismatch.
They share the same five-state philosophy and are safe to use together.
Each tool solves one step of reference work; use whichever fits the task.
Typical order: get papers (cn-med-oa), verify citations (this tool or cite-holmes), polish (paper-polisher-pro), make figures (academic-figures), translate PDFs (doc-holmes) — pick whichever step you need.
| File | Purpose |
|---|---|
scripts/verify_pmids.py | Main verifier (v3.9.0, stdlib-only) |
references/api_examples.md | PubMed / Europe PMC / Crossref / arXiv API notes |
references/python_api.md | Calling the verifier from Python (stable surfaces + examples) |
examples/claims.sample.csv | Reference format for --claims-file (incl. a DOI-only row) |
examples/refs.sample.bib | Sample bibliography for --bibliography (DOI/PMID/arXiv routing) |
examples/refs.sample.ris | Sample RIS bibliography (Zotero/EndNote/Mendeley) |
tests/ | Offline matrix + real-network acceptance (repo only, not in the package) |
MIT-0 — free to use, modify and redistribute, no attribution required.