T09 · Insecure Skill Coding Practices
- Location
scripts/zotero.py:918- Finding
Unrestricted Retrieval of Externally Supplied PDF URLs Enables SSRF
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The Zotero skill is mostly transparent, but its PDF-fetching feature can download from unvalidated third-party URLs, which warrants Review before installing.
Install only if you are comfortable giving the skill access to your Zotero library. Prefer a read-only Zotero API key unless you need add/update/delete/upload features, use --dry-run, --limit, and --collection before batch operations, and avoid fetch-pdfs on networks where the agent can reach sensitive internal services until URL validation and download-size limits are added.
scripts/zotero.py:918Unrestricted Retrieval of Externally Supplied PDF URLs Enables SSRF
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
for attempt in range(_MAX_RETRIES + 1):
req = urllib.request.Request(url, data=body, headers=headers, method=method)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
resp_body = resp.read().decode("utf-8")
resp_headers = dict(resp.headers)
return resp_body, resp_headers
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
for attempt in range(_MAX_RETRIES + 1):
req = urllib.request.Request(url, data=body, headers=headers, method=method)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
resp_body = resp.read().decode("utf-8")
resp_headers = dict(resp.headers)
return resp_body, resp_headers
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
for attempt in range(_MAX_RETRIES + 1):
req = urllib.request.Request(url, data=body, headers=headers, method=method)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
resp_body = resp.read().decode("utf-8")
resp_headers = dict(resp.headers)
return resp_body, resp_headers
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
for attempt in range(_MAX_RETRIES + 1):
req = urllib.request.Request(url, data=body, headers=headers, method=method)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
resp_body = resp.read().decode("utf-8")
resp_headers = dict(resp.headers)
return resp_body, resp_headers
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
url = f"https://api.crossref.org/works/{urllib.parse.quote(doi, safe='')}"
req = urllib.request.Request(url, headers={"Accept": "application/json"})
try:
with urllib.request.urlopen(req, timeout=15) as resp:
data = json.loads(resp.read().decode("utf-8"))
except Exception as e:
print(f"CrossRef lookup failed: {e}", file=sys.stderr)
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
url = f"https://api.crossref.org/works/{urllib.parse.quote(doi, safe='')}"
req = urllib.request.Request(url, headers={"Accept": "application/json"})
try:
with urllib.request.urlopen(req, timeout=15) as resp:
data = json.loads(resp.read().decode("utf-8"))
except Exception as e:
print(f"CrossRef lookup failed: {e}", file=sys.stderr)
The find-dois feature sends item titles and first-author names from the user's Zotero library to CrossRef to search for missing DOIs. This is an external transmission of potentially sensitive reading/library metadata to a third party, which can expose research interests or unpublished work without an explicit per-action warning.
url = "https://api.crossref.org/works?" + urllib.parse.urlencode(params)
req = urllib.request.Request(url, headers={"Accept": "application/json"})
try:
with urllib.request.urlopen(req, timeout=20) as resp:
data = json.loads(resp.read().decode("utf-8"))
return data.get("message", {}).get("items", [])
except (urllib.error.URLError, urllib.error.HTTPError, Exception) as e:
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
}
body = json.dumps({field: value}).encode("utf-8")
req = urllib.request.Request(url, data=body, headers=headers, method="PATCH")
with urllib.request.urlopen(req, timeout=30) as resp:
return resp.status
The fetch-pdfs feature sends item DOIs and a user-supplied contact email to Unpaywall. While this is part of the feature, it discloses library-derived identifiers and user identity metadata to a third party, which may be sensitive in some research contexts.
url = f"https://api.unpaywall.org/v2/{urllib.parse.quote(doi, safe='')}?email={CROSSREF_EMAIL}"
req = urllib.request.Request(url, headers={"Accept": "application/json"})
try:
with urllib.request.urlopen(req, timeout=15) as resp:
data = json.loads(resp.read().decode("utf-8"))
oa = data.get("best_oa_location") or {}
pdf_url = oa.get("url_for_pdf")
This code queries Semantic Scholar with library-derived DOIs to discover open-access PDFs. Although expected for the feature, it externally discloses what papers exist in the user's collection to another service without strong in-band consent or minimization controls.
url = f"https://api.semanticscholar.org/graph/v1/paper/DOI:{urllib.parse.quote(doi, safe='')}?fields=openAccessPdf"
req = urllib.request.Request(url, headers={"Accept": "application/json"})
try:
with urllib.request.urlopen(req, timeout=15) as resp:
data = json.loads(resp.read().decode("utf-8"))
oa = data.get("openAccessPdf") or {}
pdf_url = oa.get("url")
The tool downloads a PDF from a URL learned from third-party services and follows it directly with urllib, which can reach arbitrary hosts. Without host allowlisting or private-address blocking, a malicious or compromised metadata source could turn this into SSRF or cause downloads from untrusted internal/network locations.
"Accept": "application/pdf,*/*",
})
try:
with urllib.request.urlopen(req, timeout=60) as resp:
with open(dest_path, "wb") as f:
shutil.copyfileobj(resp, f)
# Verify it's actually a PDF (check magic bytes)
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
headers=auth_headers, method="POST",
)
try:
with urllib.request.urlopen(auth_req, timeout=30) as resp:
auth_data = json.loads(resp.read().decode("utf-8"))
except urllib.error.HTTPError as e:
print(f" ⚠ Upload auth failed: {e.code}", file=sys.stderr)
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
headers=reg_headers, method="POST",
)
try:
with urllib.request.urlopen(reg_req, timeout=30) as resp:
pass
return True
except urllib.error.HTTPError as e:
Referenced artifact was not completely inspected
All operations use `scripts/zotero.py` (Python 3, zero external dependencies).
Referenced artifact was not completely inspected
All operations use `scripts/zotero.py` (Python 3, zero external dependencies).
The skill advertises and documents capabilities that access environment secrets, read and write local files, and perform network operations, but it does not declare an explicit tool scope such as allowed-tools or permissions. That mismatch weakens policy enforcement and reviewability, increasing the risk that the skill can be invoked with broader capabilities than users or the platform expect.
The activation text is broad enough to match generic academic-reference and PDF-related requests, which can cause the skill to trigger in situations beyond explicit Zotero intent. In context, this matters because the skill supports destructive and sensitive actions such as deleting items, modifying metadata, exporting libraries, and fetching remote content, so overbroad routing raises the chance of unintended privileged actions.
CrossRef fallback transmits a user-supplied DOI to a third party. In this skill context that is expected functionality, but it still constitutes external sharing of library/query metadata and may be sensitive depending on the user's research domain.
def _doi_to_item(doi):
"""Fallback: fetch metadata from CrossRef for a DOI and convert to Zotero format."""
url = f"https://api.crossref.org/works/{urllib.parse.quote(doi, safe='')}"
req = urllib.request.Request(url, headers={"Accept": "application/json"})
try:
with urllib.request.urlopen(req, timeout=15) as resp:
The CrossRef search transmits titles and author names from the library to an external service to infer missing DOIs. That can reveal a user's reading list, manuscript references, or confidential research focus.
}
if first_author:
params["query.author"] = first_author
url = "https://api.crossref.org/works?" + urllib.parse.urlencode(params)
req = urllib.request.Request(url, headers={"Accept": "application/json"})
try:
with urllib.request.urlopen(req, timeout=20) as resp:
Unpaywall requests disclose DOIs from the user's collection and include a contact email parameter. This is not malicious, but it is a real external data-sharing behavior that may matter for privacy-sensitive users.
def _try_unpaywall(doi):
"""Try Unpaywall for an OA PDF URL. Returns (pdf_url, source_url) or None."""
url = f"https://api.unpaywall.org/v2/{urllib.parse.quote(doi, safe='')}?email={CROSSREF_EMAIL}"
req = urllib.request.Request(url, headers={"Accept": "application/json"})
try:
with urllib.request.urlopen(req, timeout=15) as resp:
Semantic Scholar lookups share collection-derived DOIs with a third party to discover PDF URLs. In a citation-management skill, that is functionally relevant but still a privacy-impacting transmission of user library metadata.
def _try_semantic_scholar(doi):
"""Try Semantic Scholar for an OA PDF URL. Returns (pdf_url, source_url) or None."""
url = f"https://api.semanticscholar.org/graph/v1/paper/DOI:{urllib.parse.quote(doi, safe='')}?fields=openAccessPdf"
req = urllib.request.Request(url, headers={"Accept": "application/json"})
try:
with urllib.request.urlopen(req, timeout=15) as resp:
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
# Save locally if requested
if download_dir:
local_path = os.path.join(download_dir, pdf_filename)
shutil.copy2(tmp_path, local_path)
print(f" 💾 Saved: {local_path}")
# Upload to Zotero storage if requested
The export command can write an entire Zotero library or collection to any user-provided path without an explicit warning about local data disclosure. In agent or scripted contexts, this increases the risk of unintentionally materializing sensitive bibliography data into shared locations or synced directories.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
# delete
p = subparsers.add_parser("delete", help="Move items to trash (default) or permanently delete")
p.add_argument("keys", nargs="+", help="Item key(s) to delete")
p.add_argument("--yes", action="store_true", help="Skip confirmation")
p.add_argument("--permanent", action="store_true", help="Permanently delete (default is recoverable trash)")
p.add_argument("--trash", action="store_true", help="Move to trash (default, kept for backwards compat)")
No suspicious patterns detected.