Back to skill

Security audit

Zotero

Security checks for vulnerabilities and agentic risk

Overview

The Zotero skill is mostly transparent, but its PDF-fetching feature can download from unvalidated third-party URLs, which warrants Review before installing.

Install only if you are comfortable giving the skill access to your Zotero library. Prefer a read-only Zotero API key unless you need add/update/delete/upload features, use --dry-run, --limit, and --collection before batch operations, and avoid fetch-pdfs on networks where the agent can reach sensitive internal services until URL validation and download-size limits are added.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/zotero.py:918
Finding

Unrestricted Retrieval of Externally Supplied PDF URLs Enables SSRF

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (24)

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/zotero.py (reported line 94)May include surrounding context.

python
for attempt in range(_MAX_RETRIES + 1):
        req = urllib.request.Request(url, data=body, headers=headers, method=method)
        try:
            with urllib.request.urlopen(req, timeout=30) as resp:
                resp_body = resp.read().decode("utf-8")
                resp_headers = dict(resp.headers)
                return resp_body, resp_headers

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/zotero.py (reported line 1353)May include surrounding context.

python
for attempt in range(_MAX_RETRIES + 1):
        req = urllib.request.Request(url, data=body, headers=headers, method=method)
        try:
            with urllib.request.urlopen(req, timeout=30) as resp:
                resp_body = resp.read().decode("utf-8")
                resp_headers = dict(resp.headers)
                return resp_body, resp_headers

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/zotero.py (reported line 1371)May include surrounding context.

python
for attempt in range(_MAX_RETRIES + 1):
        req = urllib.request.Request(url, data=body, headers=headers, method=method)
        try:
            with urllib.request.urlopen(req, timeout=30) as resp:
                resp_body = resp.read().decode("utf-8")
                resp_headers = dict(resp.headers)
                return resp_body, resp_headers

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/zotero.py (reported line 1454)May include surrounding context.

python
for attempt in range(_MAX_RETRIES + 1):
        req = urllib.request.Request(url, data=body, headers=headers, method=method)
        try:
            with urllib.request.urlopen(req, timeout=30) as resp:
                resp_body = resp.read().decode("utf-8")
                resp_headers = dict(resp.headers)
                return resp_body, resp_headers

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/zotero.py (reported line 537)May include surrounding context.

python
url = f"https://api.crossref.org/works/{urllib.parse.quote(doi, safe='')}"
    req = urllib.request.Request(url, headers={"Accept": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=15) as resp:
            data = json.loads(resp.read().decode("utf-8"))
    except Exception as e:
        print(f"CrossRef lookup failed: {e}", file=sys.stderr)

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/zotero.py (reported line 954)May include surrounding context.

python
url = f"https://api.crossref.org/works/{urllib.parse.quote(doi, safe='')}"
    req = urllib.request.Request(url, headers={"Accept": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=15) as resp:
            data = json.loads(resp.read().decode("utf-8"))
    except Exception as e:
        print(f"CrossRef lookup failed: {e}", file=sys.stderr)

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

The find-dois feature sends item titles and first-author names from the user's Zotero library to CrossRef to search for missing DOIs. This is an external transmission of potentially sensitive reading/library metadata to a third party, which can expose research interests or unpublished work without an explicit per-action warning.

Content

Scanner excerpt · scripts/zotero.py (reported line 757)May include surrounding context.

python
url = "https://api.crossref.org/works?" + urllib.parse.urlencode(params)
    req = urllib.request.Request(url, headers={"Accept": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=20) as resp:
            data = json.loads(resp.read().decode("utf-8"))
        return data.get("message", {}).get("items", [])
    except (urllib.error.URLError, urllib.error.HTTPError, Exception) as e:

Tainted flow: 'req' from os.environ.get (line 92, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/zotero.py (reported line 811)May include surrounding context.

python
}
    body = json.dumps({field: value}).encode("utf-8")
    req = urllib.request.Request(url, data=body, headers=headers, method="PATCH")
    with urllib.request.urlopen(req, timeout=30) as resp:
        return resp.status

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

The fetch-pdfs feature sends item DOIs and a user-supplied contact email to Unpaywall. While this is part of the feature, it discloses library-derived identifiers and user identity metadata to a third party, which may be sensitive in some research contexts.

Content

Scanner excerpt · scripts/zotero.py (reported line 921)May include surrounding context.

python
url = f"https://api.unpaywall.org/v2/{urllib.parse.quote(doi, safe='')}?email={CROSSREF_EMAIL}"
    req = urllib.request.Request(url, headers={"Accept": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=15) as resp:
            data = json.loads(resp.read().decode("utf-8"))
        oa = data.get("best_oa_location") or {}
        pdf_url = oa.get("url_for_pdf")

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

This code queries Semantic Scholar with library-derived DOIs to discover open-access PDFs. Although expected for the feature, it externally discloses what papers exist in the user's collection to another service without strong in-band consent or minimization controls.

Content

Scanner excerpt · scripts/zotero.py (reported line 938)May include surrounding context.

python
url = f"https://api.semanticscholar.org/graph/v1/paper/DOI:{urllib.parse.quote(doi, safe='')}?fields=openAccessPdf"
    req = urllib.request.Request(url, headers={"Accept": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=15) as resp:
            data = json.loads(resp.read().decode("utf-8"))
        oa = data.get("openAccessPdf") or {}
        pdf_url = oa.get("url")

Tainted flow: 'req' from os.environ.get (line 1369, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
91% confidence
Finding

The tool downloads a PDF from a URL learned from third-party services and follows it directly with urllib, which can reach arbitrary hosts. Without host allowlisting or private-address blocking, a malicious or compromised metadata source could turn this into SSRF or cause downloads from untrusted internal/network locations.

Content

Scanner excerpt · scripts/zotero.py (reported line 988)May include surrounding context.

python
"Accept": "application/pdf,*/*",
    })
    try:
        with urllib.request.urlopen(req, timeout=60) as resp:
            with open(dest_path, "wb") as f:
                shutil.copyfileobj(resp, f)
        # Verify it's actually a PDF (check magic bytes)

Tainted flow: 'auth_req' from os.environ.get (line 1069, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/zotero.py (reported line 1074)May include surrounding context.

python
headers=auth_headers, method="POST",
    )
    try:
        with urllib.request.urlopen(auth_req, timeout=30) as resp:
            auth_data = json.loads(resp.read().decode("utf-8"))
    except urllib.error.HTTPError as e:
        print(f"    ⚠  Upload auth failed: {e.code}", file=sys.stderr)

Tainted flow: 'reg_req' from os.environ.get (line 1112, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/zotero.py (reported line 1117)May include surrounding context.

python
headers=reg_headers, method="POST",
    )
    try:
        with urllib.request.urlopen(reg_req, timeout=30) as resp:
            pass
        return True
    except urllib.error.HTTPError as e:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 32)May include surrounding context.

md
All operations use `scripts/zotero.py` (Python 3, zero external dependencies).

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 35)May include surrounding context.

md
All operations use `scripts/zotero.py` (Python 3, zero external dependencies).

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill advertises and documents capabilities that access environment secrets, read and write local files, and perform network operations, but it does not declare an explicit tool scope such as allowed-tools or permissions. That mismatch weakens policy enforcement and reviewability, increasing the risk that the skill can be invoked with broader capabilities than users or the platform expect.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The activation text is broad enough to match generic academic-reference and PDF-related requests, which can cause the skill to trigger in situations beyond explicit Zotero intent. In context, this matters because the skill supports destructive and sensitive actions such as deleting items, modifying metadata, exporting libraries, and fetching remote content, so overbroad routing raises the chance of unintended privileged actions.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
75% confidence
Finding

CrossRef fallback transmits a user-supplied DOI to a third party. In this skill context that is expected functionality, but it still constitutes external sharing of library/query metadata and may be sensitive depending on the user's research domain.

Content

Scanner excerpt · scripts/zotero.py (reported line 534)May include surrounding context.

python
def _doi_to_item(doi):
    """Fallback: fetch metadata from CrossRef for a DOI and convert to Zotero format."""
    url = f"https://api.crossref.org/works/{urllib.parse.quote(doi, safe='')}"
    req = urllib.request.Request(url, headers={"Accept": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=15) as resp:

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

The CrossRef search transmits titles and author names from the library to an external service to infer missing DOIs. That can reveal a user's reading list, manuscript references, or confidential research focus.

Content

Scanner excerpt · scripts/zotero.py (reported line 754)May include surrounding context.

python
}
    if first_author:
        params["query.author"] = first_author
    url = "https://api.crossref.org/works?" + urllib.parse.urlencode(params)
    req = urllib.request.Request(url, headers={"Accept": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=20) as resp:

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

Unpaywall requests disclose DOIs from the user's collection and include a contact email parameter. This is not malicious, but it is a real external data-sharing behavior that may matter for privacy-sensitive users.

Content

Scanner excerpt · scripts/zotero.py (reported line 918)May include surrounding context.

python
def _try_unpaywall(doi):
    """Try Unpaywall for an OA PDF URL. Returns (pdf_url, source_url) or None."""
    url = f"https://api.unpaywall.org/v2/{urllib.parse.quote(doi, safe='')}?email={CROSSREF_EMAIL}"
    req = urllib.request.Request(url, headers={"Accept": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=15) as resp:

External Transmission

Medium
Category
Data Exfiltration
Confidence
86% confidence
Finding

Semantic Scholar lookups share collection-derived DOIs with a third party to discover PDF URLs. In a citation-management skill, that is functionally relevant but still a privacy-impacting transmission of user library metadata.

Content

Scanner excerpt · scripts/zotero.py (reported line 935)May include surrounding context.

python
def _try_semantic_scholar(doi):
    """Try Semantic Scholar for an OA PDF URL. Returns (pdf_url, source_url) or None."""
    url = f"https://api.semanticscholar.org/graph/v1/paper/DOI:{urllib.parse.quote(doi, safe='')}?fields=openAccessPdf"
    req = urllib.request.Request(url, headers={"Accept": "application/json"})
    try:
        with urllib.request.urlopen(req, timeout=15) as resp:

Tainted flow: 'local_path' from os.environ.get (line 1271, credential/environment) → shutil.copy2 (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/zotero.py (reported line 1272)May include surrounding context.

python
# Save locally if requested
                if download_dir:
                    local_path = os.path.join(download_dir, pdf_filename)
                    shutil.copy2(tmp_path, local_path)
                    print(f"    💾 Saved: {local_path}")

                # Upload to Zotero storage if requested

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The export command can write an entire Zotero library or collection to any user-provided path without an explicit warning about local data disclosure. In agent or scripted contexts, this increases the risk of unintentionally materializing sensitive bibliography data into shared locations or synced directories.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · scripts/zotero.py (reported line 1622)May include surrounding context.

python
# delete
    p = subparsers.add_parser("delete", help="Move items to trash (default) or permanently delete")
    p.add_argument("keys", nargs="+", help="Item key(s) to delete")
    p.add_argument("--yes", action="store_true", help="Skip confirmation")
    p.add_argument("--permanent", action="store_true", help="Permanently delete (default is recoverable trash)")
    p.add_argument("--trash", action="store_true", help="Move to trash (default, kept for backwards compat)")

Static analysis

No suspicious patterns detected.