Back to skill

Security audit

Semantic Scholar Library Feed

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly coherent for Semantic Scholar account automation, but it stores live session cookies and can send them to attacker-chosen URLs.

Install only if you are comfortable giving the skill access to your Semantic Scholar session. Treat the saved cookie files like passwords, keep them out of shared folders, repos, logs, and backups, and avoid using --page-url with anything except https://www.semanticscholar.org until the skill constrains authenticated requests to the official origin.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/semantic_scholar_cli.py:253
Finding

Semantic Scholar session cookies can be transmitted to arbitrary origins

Content
View full analysis
dict[str, str]: headers = { "User-Agent": DEFAULT_USER_AGENT, "Accept": accept, } if bundle is not None: headers["Cookie"] = cookie_header_from_bundle(bundle) user_agent = str(bundle.get("userAgent", "")).strip() if user_agent: headers["User-Agent"] = user_agent if referer: headers["Referer"] = referer if content_type: headers["Content-Type"] = content_type return headers ``` ```python # scripts/ss_store.py:278-291 def fetch_html( url: str, bundle: dict[str, Any], *, referer: str | None = None, timeout: int = DEFAULT_TIMEOUT, ) -> str: response = request( url, headers=build_auth_headers(bundle, accept="text/html,application/xhtml+xml", referer=referer), timeout=timeout, ) require_ok(response, f"Fetching HTML from {url}") return response.text ``` ### Technical Analysis The `ssr-dump` and `feed-crawl` commands accept `--page-url` without validating its scheme, hostname, port, or or ...[truncated 2926 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ss_store.py:160
Finding

Reusable session credentials are duplicated in plaintext files without enforced restrictive permissions

Content
View full analysis
tuple[Path, Path]: path = normalize_cookie_path(cookie_path) ensure_cookie_dir(path.parent) path.write_text(json.dumps(bundle, ensure_ascii=False, indent=2) + "\n") header_path = header_path_for_cookie_path(path) header = str(bundle.get("cookieHeader", "")).strip() header_path.write_text(f"{header}\n") return path, header_path ``` The directory helper also relies on ambient permissions rather than explicitly enforcing owner-only access: ```python def ensure_cookie_dir(path: str | Path = DEFAULT_COOKIE_DIR) -> Path: cookie_dir = Path(path).expanduser().resolve() cookie_dir.mkdir(parents=True, exist_ok=True) return cookie_dir ``` ### Technical Analysis The Skill stores reusable Semantic Scholar session cookies in two plaintext locations: a JSON bundle and a raw cookie-header file. Both files contain the full authentication values. File creation uses `Path.write_text()` without explicitly assigning owner-only permissions, and directory creation uses `mkdir()` without enforcing mode `0700`. The resulting permissions depend on the process umask and any permissions already present on the directory or files. Existing files with permissive modes are overwritten without first correcting those modes. A caller may also override `--cookie-path`, allowing the credentials to be written outside the default authentication directory. Plaintext storage is partly inherent to the CLI's session-reuse design, but duplicating credentials and failing to enforce restrictive access controls are unnecessary. The raw-header copy expands the number of local locations from which account credentials can be recovered. ### Attack Path 1. ...[truncated 1085 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (10)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill exposes meaningful capabilities—shell execution, network access, and file read/write—yet declares no explicit tool scope or permission boundaries. That increases the chance an agent can invoke broader-than-necessary actions, including handling local credential files and making authenticated requests, without a clear least-privilege contract.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Importing a browser-copied authenticated curl command can capture full session cookies and possibly other sensitive headers, then normalize them for reuse by the skill. Without a strong privacy warning and minimization guidance, users may unknowingly hand over active session tokens that permit unauthorized access if logged, stored insecurely, or reused outside the intended task.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs storing authenticated Semantic Scholar cookies in predictable local files, including reusable session material (sid, s2), without prominently warning that these are equivalent to account credentials. If other tools, users, or processes can read those files, an attacker could hijack the user's private Library and Research Feed access.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/library.md (reported line 52)May include surrounding context.

md
When BibTeX entries already contain arXiv identifiers, prefer:

- `POST https://api.semanticscholar.org/graph/v1/paper/batch?fields=paperId,title,year,url,externalIds`

Pass identifiers like:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The workflow explicitly instructs an authenticated crawl of a user's private Semantic Scholar Research Feed and incremental export to a local file, but it provides no consent, scope-limiting, or data-handling safeguards. In an agent setting, this creates a real privacy and data-exfiltration risk because the skill operationalizes access to account-scoped content and persistence of that content outside the service.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The CLI persists a raw Cookie header containing live authentication cookies to disk, which creates a durable local secret that can be stolen by other local users, malware, backups, logs, or accidental file sharing. In this skill’s context, those cookies grant access to a user’s private Semantic Scholar account data, so compromise of the stored file can lead to account/session takeover and exposure of private library and feed information.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This command can write decoded contents from an authenticated private Semantic Scholar page to an arbitrary local file without clearly warning that the output may contain sensitive account data. That increases the chance of unintentional local data exposure through shared directories, source repositories, cloud sync, or backups.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The feed crawler persists aggregated recommendation/feed data from authenticated requests to disk during execution, and this data can reflect a user’s private reading interests, folders, and recommendation history. Because the skill is specifically designed to access private Library and Research Feed content, silent persistence meaningfully raises privacy risk if the file is later exposed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code persists an authenticated Semantic Scholar cookie bundle and also writes a raw Cookie header string to disk, which materially increases the chance of credential/session theft if the files are read by another local user, backup system, malware, or logs. In this skill’s context, the cookies grant access to private library/feed data and likely account actions, so storing them in plaintext without restrictive permissions or explicit consent is a real security weakness.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

When a bundle is provided, build_auth_headers injects the Cookie header into outbound requests, and the request/fetch helpers send those headers over the network. This is a network operation involving credential material, but the file contains no confirmation prompt, user-visible logging, or explanatory comment/docstring disclosing that saved authentication cookies will be transmitted.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.