T09 · Insecure Skill Coding Practices
- Location
scripts/extract.py:62- Finding
Scientific Paper Content Is Retrieved over Plaintext HTTP
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is a disclosed paper-discovery and extraction workflow with some security-quality caveats, but no evidence of hidden, deceptive, destructive, or purpose-incompatible behavior.
Install only if you are comfortable with the skill querying public paper APIs, writing batch results under ~/.discovery/data/batch, and using your GitHub account for the optional PR workflow. Review extracted JSON before submission, and prefer fixing the HTTP arXiv endpoint and removing the unverified TLS fallback before relying on results for public data contribution.
scripts/extract.py:62Scientific Paper Content Is Retrieved over Plaintext HTTP
scripts/extract.py:28TLS Certificate Verification Can Silently Fall Back to an Unverified Context
Referenced artifact was not completely inspected
d the paper text and produce the structured JSON yourself. The bundled prompt (`references/prompt.txt`) is your extraction specification.
Referenced artifact was not completely inspected
d the paper text and produce the structured JSON yourself. The bundled prompt (`references/prompt.txt`) is your extraction specification.
Referenced artifact was not completely inspected
d the paper text and produce the structured JSON yourself. The bundled prompt (`references/prompt.txt`) is your extraction specification.
Referenced artifact was not completely inspected
d the paper text and produce the structured JSON yourself. The bundled prompt (`references/prompt.txt`) is your extraction specification.
Referenced artifact was not completely inspected
d the paper text and produce the structured JSON yourself. The bundled prompt (`references/prompt.txt`) is your extraction specification.
The skill explicitly instructs use of local file reads/writes and the GitHub CLI, but it does not declare any permissions or allowed-tools scope. That creates an overbroad execution surface where an agent may invoke filesystem and network-capable tooling without clear user consent boundaries, increasing the risk of unintended data access or repository actions.
The trigger phrase "Extract some papers" is broad natural language that can overlap with ordinary conversation and may activate the skill unexpectedly. Because the skill can discover papers, read local bundled files, write results, and optionally lead into GitHub submission, accidental invocation can cause unintended actions or state changes.
The trigger phrase "Submit my extractions" is ambiguous because it does not clearly signal that networked GitHub operations, repository cloning, branch creation, pushing, and PR creation may follow. Users could reasonably interpret it as a local save action, so the ambiguity increases the chance of unintended external data transfer.
The skill instructs writing extracted paper analyses to /tmp/result.json or any local path, creating persistent local artifacts outside the immediate chat/session. Session persistence can expose sensitive or proprietary research inputs to other local users/processes, allow accidental reuse across tasks, or leave stale files that are later submitted or processed unintentionally.
For each paper from Step 1, produce a JSON object following the schema in
`references/prompt.txt`. The paper's abstract is your input text.
Write the JSON to a temporary file (e.g., `/tmp/result.json` or any local path).
**Key requirements:**
- Output ONLY valid JSON (no markdown wrapping, no commentary)
The manifest frames the skill as extracting provides/requires relationships from scientific publications to surface cross-domain connections. In contrast, the module docstring and command set show broader behavior: it discovers papers from multiple remote sources, deduplicates against remote/local tracking data, validates outputs, and saves files locally, which is materially broader than extraction alone.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
"""Discover recent papers from OpenAlex."""
papers = []
url = (
f"https://api.openalex.org/works"
f"?filter=is_oa:true,type:article,has_abstract:true"
f"&sort=publication_date:desc&per_page={count}"
f"&mailto=discovery-engine@proton.me"
The docstring states that extraction is done by the agent and that this script only handles non-LLM tasks like discovery, validation, and saving. However, normalize_result actively transforms extraction payload semantics by moving fields, synthesizing structured objects from strings, inferring constraint classes, and generating operation names, which goes beyond passive handling and changes the extraction content itself.
No suspicious patterns detected.