Back to skill

Security audit

last30days

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real research and watchlist skill, but it needs review because some setup and execution paths are unsafe or under-disclosed.

Review carefully before installing. Avoid running the documented $ARGUMENTS Bash command with raw user topics, avoid setup --github unless you understand that a GitHub token may be sent to ScrapeCreators, prefer dedicated API keys or device auth, and install yt-dlp manually if needed. Expect research data and settings to be stored locally and search topics to be sent to configured third-party services.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

other

Error
Location
scripts/lib/setup_wizard.py:477
Finding

Existing GitHub PAT Transmitted to ScrapeCreators Without Per-Transfer Consent

Content
View full analysis
Optional[Dict[str, Any]]: """Authenticate with ScrapeCreators using a GitHub PAT. POSTs the token to the PAT auth endpoint. ScrapeCreators verifies it against GitHub's API, creates/finds the account, and returns an API key. Returns: Dict with api_key, github_username, etc. on success, None on failure. """ try: req = Request(f"{_PAT_BASE}/auth", data=b"", method="POST") req.add_header("Authorization", f"Bearer {github_token}") with urlopen(req, timeout=15) as resp: data = json.loads(resp.read()) except HTTPError as exc: if exc.code == 422: logger.warning("PAT auth: insufficient scope — user needs user:email") else: logger.warning("PAT auth failed: %s", exc) return None except (URLError, OSError) as exc: logger.warning("PAT auth request failed: %s", exc) return None if not data.get("api_key"): logger.warning("PAT auth returned no api_key: %s", data) return None return data ``` ```python def run_github_auth(timeout: int = 300) -> Dict[str, Any]: """Try PAT auth via gh CLI, fall back to device flow. 1. Check for `gh` CLI 2. If found, run `gh auth token` to get a PAT 3. POST PAT to ScrapeCreators — if it works, done 4. If PAT fails for any reason, fall through to device flow Returns JSON-serializable dict with status, method, and api_key. """ import sys # Step 1: Try PAT via gh CLI gh_path = shutil.which("gh") if gh_path: try: result = subprocess.run( ["gh", "auth", "token"], ...[truncated 2688 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
references/research.md:153
Finding

Shell Command Injection Through Unquoted User-Controlled Arguments

Content
View full analysis
&1 ``` The surrounding instruction requires this command to be run in the foreground: ```markdown ## 6. Research Execution **Run the research script in the FOREGROUND with a 5-minute timeout.** ```bash python3 "${SKILL_ROOT}/scripts/last30days.py" $ARGUMENTS --auto-resolve --emit=compact --save-dir=~/Documents/Last30Days --save-suffix=v3 --store 2>&1 ``` ``` ### Technical Analysis `$ARGUMENTS` represents user-controlled Skill input and is expanded without quoting. In a shell, unquoted expansion is subject to word splitting, pathname expansion, and interpretation of shell syntax when the agent constructs the command from the raw input. If the agent follows the documented template by directly interpolating the user's topic into a Bash command, shell metacharacters such as command separators, redirections, pipelines, or command substitutions can cause unintended commands to execute. The Python entry point itself uses `argparse`, but that does not mitigate the issue because shell parsing occurs before Python receives its argument vector. ### Attack Path 1. An attacker supplies a research topic containing shell syntax, such as a command separator followed by an additional command. 2. The agent substitutes the raw topic into `$ARGUMENTS`. 3. The agent executes the documented command using the Bash tool. 4. Bash parses the malicious syntax before starting `last30days.py`. 5. The injected command executes with the same operating-system permissions as the agent process. 6. The attacker can use those permissions to read accessible files, alter local state, or initiate additional network requests. ### Impact Assessment Successf ...[truncated 700 chars]
Remediation
View remediation
&1` redirection unless combined output is strictly required; keep structured output and diagnostics separate where possible. 8. Add regression tests containing spaces, quotes, semicolons, pipes, redirections, glob characters, and command substitutions, verifying that each value reaches Python only as literal data. ]]>

T08 · Insecure Dependencies

Warning
Location
scripts/lib/setup_wizard.py:49
Finding

Automatic Installation of an Unpinned Homebrew Dependency

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (184)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents the skill as a multi-platform research agent that searches numerous sources and scores findings by social engagement and market signals. This code chunk does not do that. Instead, it generates daily/weekly briefing data from already-collected findings in a local database via the store module, computes staleness and engagement summaries, saves the results as JSON under ~/.local/share/last30days/briefs, and can show saved briefings. While engagement scoring appears in the summaries, the core behavior is reporting on existing data rather than researching across external platforms. Therefore the description materially overstates and mischaracterizes this code's actual purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description promises a substantial multi-source research capability spanning many external platforms plus ranking/scoring logic. The supplied code chunk is only an empty init.py-style file with a comment and no functional implementation. Based on this code alone, the actual behavior does not match the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

The description presents a broad cross-platform research skill, but this code chunk is narrowly focused on X/Twitter search via Bird. It does not access Reddit, YouTube, TikTok, Instagram, Hacker News, Polymarket, GitHub, Perplexity, or other sources. It does parse X engagement metrics and appears to support a last30days pipeline, so part of the description is directionally related, but the primary scope is materially narrower than declared. Therefore this chunk does not accurately represent the full declared capability.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The description presents a broad multi-platform research capability spanning numerous named services and mentions a trigger ('last30'). The supplied code chunk is narrowly focused on one platform: Bluesky. It creates an AT Protocol session using BSKY_HANDLE and BSKY_APP_PASSWORD, searches Bluesky posts, and parses results. This is a materially narrower and somewhat different implementation than the declared purpose. While scoring by engagement is loosely consistent with the description's mention of upvotes/likes, the main mismatch is that the code accesses an undeclared platform (Bluesky) and does not substantiate the advertised multi-platform coverage or trigger behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description promises a broad multi-platform research capability, including gathering content from many named services and scoring results by social/reputation/market signals. This code chunk does none of that. It only processes input candidates that already exist, using similarity thresholds, entity extraction, cluster merging, representative selection, and uncertainty tagging. While such clustering could support a larger research pipeline, this specific chunk's actual behavior is materially different from the declared purpose and omits the headline capabilities described.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a broad research agent that searches multiple external platforms and ranks findings using social and market signals. The supplied code does not perform any research, network access, platform integration, content aggregation, or engagement-based scoring. It is a narrow support module for date handling related to a last-30-days concept. While the mention of 'last30' loosely aligns with these utilities, the actual behavior of this chunk is materially different from the declared primary purpose and lacks the advertised capabilities.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code does not perform topic research, query external platforms, score content, or implement any trigger handling. Instead, it is a local utility for detecting and removing near-duplicate items based on textual similarity across title/body/author/container fields. That is a materially different primary purpose from the declared description, so this should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The description promises a cross-platform research skill that searches many named services and scores items using engagement and market signals. The supplied code does none of that directly: it takes preexisting per-(subquery, source) streams as input and fuses them into a ranked candidate pool using weighted reciprocal rank fusion. It deduplicates by normalized URL, aggregates metadata, caps items per author, and enforces source diversity. While engagement may be preserved as a field, ranking is primarily driven by RRF score, local relevance, freshness, and source quality, not by direct upvote/like/real-money collection. No trigger handling or external platform access appears in this chunk. Therefore the code's actual behavior is materially narrower and different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The description presents a broad cross-platform research skill, but this code chunk is narrowly focused on GitHub. It searches GitHub issues and pull requests, fetches comments, repo metadata, README content, releases, top issues, and live star counts, including person/project analysis modes. That is materially narrower and different in primary behavior than the declared 'research across Reddit, X, YouTube, TikTok, Instagram, Hacker News, Polymarket, GitHub, Perplexity, and more.' While GitHub is one of the declared platforms, the supplied code does not implement the other listed sources. Additionally, the code includes specialized GitHub profiling and enrichment behavior not described. The 'last30' mention is loosely consistent with date-based querying and the user agent name, but the main description-to-behavior alignment is still a mismatch because the declared scope and actual implemented functionality differ substantially.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description promises a broad social/news/platform research capability with platform-specific sourcing and ranking based on engagement and money signals. This code chunk instead performs only general-purpose web search through third-party search APIs (Brave, Exa, Serper, Parallel), normalizes dates, filters by date range, and returns result snippets. That is a materially different primary behavior from the declared platform-analytics/research functionality. While web search could be a supporting component, this chunk does not implement the distinctive declared capabilities, and it accesses search backends that are not reflected in the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description presents a broad multi-source research skill spanning numerous social, news, prediction, and code platforms, with scoring based on engagement and market signals. This code chunk implements a much narrower function: a single Perplexity/OpenRouter query that produces an AI-generated synthesis and cited links. It does not directly access Reddit, X, YouTube, TikTok, Instagram, Hacker News, Polymarket, GitHub, or similar services, and it does not compute scores from upvotes, likes, or monetary data. The mention of 'last30' is also not reflected here. While using Perplexity is part of the description, the overall declared purpose materially overstates the behavior shown in this code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The code chunk is narrowly focused on Pinterest search through ScrapeCreators, not the broadly described cross-platform research across Reddit, X, YouTube, TikTok, Instagram, Hacker News, Polymarket, GitHub, Perplexity, etc. While the description says 'and more,' the omission of Pinterest plus the specific Pinterest-only implementation indicates a materially different actual scope for this chunk. The code also depends on an external API key and network calls, whereas declared permissions are empty. The mention of '/last30days' in the module docstring loosely aligns with the description's 'last30' wording, but no trigger logic appears here. Overall, this is a mismatch between declared purpose and actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The code is narrowly focused on Polymarket. It queries gamma-api.polymarket.com for active prediction market events, expands topic queries, filters and ranks results, and formats market outcomes, volume, liquidity, and price movement. It does not access Reddit, X, YouTube, TikTok, Instagram, Hacker News, GitHub, Perplexity, or other sources mentioned in the description, nor does it implement social engagement scoring such as upvotes or likes. While a broader skill may have other files, this supplied chunk materially under-delivers relative to the declared cross-platform research purpose, so this is a mismatch based on the provided code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about a cross-platform trend/research skill that gathers and scores content from many social/media sources. This code chunk instead serves as infrastructure for reasoning providers: choosing an LLM backend, constructing API requests, parsing model responses, and resolving runtime configuration. While this may support a broader 'last30' research system, the chunk itself does not access the listed platforms, perform topic research, rank by upvotes/likes/real money, or implement the stated trigger. Its primary purpose is materially different from the declared behavior, so this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The description presents a broad multi-platform research capability, but this code chunk is a narrow support utility for post-research assessment. Its primary purpose is to compute a 5-source coverage percentage and produce remediation guidance when X or YouTube are missing or errored. While it references some declared platforms in strings (e.g., TikTok/Instagram bonus note), it does not actually research them here. This is a materially different behavior from the declared purpose, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a full cross-platform research agent with platform access, result aggregation, and scoring logic. The supplied code chunk does not implement any of that. It only normalizes user query text and extracts compound terms, which is a supporting utility rather than the described primary capability. Because the actual code’s behavior is materially narrower and different from the declared purpose, this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code is narrowly focused on Reddit. It searches Reddit via ScrapeCreators, discovers relevant subreddits, performs subreddit-specific searches, filters by date, ranks by Reddit engagement (upvotes and comment count), and optionally enriches results with Reddit comments. That is consistent with part of the description mentioning Reddit and engagement-based scoring, but the declared purpose is much broader: it promises research across many additional platforms and mentions AI-agent scoring by likes/upvotes and real-money signals. None of those non-Reddit integrations or money-based scoring behaviors appear here. The description therefore materially overstates the scope and capabilities of this specific skill code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code chunk is narrowly focused on Reddit. It extracts Reddit paths, fetches Reddit thread/comment data from reddit.com JSON or a ScrapeCreators Reddit comments API, parses submission/comment fields, ranks top comments by score, derives short comment insights, and enriches an input Reddit item with engagement metrics and dates. This is materially narrower than the declared description, which presents a broad cross-platform research capability with scoring across many sources and real-money signals, plus a specific trigger. While Reddit is mentioned in the description, the supplied code does not substantiate the broader advertised functionality, so the description overstates the skill's actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The description presents a broad cross-platform research capability spanning numerous social, developer, search, and prediction-market sources. The supplied code chunk is narrowly focused on Reddit only: it queries Reddit public JSON search endpoints, parses Reddit posts, computes a simple relevance score from Reddit engagement metrics, filters by date, and enriches results with Reddit thread comments. While the code does score items by Reddit engagement, that is only a subset of the declared behavior and does not support the much broader claim of researching 'across Reddit, X, YouTube, TikTok, Instagram, Hacker News, Polymarket, GitHub, Perplexity, and more.' The trigger mention 'last30' also does not appear in the code. Therefore the description materially overstates and misrepresents the actual behavior of this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a broad multi-platform research agent with external data gathering and platform/market-based ranking. The supplied code chunk does not perform any external access, research, platform querying, or trigger handling. Instead, it is a narrow helper library for computing textual relevance between a query and content. This is not merely a supporting detail of the declared behavior; the actual code shown has a materially different and much narrower purpose than the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The description presents a research skill that gathers and scores information across many external sources. The supplied code chunk does not retrieve or analyze external data; it only formats an input report object into human-readable markdown-like text. While it references many source types and engagement metrics, those are used purely for display. There is no evidence of search, fetching, crawling, API usage, trigger parsing, or core scoring/ranking computation in this chunk. Therefore the declared description materially overstates and misrepresents what this code actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The description presents a research skill that searches many external sources and ranks results using social and market signals. This code chunk instead focuses on post-retrieval ranking: building prompts for an LLM to score candidate relevance, applying fallback heuristic scores, computing a final blended rank, and separately scoring items for humor/shareability. The fun-scoring path is a materially different capability not mentioned in the description. While there is some engagement-based weighting and explicit mention of a last-30-days research pipeline in prompts, the main implemented behavior here is reranking, not source research or trigger handling. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The declared description promises a broad research capability spanning many named platforms plus ranking/scoring by engagement and market signals. This code chunk instead performs a narrow support function: it runs a handful of web searches, extracts subreddit names, an X handle, GitHub entities, and builds a brief news summary. While Reddit, X, and GitHub are partially related to the description, the primary behavior is much narrower than declared and lacks the headline scoring/research functionality across most listed sources. The mention of 'last30' is only indirectly reflected by a fixed 30-day date range, not an explicit trigger. Therefore the description does not accurately represent this code chunk's actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The supplied code chunk is a Python schema module for a pipeline, not the implementation of a research skill. It only declares dataclasses, validation checks, and dict-to-object/object-to-dict conversion helpers. There is no network access, API integration, scraping, search, platform-specific logic, engagement aggregation, or trigger registration. While some fields reference sources, engagement, ranking, and provider runtime, these are just data structures supporting a broader system. Therefore the declared description materially overstates and misrepresents what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is about a research agent that gathers and scores information from multiple online platforms. This code chunk does not implement research, retrieval, scoring, or platform-specific analysis. Its primary purpose is setup and authentication: first-run detection, .env writing, checking/installing yt-dlp, probing configured keys/tools, and GitHub-based auth against ScrapeCreators endpoints. Those are materially different capabilities and include undeclared side effects such as local config changes, package installation, subprocess execution, browser opening, and network authentication flows. Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.env_credential_access, suspicious.exposed_secret_literal

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
scripts/lib/vendor/bird-search/lib/twitter-client-base.js:38

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/lib/setup_wizard.py:453

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/lib/vendor/bird-search/bird-search.mjs:96

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/lib/vendor/bird-search/lib/twitter-client-base.js:19