Back to skill

Security audit

Reddit Scraper

Security checks for vulnerabilities and agentic risk

Overview

This is a read-only Reddit search helper with documentation and output-handling issues, but no hidden destructive behavior, credential use, or persistence was found.

Install only if you are comfortable with subreddit names and search queries being sent to Reddit. Treat all returned Reddit titles and post bodies as untrusted external content, and be aware that the current script may print full post text even though the documentation describes a shorter preview.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Warning
Location
scripts/reddit_scraper.py:108
Finding

Untrusted Reddit Content Is Emitted Without Prompt-Injection Safeguards

Content
View full analysis

Vulnerability Details

File Location: scripts/reddit_scraper.py:108-110, 132-150
Vulnerability Type: Indirect prompt injection through untrusted external content
Risk Level: Medium

Complete Code Snippet

python
            # Get selftext (no truncation)
            selftext = data.get('selftext', '') or ''

            return {
                'title': data.get('title', ''),
                'author': data.get('author', '[deleted]'),
                'score': data.get('score', 0),
                'num_comments': data.get('num_comments', 0),
                'url': short_url,
                'subreddit': data.get('subreddit', ''),
                'created_utc': int(data.get('created_utc', time.time())),
                'flair': data.get('link_flair_text', ''),
                'selftext': selftext,
                'is_self': data.get('is_self', False),
                'upvote_ratio': data.get('upvote_ratio', 0),
            }
python
            # Title with flair
            title = post.get('title', 'N/A')
            if flair:
                print(f"\n{i}. [{flair}] {title}")
            else:
                print(f"\n{i}. {title}")

            print(f"   r/{subreddit}")
            print(f"   🔼 {post.get('score', 0)} ({int(ratio*100)}%) • 💬 {post.get('num_comments', 0)} • {created}")

            # Always show post URL (permalink)
            print(f"   {post.get('url', '')}")

            # Show selftext summary if available
            selftext = post.get('selftext', '')
            if selftext:
                preview = selftext.replace('\n', ' ').strip()
                if preview:
                    print(f"   📝 {preview}")

Technical Analysis

Reddit post titles, flair, and self-text are controlled by external Reddit users. The scraper retrieves these fields and emits them directly without explicit trust-boundary markers, content-l ...[truncated 2607 chars]

Remediation
View remediation

Remediation Suggestions

  1. Mark every Reddit-derived field as untrusted external content in both text and structured output. Use explicit boundaries that cannot be confused with Skill instructions.

  2. Add an instruction for consuming agents stating that Reddit content must be treated only as quoted data and that commands, links, or requests contained in it must never be followed automatically.

  3. Honor the verbose argument so self-text is omitted by default:

    python
    if verbose:
        selftext = post.get('selftext', '')
        if selftext:
            preview = selftext.replace('\n', ' ').strip()[:200]
            if preview:
                print(f"   📝 [UNTRUSTED REDDIT CONTENT] {preview}")
    
  4. Enforce the documented self-text length limit, preferably before output formatting, to reduce exposure to lengthy adversarial payloads.

  5. For JSON output, include provenance and trust metadata such as "source": "reddit" and "trusted": false.

  6. Keep external content separate from system, developer, and Skill instructions when integrating scraper output into an agent prompt.

  7. Require explicit user confirmation before a consuming agent follows links or performs tool actions suggested by retrieved Reddit content.

  8. Add tests confirming that self-text is hidden without --verbose, truncated when enabled, and clearly labeled as untrusted.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (11)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The manifest says the skill uses web scraping of old.reddit.com, while the documentation says it uses the public JSON API. This behavior mismatch is dangerous because reviewers and policy systems may approve one data-access method while the implementation actually uses another, undermining trust, auditing, and network restrictions.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill advertises network-capable behavior but does not declare any explicit tool scope such as permissions or allowed-tools. In an agent environment, undocumented network access weakens policy enforcement and makes it harder to constrain or review what external communication the skill can perform.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest says the skill operates by web scraping old.reddit.com, implying HTML scraping of that site. This technical file states the current method is direct access to Reddit's public JSON API endpoints on www.reddit.com, which is a materially different behavior and data-access method than the manifest describes.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This document explicitly says the previous version used web scraping, while the current version uses the public JSON API. That directly conflicts with the manifest description asserting current behavior is scraping old.reddit.com, creating an intent/documentation mismatch about how the skill actually accesses Reddit.

Content

No source excerpt is available for this finding.

Unbounded Output

Medium
Category
Output Handling
Confidence
80% confidence
Finding

The scraper preserves full Reddit selftext and later prints it without truncation, allowing untrusted remote content to produce arbitrarily large output. In an agent context, this can cause token/window exhaustion, degraded availability, log flooding, or prompt-context pollution from hostile post content, especially when large posts are returned in bulk.

Content

Scanner excerpt · scripts/reddit_scraper.py (reported line 108)May include surrounding context.

python
if data.get('promoted', False):
                return None
            
            # Get selftext (no truncation)
            selftext = data.get('selftext', '') or ''
                
            # Build short URL using redd.it

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest says the skill reads Reddit via web scraping of old.reddit.com, but the SKILL.md documentation says it uses the public JSON API. These are different access methods, so the documentation contradicts the stated implementation approach and intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The markdown instructs users to use Spanish terms such as "nuevos" and "populares" for the --sort option, while the rest of the file is in English and there is no explanation, locale choice, or opt-in. This creates a language/locale policy concern because the skill appears to force non-English command values without documenting why.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The file docstring and CLI description explicitly describe the tool as using Reddit's public JSON API with no API key needed, which conflicts with the manifest's claim that it works via web scraping of old.reddit.com. This is an intent/documentation divergence across the skill's declared documentation surfaces.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest says this skill reads and searches Reddit posts via web scraping of old.reddit.com, but the implementation targets https://www.reddit.com JSON endpoints and parses API responses. This is a genuine description-to-behavior mismatch, even though the overall read-only Reddit browsing purpose remains the same.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The code accepts Spanish-specific values ('populares', 'nuevos') alongside English options, which embeds a partial language policy in the interface without explaining locale behavior. This can create inconsistent language expectations rather than a clearly user-selected localization mode.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

This code sends query parameters derived from user input to Reddit's public API, including search terms and subreddit values. Although network access is central to the tool's purpose, the CLI does not explicitly warn users that their inputs will be transmitted to an external service.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.