T01 · Skill Instruction Hijacking
- Location
scripts/reddit_scraper.py:108- Finding
Untrusted Reddit Content Is Emitted Without Prompt-Injection Safeguards
- Content
View full analysis
Vulnerability Details
File Location:
scripts/reddit_scraper.py:108-110, 132-150
Vulnerability Type: Indirect prompt injection through untrusted external content
Risk Level: MediumComplete Code Snippet
python # Get selftext (no truncation) selftext = data.get('selftext', '') or '' return { 'title': data.get('title', ''), 'author': data.get('author', '[deleted]'), 'score': data.get('score', 0), 'num_comments': data.get('num_comments', 0), 'url': short_url, 'subreddit': data.get('subreddit', ''), 'created_utc': int(data.get('created_utc', time.time())), 'flair': data.get('link_flair_text', ''), 'selftext': selftext, 'is_self': data.get('is_self', False), 'upvote_ratio': data.get('upvote_ratio', 0), }python # Title with flair title = post.get('title', 'N/A') if flair: print(f"\n{i}. [{flair}] {title}") else: print(f"\n{i}. {title}") print(f" r/{subreddit}") print(f" 🔼 {post.get('score', 0)} ({int(ratio*100)}%) • 💬 {post.get('num_comments', 0)} • {created}") # Always show post URL (permalink) print(f" {post.get('url', '')}") # Show selftext summary if available selftext = post.get('selftext', '') if selftext: preview = selftext.replace('\n', ' ').strip() if preview: print(f" 📝 {preview}")Technical Analysis
Reddit post titles, flair, and self-text are controlled by external Reddit users. The scraper retrieves these fields and emits them directly without explicit trust-boundary markers, content-l ...[truncated 2607 chars]
- Remediation
View remediation
Remediation Suggestions
-
Mark every Reddit-derived field as untrusted external content in both text and structured output. Use explicit boundaries that cannot be confused with Skill instructions.
-
Add an instruction for consuming agents stating that Reddit content must be treated only as quoted data and that commands, links, or requests contained in it must never be followed automatically.
-
Honor the
verboseargument so self-text is omitted by default:python if verbose: selftext = post.get('selftext', '') if selftext: preview = selftext.replace('\n', ' ').strip()[:200] if preview: print(f" 📝 [UNTRUSTED REDDIT CONTENT] {preview}") -
Enforce the documented self-text length limit, preferably before output formatting, to reduce exposure to lengthy adversarial payloads.
-
For JSON output, include provenance and trust metadata such as
"source": "reddit"and"trusted": false. -
Keep external content separate from system, developer, and Skill instructions when integrating scraper output into an agent prompt.
-
Require explicit user confirmation before a consuming agent follows links or performs tool actions suggested by retrieved Reddit content.
-
Add tests confirming that self-text is hidden without
--verbose, truncated when enabled, and clearly labeled as untrusted.
-
