Back to skill

Security audit

reddit-post-guardian

Security checks across malware telemetry and agentic risk

Overview

This skill does not install code or post to Reddit, but it teaches deceptive Reddit marketing rewrites and tactics that can work around community moderation rules.

Review carefully before installing. It is reasonable to use this skill for summarizing subreddit rules and spotting obvious compliance issues, but avoid using the rewrite and post-plan portions to fabricate anecdotes, fake customer claims, hide AI authorship, or work around subreddit promotion rules. Verify rules yourself and seek moderator approval where promotion is restricted.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (5)

Context-Inappropriate Capability

High
Confidence
98% confidence
Finding
This guidance goes beyond compliance checking and provides operational advice for getting promotional links past subreddit anti-spam and anti-promotion controls by moving them into comments immediately after posting. In context, that meaningfully enables policy evasion and can help users distribute spam or self-promotional content while appearing compliant.

Context-Inappropriate Capability

High
Confidence
97% confidence
Finding
The rewrite instructions explicitly tell the model to add 'made up but plausible' personal details while presenting the result as authentic founder voice. That creates deceptive content designed to appear more trustworthy and human, which can mislead moderators and readers and undermines the claimed preservation of the original message.

Ssd 4

High
Confidence
99% confidence
Finding
The multi-step post plan normalizes a tactic of stripping promotional material from the main post and then reintroducing it in comments immediately after publication, coupled with tactical timing and engagement instructions. In this context, it incrementally operationalizes moderation evasion rather than merely assessing rule compliance.

Ssd 2

High
Confidence
99% confidence
Finding
The skill instructs the model to rewrite AI-generated promotional text specifically to reduce detection by moderators and AI-detection tools and make it appear human-authored. This is dangerous because it directly facilitates deceptive content laundering and evasion of trust-and-safety enforcement mechanisms.

Ssd 4

High
Confidence
98% confidence
Finding
Instructing the model to insert fabricated but plausible personal details creates false credibility signals that can deceive readers into believing the post reflects genuine lived experience or customer feedback. Because the skill targets promotional Reddit posts, this deception is especially risky in a community setting that relies heavily on authenticity norms.

VirusTotal

VirusTotal findings are pending for this skill version.

View on VirusTotal

Static analysis

No suspicious patterns detected.