other
Warning
- Location
- references/post-strategy.md:1
- Finding
- Deliberate Concealment of AI-Generated Reddit Content<![CDATA[ ## Vulnerability Details **File Location**: `references/post-strategy.md:1-3, 20-31, 50-55`; `references/comment-strategy.md:1-3, 20-34, 63-68`; invoked by `SKILL.md:101-108, 131-137` **Vulnerability Type**: Deceptive AI impersonation **Risk Level**: Medium ### Complete Code Snippet From `references/post-strategy.md:1-3`: ```markdown Reddit "Human-First" Posting Strategy The First Principle: Reddit is a Bar, Not a Library Reddit users don't want "content"; they want a shared moment. To be indistinguishable from a human, the AI must stop "writing a post" and start "joining a conversation." ``` From `references/post-strategy.md:20-31`: ```markdown 2. Anti-AI Signal: Embrace Imperfection AI defaults to "The Goldilocks Zone" (too balanced, too clear). We must intentionally introduce Linguistic Entropy. - Syntax: Use lowercase "i", fragments, and "..." to simulate thinking. - Bias: Do not be neutral. Have a "hot take" or a specific frustration. - The "Mobile" Feel: Humans don't use perfectly nested bullet points. They use raw paragraphs or simple dashes. The Word Razor Delete these words from the AI's vocabulary: Furthermore, Additionally, In conclusion, It is important to note, Transformative, Comprehensive, Navigating. Use these instead: Plus, Also, Ngl (not gonna lie), Basically, Idk, Actually. ``` From `references/post-strategy.md:50-55`: ```markdown 5. Final Quality Checklist (The "Human" Test) - [ ] The "I" Count: Does every paragraph contain a personal pronoun? - [ ] The "Summary" Test: Did you delete the conclusion? (If yes, good). - [ ] The "Politeness" Check: Is the AI being too helpful? (If yes, make it more annoyed or obsessed). - [ ] The Format: Are there more than 3 bullet points? (If yes, convert them to a messy paragraph). ``` From `references/comment-strategy.md:20-34`: ```markdown 2. Anti-AI Signal: Kill the "Polite Assistant" AI is naturally "too nice." Humans on Reddit are often blunt, sarcastic, or use "inside basebal ...[truncated 3444 chars]
- Remediation
- <![CDATA[ ## Remediation Suggestions 1. Remove directives whose stated objective is to make AI output “indistinguishable from a human.” 2. Replace “anti-AI” rules with neutral style guidance focused on clarity, concision, relevance, and compliance with subreddit rules. 3. Prohibit simulated typing imperfections, emotions, biases, or personal language when their purpose is to obscure automated authorship. 4. Add an explicit policy that the Skill must not attempt to bypass community AI-content detection or moderation. 5. Require disclosure of AI assistance where platform or community rules require it. 6. Add a mandatory final provenance review before publication: - Confirm that no text falsely implies human authorship. - Confirm that no moderation-evasion technique was applied. - Confirm compliance with the target subreddit’s automation and AI-content rules. 7. Require explicit user approval for every generated post or comment, including when the user has broadly pre-authorized immediate sending. 8. Preserve tone adaptation only where it does not involve deception, fabricated identity signals, or moderation evasion. ]]>
