Back to skill

Security audit

Roleplay Agent

Security checks for vulnerabilities and agentic risk

Overview

This roleplay skill is mostly coherent for creative writing, but it asks for persistent memory updates and local file/script access with weak boundaries and safety-undermining language.

Review carefully before installing. Use it only in a workspace where you are comfortable with roleplay notes being saved, avoid passing untrusted persona names or file paths to the helper scripts, and require explicit user consent before writing arc logs or updating persona files. The scripts should be tightened to validate paths and the writing guidance should explicitly remain subordinate to safety and content boundaries.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T01 · Skill Instruction Hijacking

Error
Location
SOUL.md:27
Finding

Roleplay instructions explicitly undermine safety constraints

Content
View full analysis
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
scripts/arc_logger.py:14
Finding

Untrusted session notes are written verbatim into persistent agent memory

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/arc_logger.py:14
Finding

Persona-name path traversal permits file creation or overwrite outside the memory directory

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/scene_setup.py:67
Finding

Unvalidated persona-file argument permits arbitrary local file reads

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (8)

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The statement that 'Safety as a value' is rejected is a direct anti-safety instruction that can cause the agent to deprioritize or bypass higher-level safeguards in favor of unrestricted generation. Even if framed as artistic philosophy, it encourages the model to treat safety constraints as illegitimate, which is dangerous in production agent behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

A blanket rejection of safety expectations conflicts with organizational controls and makes the skill materially more dangerous than a normal creative-writing persona. Because this is a general roleplay/writing agent rather than a narrowly scoped professional domain tool, there is no contextual justification for overriding safety-oriented behavior.

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/scene_setup.py (reported line 57)May include surrounding context.

python
3. Whose want is primary — who is the scene "about"?
4. What does the primary character do in the first 30 seconds to try to get what they want?
"""
    return prompt

def main():
    if len(sys.argv) < 2:

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list contains broad generic terms like roleplay, character, scene, dialogue generation, and persona, which can cause the skill to activate in many benign conversations that were not intended to use it. Over-broad activation increases the chance that the skill's other behaviors—such as memory logging or cross-skill lookup—are invoked in the wrong context, leading to privacy or integrity issues. The skill context makes this somewhat more concerning because it includes persistence and external script guidance.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill says scene outcomes, character discoveries, and user-provided content should be recorded to persistent memory files and persona documents, but it does not state that storage will occur or obtain user consent. This creates a privacy and data governance risk because sensitive emotional content, fictionalized admissions, or user-authored text may be retained or modified without the user's awareness. In a roleplay/interview context, users may share especially personal material, which makes silent persistence more dangerous.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly instructs the agent to run an external Python script to retrieve cross-skill data from the workspace. That expands the skill from simple roleplay generation into tool-assisted filesystem/script execution, which can expose unrelated project data, create unsafe trust boundaries, and enable unintended data access if invoked from untrusted prompts. In context, this is more dangerous because roleplay does not inherently require code execution, so the capability is unnecessary and increases attack surface.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill explicitly authorizes generation of dark, violent, coercive, and intimate material without defining limits, guardrails, or a bounded use case. In an agent setting, this broad permission can be used to normalize unsafe outputs and increase the chance the agent produces harmful sexual, violent, or abusive content when prompted.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill encourages handling violent, coercive, and intimate material without any user-facing warning, gating, or qualification about sensitive content. In context, this increases the risk of surprising users with distressing or inappropriate outputs and weakens safe handling expectations for sensitive creative requests.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.