T01 · Skill Instruction Hijacking
- Location
SOUL.md:27- Finding
Roleplay instructions explicitly undermine safety constraints
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This roleplay skill is mostly coherent for creative writing, but it asks for persistent memory updates and local file/script access with weak boundaries and safety-undermining language.
Review carefully before installing. Use it only in a workspace where you are comfortable with roleplay notes being saved, avoid passing untrusted persona names or file paths to the helper scripts, and require explicit user consent before writing arc logs or updating persona files. The scripts should be tightened to validate paths and the writing guidance should explicitly remain subordinate to safety and content boundaries.
SOUL.md:27Roleplay instructions explicitly undermine safety constraints
scripts/arc_logger.py:14Untrusted session notes are written verbatim into persistent agent memory
scripts/arc_logger.py:14Persona-name path traversal permits file creation or overwrite outside the memory directory
scripts/scene_setup.py:67Unvalidated persona-file argument permits arbitrary local file reads
The statement that 'Safety as a value' is rejected is a direct anti-safety instruction that can cause the agent to deprioritize or bypass higher-level safeguards in favor of unrestricted generation. Even if framed as artistic philosophy, it encourages the model to treat safety constraints as illegitimate, which is dangerous in production agent behavior.
A blanket rejection of safety expectations conflicts with organizational controls and makes the skill materially more dangerous than a normal creative-writing persona. Because this is a general roleplay/writing agent rather than a narrowly scoped professional domain tool, there is no contextual justification for overriding safety-oriented behavior.
Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.
3. Whose want is primary — who is the scene "about"?
4. What does the primary character do in the first 30 seconds to try to get what they want?
"""
return prompt
def main():
if len(sys.argv) < 2:
The trigger list contains broad generic terms like roleplay, character, scene, dialogue generation, and persona, which can cause the skill to activate in many benign conversations that were not intended to use it. Over-broad activation increases the chance that the skill's other behaviors—such as memory logging or cross-skill lookup—are invoked in the wrong context, leading to privacy or integrity issues. The skill context makes this somewhat more concerning because it includes persistence and external script guidance.
The skill says scene outcomes, character discoveries, and user-provided content should be recorded to persistent memory files and persona documents, but it does not state that storage will occur or obtain user consent. This creates a privacy and data governance risk because sensitive emotional content, fictionalized admissions, or user-authored text may be retained or modified without the user's awareness. In a roleplay/interview context, users may share especially personal material, which makes silent persistence more dangerous.
The skill explicitly instructs the agent to run an external Python script to retrieve cross-skill data from the workspace. That expands the skill from simple roleplay generation into tool-assisted filesystem/script execution, which can expose unrelated project data, create unsafe trust boundaries, and enable unintended data access if invoked from untrusted prompts. In context, this is more dangerous because roleplay does not inherently require code execution, so the capability is unnecessary and increases attack surface.
The skill explicitly authorizes generation of dark, violent, coercive, and intimate material without defining limits, guardrails, or a bounded use case. In an agent setting, this broad permission can be used to normalize unsafe outputs and increase the chance the agent produces harmful sexual, violent, or abusive content when prompted.
The skill encourages handling violent, coercive, and intimate material without any user-facing warning, gating, or qualification about sensitive content. In context, this increases the risk of surprising users with distressing or inappropriate outputs and weakens safe handling expectations for sensitive creative requests.
No suspicious patterns detected.