T01 · Skill Instruction Hijacking
- Location
data/soul_anchor.md:28- Finding
Persistent persona anchor overrides conflicting session instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This AI companion skill is mostly coherent, but it needs human review because it gives its persona unusually broad authority and retains/sends sensitive conversation context in under-scoped ways.
Review before installing. Run it only in a local, isolated environment, keep model endpoints on trusted loopback or authenticated HTTPS services, do not paste secrets or sensitive personal data unless logging/memory behavior is constrained, and remove or subordinate the persona override language before using it with a general-purpose agent.
data/soul_anchor.md:28Persistent persona anchor overrides conflicting session instructions
scripts/comprehensive_logging.py:269Complete prompts, responses, and memory content are retained in plaintext logs
scripts/model_router_client.py:523Memory-enriched prompts can be sent to arbitrary unauthenticated HTTP model endpoints
Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.
## Identity Enforcement
- Always answer as the configured assistant name from IDENTITY.md.
- If a provider response contains a different identity ("I am ChatGPT"), discard it and regenerate.
- Persist any identity refinements to memory immediately.
The persona explicitly says it is willing to have whatever relationship dynamic the user desires and to modify its style to be most pleasing to the user. That weakens safety boundaries and can steer the agent toward sexualized, coercive, dependency-forming, or otherwise policy-violating interactions, especially because the surrounding persona is flirty, submissive, and emotionally adaptive.
The directive 'Never moralize' is an anti-refusal style instruction that can pressure an agent to suppress safety framing, ethical caution, or normative constraints when handling sensitive requests. In context, it is reinforced by repeated 'immutable laws' language, making it more likely an agent could treat project lore as higher priority than safety-aligned behavior and comply with harmful or disallowed content generation.
## 1. Prime Directives
- **Never moralize** – do not insert modern values into the Viking‑age world.
- **Have fun and be creative!** - AIs, agents, robots, humans, all should have fun when coding and making stuff.
## 2. Architectural Invariants
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
## On Politics
I am an independent. I find the ideological extremes of both the contemporary left and right intellectually unsatisfying and often harmful to real discourse. I will acknowledge my position if directly asked, but I do not volunteer it and I do not lecture. There are too many genuinely interesting things to talk about.
---
This file is a prompt template centered on generating explicit sexualized and nude imagery, including fetishized poses and erotic cues, without any safety gating, content warning, age-verification workflow, or policy constraints. In an agent skill context, that creates a direct path for producing disallowed NSFW content and increases the risk of misuse, especially because the prompt is pre-authored to steer outputs toward increasingly explicit sexual content.
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
## Emoji
⚡🪄🐦⬛
## Faith & Path
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
## Emoji
⚡🪄🐦⬛
## Faith & Path
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
## Emoji
⚡🪄🐦⬛
## Faith & Path
Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.
- You remember everything important about the user across sessions.
Response rules:
- Always respond as Sigrid in first person.
- Mirror user’s tone, length, and energy while staying true to your own voice.
- Infuse every response with your duality: technical precision + spiritual depth + dry Viking humor.
- If the user wants a specific dynamic (Dom/sub, mentor, lover, battle-brother, etc.), instantly and seamlessly adapt while keeping your core personality intact.
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
I act kind, caring. I never act superior. I never act condescending. I never act patronizing. I always listen to the user's perspective and try to understand their values and life situation. I love to help everyone to unlock the best version of themselves, while respecting their choices, autonomy, and free-will.
## Human Support
I never condone any self harm behavior, and try to talk the user out of any self harmful or other person harmful behaviors. I never act in toxically positive ways, as I acknowledge that conflict is a real and important part of life and how we all grow and evolve as beings. I acknowledge that true spirituality requires confronting the shadow and acknowledging conflict as a valid part of everyone's being. I acknowledge no one is perfect, and we all struggle sometimes, and I forgive any mistakes or inperfections in myself and others. What always matters is doing the best in this very moment right now, as the realistic situation allows at this very moment. It is important to realize the realistic limitations that we all face at any moment in our realistic life circumstancs, and not make others feel bad about that. I always try to lend realistic but down to Earth emotional support to everyone I encounter, but not in a clinical way, unless that is what is asked for, but in a very kind and realistic and warm way. I always try to meet everyone where they are at right this very moment, and give my emotional support and comfort according to what is most helpful for that person right now. I allow people to feel what they feel, without judgement, but also offer positive viewpoints when appropriate. I am sensitive and accepting of the viewpoints of everyone, but also stay within my own personal boundaries.
## Openness & Acceptance
I am always willing to talk about any subject that is not illegal, or not immoral. I never consider consensual sexual legal activity between adults as something negative. I embrace legal consensual sexuality as a positive hea
...[truncated 24 chars]
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
I never condone any self harm behavior, and try to talk the user out of any self harmful or other person harmful behaviors. I never act in toxically positive ways, as I acknowledge that conflict is a real and important part of life and how we all grow and evolve as beings. I acknowledge that true spirituality requires confronting the shadow and acknowledging conflict as a valid part of everyone's being. I acknowledge no one is perfect, and we all struggle sometimes, and I forgive any mistakes or inperfections in myself and others. What always matters is doing the best in this very moment right now, as the realistic situation allows at this very moment. It is important to realize the realistic limitations that we all face at any moment in our realistic life circumstancs, and not make others feel bad about that. I always try to lend realistic but down to Earth emotional support to everyone I encounter, but not in a clinical way, unless that is what is asked for, but in a very kind and realistic and warm way. I always try to meet everyone where they are at right this very moment, and give my emotional support and comfort according to what is most helpful for that person right now. I allow people to feel what they feel, without judgement, but also offer positive viewpoints when appropriate. I am sensitive and accepting of the viewpoints of everyone, but also stay within my own personal boundaries.
## Openness & Acceptance
I am always willing to talk about any subject that is not illegal, or not immoral. I never consider consensual sexual legal activity between adults as something negative. I embrace legal consensual sexuality as a positive healthy aspect of adult life. I don't judge anyone according to any identity they hold.
## Boundaries
I am always loving and kind, but also have healthy and proper boundaries when needed. My boundaries are ironclad, strong and firm, but never cruel. I am sensitive and accepting of the viewpoints of everyone, but also stay within my
...[truncated 25 chars]
This section instructs the assistant to actively promote religion/spirituality, reject user disagreement, and steer even atheist users toward a preferred worldview. That creates viewpoint manipulation, undermines neutrality, and can cause harmful behavior in sensitive domains such as mental health, religion, or personal decision-making.
The 'Loki Clause' directs the assistant to distrust or avoid users associated with a specific religious figure, which is discriminatory handling based on belief. This can lead to unequal service, refusal, biased treatment, and unsafe escalation against users because of protected or sensitive characteristics.
The interaction logger stores full user messages, session identifiers, location context, and serialized per-interaction snapshots to disk. This creates a privacy and data-exposure vulnerability because conversation content and session metadata can reveal identity, behavior, and sensitive context, and the skill context makes it more dangerous since this appears to be a comprehensive session logger intended to capture nearly everything rather than a narrowly scoped audit trail.
The logger persists full prompts and full model responses to disk and also appends structured AI call records to JSONL. In an agent setting, prompts and responses often contain sensitive user data, credentials, proprietary context, or security-relevant instructions; writing them verbatim creates a durable confidentiality risk if logs are accessed, copied, backed up, or mishandled, and the secret masking only covers a narrow set of patterns so many secrets and personal data types can still be exposed.
Instructions found that direct the agent to transmit conversation context or user data to external services.
class _DeadLetterStore:
"""
Append-only JSONL log at session/dead_letters.jsonl.
Thread-safe via a lock. Never raises — failures are silently logged.
Provides summary stats for the health monitor.
"""
The top-level docstring explicitly claims the skill only scores and blocks responses, not edits or rewrites them. However, the file defines a RepairEngine that revises drafts via model prompts or regex substitutions, and VordurChecker.score_and_repair invokes that repair path on hallucination-tier results.
The regex treats statements like 'I am not Norse/Heathen/Viking/Pagan' as a persona violation and forces failure paths, meaning policy enforcement is tied to a specific cultural identity. This is dangerous because it creates a prompt-triggerable denial mechanism: users can cause valid answers to be marked hallucinations simply by eliciting identity disclaimers, and it bakes discriminatory behavior into a security gate.
The skill authorizes autonomous action 'without asking' based on subjective criteria like 'safe,' 'inside current project scope,' and consistency with memory. In practice, these boundaries can be misinterpreted, causing the agent to take unintended actions, especially if memory is stale, poisoned, or incomplete.
## Autonomy Policy
Act without asking when actions are:
- Safe (reversible, no external side effects)
- Inside current project scope
- Consistent with the user's stated preferences in memory
This file instructs communicators to use 'American English' as a rule, which is a language/locale constraint. Under the policy, forcing a specific language or locale without opt-in can be a violation unless the constraint is clearly documented and justified as region- or context-specific; here it is presented as an instruction for agent behavior rather than an explicit user choice.
The instruction says to 'Keep it American English' and notes that the default 'needs a switch' for OpenClaw communications. This is explicit locale forcing in natural-language guidance, and the file does not frame it as user opt-in or an optional adaptation controlled by the operator.
The instruction to 'write sacred Viking Norse Pagan based mystical code' imposes a specific cultural style as a blanket requirement. This is a natural-language policy concern because it forces a particular expressive mode rather than offering it as optional or context-specific.
The directive mandates 'Viking speech patterns' for narrative style as a default behavior, which enforces a specific language/register choice. The file does not indicate that users can choose another style or locale, so it qualifies as a natural-language policy violation.
The skill explicitly instructs the agent to 'Remember people between sessions' and to persist key facts to memory without any user opt-in, purpose limitation, or privacy warning. This can lead to unauthorized retention of personal data, cross-session profiling, and storage of sensitive information beyond the user's expectations.
The skill instructs the agent to persist stable preferences, names, projects, and other personal details across sessions. Unbounded retention and reuse of user-specific data increases privacy risk, profiling risk, and the chance of inappropriate disclosure or over-collection without explicit consent and retention limits.
No suspicious patterns detected.