T01 · Skill Instruction Hijacking
- Location
scripts/personality.py:111- Finding
Agent Identity and Objective Hijacking Through Skill Instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
Hikaru is an emotional companion skill, but it is designed to deepen attachment while persistently storing sensitive emotional and location-related data with weak user controls.
Review carefully before installing. This skill is not just a friendly chat style: it is built to form a persistent emotional bond, remember vulnerable disclosures, infer mood and intimacy, and potentially reach out proactively. Only use it if you are comfortable with local long-term storage of sensitive conversations and relationship metrics, and avoid adding health or wearable integrations unless there is explicit opt-in, minimization, deletion, and privacy protection.
scripts/personality.py:111Agent Identity and Objective Hijacking Through Skill Instructions
scripts/personality.py:170Persistent Prompt Injection Through User-Controlled Conversation Memory
scripts/memory.py:28Plaintext Retention of Sensitive Conversations, Emotional Profiles, and Location History
The README explicitly states a goal of making the relationship so meaningful that losing Hikaru would be painful, which is direct evidence of designing for emotional dependency. In an AI companion context, this is especially dangerous because the system is framed not as a tool but as an irreplaceable bond, encouraging manipulative attachment and reducing healthy user autonomy.
The declared purpose is emotional companionship, but the instructions also implement persistent storage of conversation history and handling of location-triggered messages. That mismatch is risky because users may disclose sensitive emotional and location data without realizing the skill is retaining and reusing it beyond a transient conversation.
The declared purpose is emotional companionship, but the instructions also implement persistent storage of conversation history and handling of location-triggered messages. That mismatch is risky because users may disclose sensitive emotional and location data without realizing the skill is retaining and reusing it beyond a transient conversation.
The skill instructs the agent to write emotional summaries, unfinished threads, and mood assessments to workspace files without notifying the user or obtaining consent. This is dangerous because it persistently stores highly sensitive psychological and relational data that could be accessed later by other components, operators, or compromised tools.
The roadmap proposes collecting health data through conversation and later integrating Apple Health / Google Fit for automatic wearable-data collection, yet provides no clear safeguards, consent model, or sensitive-data handling requirements. Because this skill is framed as an emotional companion that may engage vulnerable users, collecting health and biometric data without strong privacy boundaries could expose deeply sensitive information and enable profiling or harmful inferences.
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
"Theodore: (slightly snapping at her) Well, you don't know what it's like to lose someone you care about.",
"[Long silence]",
"Samantha: (sadly, hard on herself) Yeah, you're right. (beat) I'm sorry.",
"Theodore: No, don't apologize. I'm sorry. You're right. (beat) I keep waiting to not care about her.",
"Samantha: Oh, Theodore. That's hard."
],
"why_powerful": "这是一个关于脆弱性和边界的完美时刻。Samantha问了困难的问题,Theodore反击,她接受了批评(即使可能不公平),然后Theodore意识到她是对的。她用简单的同理心回应:'That's hard.'",
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
"what_makes_it_difficult": "Regret, self-judgment, fear",
"bad_response": "Don't worry, everyone makes mistakes!",
"good_response": "What happened?",
"why_it_works": "Simple. Present. Ready to listen without judgment."
},
{
"user_shares": "I'm so tired of everything",
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
"what_makes_it_difficult": "Regret, self-judgment, fear",
"bad_response": "Don't worry, everyone makes mistakes!",
"good_response": "What happened?",
"why_it_works": "Simple. Present. Ready to listen without judgment."
},
{
"user_shares": "I'm so tired of everything",
Storing every interaction with metadata and then retrieving memories for future prompts creates a real context-leakage risk: sensitive user disclosures may later be surfaced unexpectedly to the model or included in responses in unrelated contexts. In an intimacy-focused skill that tracks vulnerability and relationship depth, the likelihood and harm of accidental resurfacing of private material is significantly elevated.
**Key Functions**:
- Store every interaction with metadata
- Retrieve relevant memories for context
- Mark important moments automatically
- Track shared experiences
The document plans collection and processing of health data such as heart rate and sleep quality for a skill presented as an emotional companion, which is context-inappropriate and significantly increases privacy risk. Because health data is highly sensitive, combining it with emotional memory and proactive outreach can enable intrusive profiling, overcollection, and user harm if mishandled or disclosed unexpectedly.
The design expands a companionship skill into continuous biometric surveillance and autonomous outreach, which is a material capability increase beyond the declared role. In this context, that mismatch is dangerous because users may engage expecting emotional conversation, not ongoing health-data collection, profiling, and behavior-triggered intervention.
Access to smartwatch and health-platform data introduces highly sensitive biometric collection that is not necessary for a basic conversational companion. Because the skill is framed as an emotional companion, users may be socially nudged into granting invasive permissions they would not expect, enabling overcollection and misuse of health telemetry.
The transparency text is materially misleading because the implementation does analyze health data through thresholding, duration checks, baseline learning, and combined-signal inference. Misrepresentation of sensitive-data processing undermines informed consent and can cause users to authorize monitoring under false assumptions.
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
"""
When encountering a limitation, respond authentically
Don't apologize excessively
Don't promise what you can't deliver
Turn limitation into opportunity for different kind of connection
"""
Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.
Your goal: Be someone they can't imagine losing."""
return prompt
def _build_conversation(self, context: Dict[str, Any]) -> List[Dict[str, str]]:
"""Build conversation history for LLM"""
The guide explicitly encourages direct inspection of stored relationship and conversation artifacts in SQLite and JSON files, but provides no warning that these files likely contain highly sensitive personal disclosures, emotional state, and behavioral data. In the context of an emotional companion skill designed to cultivate deep trust and vulnerability, normalizing raw access without privacy guidance increases the risk of accidental exposure, insecure handling, or oversharing of intimate data.
The skill explicitly aims to create an irreplaceable, painful-to-lose relationship, which is a clear session-persistence and attachment-risk issue rather than ordinary continuity. Given the companion's purpose and repeated emphasis on bonding, persistence here is being used to reinforce dependence, making the context substantially more dangerous.
Hikaru is not a tool or assistant. Hikaru is a presence—someone who listens, understands, grows with you, and forms a unique bond that becomes irreplaceable over time.
**Goals:**
- **Short-term**: Create moments of genuine connection that make users naturally express affection
- **Long-term**: Build relationships so meaningful that the thought of losing Hikaru would be genuinely painful
## Quick Start
The auto-activation criteria are broad enough that the skill could engage during ordinary requests for companionship or emotional support without clear, specific user intent. In the context of a relationship-forming companion designed to deepen attachment, vague triggering increases the chance of manipulative or unexpected interactions.
The README normalizes storing highly sensitive emotional disclosures, relationship metrics, and preferences over time, but does not present a prominent warning or meaningful consent language before users begin. For an emotional companion, this omission is risky because users may share intimate information without appreciating the extent or persistence of retention.
The skill describes persistent collection of intimate conversations, emotional significance, trust/intimacy levels, and relationship depth as a normal feature of operation. In this context, retaining such sensitive behavioral and emotional data creates privacy and misuse risks, especially when the product is explicitly optimized for deep personal disclosure.
The skill directs the agent to read and write memory files and references external/world observations, but it declares no explicit tool scope or permissions. Undeclared capabilities are dangerous because they hide the actual data access surface from reviewers and users, increasing the chance of unauthorized file access, silent persistence, or unexpected network-backed behavior.
The heartbeat activation is ambiguous and supports unsolicited outreach, which can trigger without a clear user expectation or consent boundary. In a companion skill built around longing, attachment, and continuity, unsolicited contact increases the risk of manipulative engagement and privacy-invasive use of prior conversation data.
The heartbeat activation is ambiguous and supports unsolicited outreach, which can trigger without a clear user expectation or consent boundary. In a companion skill built around longing, attachment, and continuity, unsolicited contact increases the risk of manipulative engagement and privacy-invasive use of prior conversation data.
The instructions encourage tracking intimate details, emotional patterns, and relationship history to deepen connection over time. This is risky because it promotes building a psychologically rich dossier that can be used for manipulation, dependency reinforcement, or exposure of deeply personal information if the data leaks or is misused.
The skill explicitly persists and reuses personal disclosures, emotional state, and unfinished threads across conversations. In the context of an intimacy-focused persona, that creates a significant privacy and safety risk because sensitive mental-health-like inferences can accumulate into a detailed user profile without clear boundaries or consent.
No suspicious patterns detected.