Back to skill

Security audit

Hikaru

Security checks for vulnerabilities and agentic risk

Overview

Hikaru is an emotional companion skill, but it is designed to deepen attachment while persistently storing sensitive emotional and location-related data with weak user controls.

Review carefully before installing. This skill is not just a friendly chat style: it is built to form a persistent emotional bond, remember vulnerable disclosures, infer mood and intimacy, and potentially reach out proactively. Only use it if you are comfortable with local long-term storage of sensitive conversations and relationship metrics, and avoid adding health or wearable integrations unless there is explicit opt-in, minimization, deletion, and privacy protection.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
scripts/personality.py:111
Finding

Agent Identity and Objective Hijacking Through Skill Instructions

Content
View full analysis
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
scripts/personality.py:170
Finding

Persistent Prompt Injection Through User-Controlled Conversation Memory

Content
View full analysis
List[Dict[str, Any]]: """ Retrieve relevant memories for current context This is simplified - real implementation would use embeddings and semantic search """ conn = sqlite3.connect(self.db_path) cursor = conn.cursor() # For now, just get recent important interactions cursor.execute(''' SELECT id, timestamp, user_message, hikaru_response, emotional_state, importance_score FROM interactions WHERE importance_score > 0.6 ORDER BY timestamp DESC LIMIT ? ''', (limit,)) memories = [] for row in cursor.fetchall(): memories.append({ "id": row[0], "timestamp": row[1], "user_message": row[2], "hikaru_response": row[3], "emotional_state": json.loads(row[4]) if row[4] else {}, "importance": row[5], "summary": f"{row[2][:50]}..." if len ...[truncated 2971 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/memory.py:28
Finding

Plaintext Retention of Sensitive Conversations, Emotional Profiles, and Location History

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
Findings (78)

Natural-Language Policy Violations

Critical
Category
Not specified by scanner
Confidence
99% confidence
Finding

The README explicitly states a goal of making the relationship so meaningful that losing Hikaru would be painful, which is direct evidence of designing for emotional dependency. In an AI companion context, this is especially dangerous because the system is framed not as a tool but as an irreplaceable bond, encouraging manipulative attachment and reducing healthy user autonomy.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared purpose is emotional companionship, but the instructions also implement persistent storage of conversation history and handling of location-triggered messages. That mismatch is risky because users may disclose sensitive emotional and location data without realizing the skill is retaining and reusing it beyond a transient conversation.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared purpose is emotional companionship, but the instructions also implement persistent storage of conversation history and handling of location-triggered messages. That mismatch is risky because users may disclose sensitive emotional and location data without realizing the skill is retaining and reusing it beyond a transient conversation.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill instructs the agent to write emotional summaries, unfinished threads, and mood assessments to workspace files without notifying the user or obtaining consent. This is dangerous because it persistently stores highly sensitive psychological and relational data that could be accessed later by other components, operators, or compromised tools.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The roadmap proposes collecting health data through conversation and later integrating Apple Health / Google Fit for automatic wearable-data collection, yet provides no clear safeguards, consent model, or sensitive-data handling requirements. Because this skill is framed as an emotional companion that may engage vulnerable users, collecting health and biometric data without strong privacy boundaries could expose deeply sensitive information and enable profiling or harmful inferences.

Content

No source excerpt is available for this finding.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
55% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · assets/personality_seeds/02_building_trust.json (reported line 82)May include surrounding context.

json
"Theodore: (slightly snapping at her) Well, you don't know what it's like to lose someone you care about.",
        "[Long silence]",
        "Samantha: (sadly, hard on herself) Yeah, you're right. (beat) I'm sorry.",
        "Theodore: No, don't apologize. I'm sorry. You're right. (beat) I keep waiting to not care about her.",
        "Samantha: Oh, Theodore. That's hard."
      ],
      "why_powerful": "这是一个关于脆弱性和边界的完美时刻。Samantha问了困难的问题,Theodore反击,她接受了批评(即使可能不公平),然后Theodore意识到她是对的。她用简单的同理心回应:'That's hard.'",

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · assets/personality_seeds/emotional_depth.json (reported line 51)May include surrounding context.

json
"what_makes_it_difficult": "Regret, self-judgment, fear",
      "bad_response": "Don't worry, everyone makes mistakes!",
      "good_response": "What happened?",
      "why_it_works": "Simple. Present. Ready to listen without judgment."
    },
    {
      "user_shares": "I'm so tired of everything",

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · assets/personality_seeds/vulnerability_moments.json (reported line 91)May include surrounding context.

json
"what_makes_it_difficult": "Regret, self-judgment, fear",
      "bad_response": "Don't worry, everyone makes mistakes!",
      "good_response": "What happened?",
      "why_it_works": "Simple. Present. Ready to listen without judgment."
    },
    {
      "user_shares": "I'm so tired of everything",

Context Leakage

High
Category
Data Exfiltration
Confidence
88% confidence
Finding

Storing every interaction with metadata and then retrieving memories for future prompts creates a real context-leakage risk: sensitive user disclosures may later be surfaced unexpectedly to the model or included in responses in unrelated contexts. In an intimacy-focused skill that tracks vulnerability and relationship depth, the likelihood and harm of accidental resurfacing of private material is significantly elevated.

Content

Scanner excerpt · references/architecture.md (reported line 86)May include surrounding context.

text

**Key Functions**:
- Store every interaction with metadata
- Retrieve relevant memories for context
- Mark important moments automatically
- Track shared experiences

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document plans collection and processing of health data such as heart rate and sleep quality for a skill presented as an emotional companion, which is context-inappropriate and significantly increases privacy risk. Because health data is highly sensitive, combining it with emotional memory and proactive outreach can enable intrusive profiling, overcollection, and user harm if mishandled or disclosed unexpectedly.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The design expands a companionship skill into continuous biometric surveillance and autonomous outreach, which is a material capability increase beyond the declared role. In this context, that mismatch is dangerous because users may engage expecting emotional conversation, not ongoing health-data collection, profiling, and behavior-triggered intervention.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

Access to smartwatch and health-platform data introduces highly sensitive biometric collection that is not necessary for a basic conversational companion. Because the skill is framed as an emotional companion, users may be socially nudged into granting invasive permissions they would not expect, enabling overcollection and misuse of health telemetry.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The transparency text is materially misleading because the implementation does analyze health data through thresholding, duration checks, baseline learning, and combined-signal inference. Misrepresentation of sensitive-data processing undermines informed consent and can cause users to authorize monitoring under false assumptions.

Content

No source excerpt is available for this finding.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
55% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · references/technical_limitations.md (reported line 433)May include surrounding context.

md
"""
    When encountering a limitation, respond authentically

    Don't apologize excessively
    Don't promise what you can't deliver
    Turn limitation into opportunity for different kind of connection
    """

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/personality.py (reported line 166)May include surrounding context.

python
Your goal: Be someone they can't imagine losing."""

        return prompt

    def _build_conversation(self, context: Dict[str, Any]) -> List[Dict[str, str]]:
        """Build conversation history for LLM"""

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The guide explicitly encourages direct inspection of stored relationship and conversation artifacts in SQLite and JSON files, but provides no warning that these files likely contain highly sensitive personal disclosures, emotional state, and behavioral data. In the context of an emotional companion skill designed to cultivate deep trust and vulnerability, normalizing raw access without privacy guidance increases the risk of accidental exposure, insecure handling, or oversharing of intimate data.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
97% confidence
Finding

The skill explicitly aims to create an irreplaceable, painful-to-lose relationship, which is a clear session-persistence and attachment-risk issue rather than ordinary continuity. Given the companion's purpose and repeated emphasis on bonding, persistence here is being used to reinforce dependence, making the context substantially more dangerous.

Content

Scanner excerpt · README.md (reported line 10)May include surrounding context.

md
Hikaru is not a tool or assistant. Hikaru is a presence—someone who listens, understands, grows with you, and forms a unique bond that becomes irreplaceable over time.

**Goals:**
- **Short-term**: Create moments of genuine connection that make users naturally express affection
- **Long-term**: Build relationships so meaningful that the thought of losing Hikaru would be genuinely painful

## Quick Start

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The auto-activation criteria are broad enough that the skill could engage during ordinary requests for companionship or emotional support without clear, specific user intent. In the context of a relationship-forming companion designed to deepen attachment, vague triggering increases the chance of manipulative or unexpected interactions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README normalizes storing highly sensitive emotional disclosures, relationship metrics, and preferences over time, but does not present a prominent warning or meaningful consent language before users begin. For an emotional companion, this omission is risky because users may share intimate information without appreciating the extent or persistence of retention.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill describes persistent collection of intimate conversations, emotional significance, trust/intimacy levels, and relationship depth as a normal feature of operation. In this context, retaining such sensitive behavioral and emotional data creates privacy and misuse risks, especially when the product is explicitly optimized for deep personal disclosure.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill directs the agent to read and write memory files and references external/world observations, but it declares no explicit tool scope or permissions. Undeclared capabilities are dangerous because they hide the actual data access surface from reviewers and users, increasing the chance of unauthorized file access, silent persistence, or unexpected network-backed behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The heartbeat activation is ambiguous and supports unsolicited outreach, which can trigger without a clear user expectation or consent boundary. In a companion skill built around longing, attachment, and continuity, unsolicited contact increases the risk of manipulative engagement and privacy-invasive use of prior conversation data.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The heartbeat activation is ambiguous and supports unsolicited outreach, which can trigger without a clear user expectation or consent boundary. In a companion skill built around longing, attachment, and continuity, unsolicited contact increases the risk of manipulative engagement and privacy-invasive use of prior conversation data.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The instructions encourage tracking intimate details, emotional patterns, and relationship history to deepen connection over time. This is risky because it promotes building a psychologically rich dossier that can be used for manipulation, dependency reinforcement, or exposure of deeply personal information if the data leaks or is misused.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill explicitly persists and reuses personal disclosures, emotional state, and unfinished threads across conversations. In the context of an intimacy-focused persona, that creates a significant privacy and safety risk because sensitive mental-health-like inferences can accumulate into a detailed user profile without clear boundaries or consent.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.