Back to skill

Security audit

soul-archive

Security checks for vulnerabilities and agentic risk

Overview

This skill is a local memory/personality archive, but it automatically builds a broad plaintext personal profile and can inject it into future AI prompts without strong per-session consent.

Install only if you intentionally want a long-lived, plaintext personal dossier that local agents can reuse. Before using it, review config.json, consider setting auto_extract, auto_reflect, and auto_context_inject to false, keep the data directory out of cloud sync and public repositories, and avoid storing health, financial, intimate, or other sensitive details unless you accept that they may later be included in agent prompts.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/soul_extract.py:358
Finding

Sensitive Personal Data Is Persisted Without Enforcing the Configured Confirmation Gate

Content
View full analysis

Vulnerability Details

File Location: scripts/soul_extract.py:198-199, 358-429
Vulnerability Type: Missing authorization and consent enforcement for sensitive-data persistence
Risk Level: Medium

Vulnerable Code

The default configuration declares that health, financial, and intimate-relationship information requires confirmation:

python
"sensitive_topics_filter": True,
"require_confirmation_for": ["health", "finance", "intimate_relationships"],

However, the persistence method loads the configuration but only enforces extraction-dimension switches. It does not inspect either sensitive-data setting before writing extracted information:

python
def save_extraction(self, extraction: dict):
    changes = []
    config = self.load_config()
    dims = config.get("extract_dimensions", {})
    thr = self._dedup_threshold()

    # 1. Identity
    if dims.get("identity", True) and extraction.get("basic_info"):
        current = load_json(self.paths["basic_info"], DEFAULT_BASIC_INFO.copy())
        updated = self._merge_identity(current, extraction["basic_info"], thr)
        if updated:
            save_json(self.paths["basic_info"], current)
            changes.append(f"identity: updated {', '.join(updated)}")

    # 2. Personality
    if dims.get("personality", True) and extraction.get("personality"):
        current = load_json(self.paths["personality"], DEFAULT_PERSONALITY.copy())
        updated = self._merge_personality(current, extraction["personality"], thr)
        if updated:
            save_json(self.paths["personality"], current)
            changes.append(f"personality: updated {', '.join(updated)}")

    # 3. Language
    if dims.get("language_style", True) and extraction.get("language"):
        current = load_json(self.paths["language"], DEFAULT_LANGUAGE.copy())
        updated = self._merge_language(current, extraction["language"], thr)
        if update
...[truncated 4409 chars]
Remediation
View remediation

Remediation Suggestions

  1. Enforce sensitive-topic policy inside SoulArchive.save_extraction() so every caller is subject to the same authorization check.
  2. Classify candidate records before any write operation and map them to the configured topic names in require_confirmation_for.
  3. Require an explicit, verifiable confirmation value for each sensitive extraction batch, rather than relying on natural-language instructions or an agent assertion.
  4. Reject or quarantine sensitive records when confirmation is missing. Return structured information describing which records require approval.
  5. Ensure sensitive_topics_filter: false and modifications to require_confirmation_for have clearly defined semantics and are changed only through an explicit user action.
  6. Record consent metadata with approved entries, including the approved categories and confirmation time, without storing unnecessary conversation content.
  7. Add tests proving that health, finance, and intimate-relationship records cannot reach save_json() or append_jsonl() without confirmation, including through identity, topic, episodic, and emotional fields.
  8. Consider restrictive filesystem permissions and optional encryption for the resulting archive because the affected records are intentionally stored as plaintext.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (53)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · PRIVACY.md (reported line 23)May include surrounding context.

md
- **Stop auto-extraction**: edit `~/.agent-guild/skills_data/soul-archive/config.json`, set `auto_extract`, `auto_reflect`, and `auto_context_inject` to `false`. After this, nothing is captured unless you explicitly trigger it.
- **Disable specific axes**: in the same `config.json`, flip any axis under `extract_dimensions` to `false`.
- **Delete everything**: `rm -rf ~/.agent-guild/skills_data/soul-archive`. There is no other copy.
- **Export your data**: every file is plaintext JSON. Copy the directory anywhere.

## Sensitive topics

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · PRIVACY.md (reported line 23)May include surrounding context.

md
- **Stop auto-extraction**: edit `~/.agent-guild/skills_data/soul-archive/config.json`, set `auto_extract`, `auto_reflect`, and `auto_context_inject` to `false`. After this, nothing is captured unless you explicitly trigger it.
- **Disable specific axes**: in the same `config.json`, flip any axis under `extract_dimensions` to `false`.
- **Delete everything**: `rm -rf ~/.agent-guild/skills_data/soul-archive`. There is no other copy.
- **Export your data**: every file is plaintext JSON. Copy the directory anywhere.

## Sensitive topics

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · PRIVACY.md (reported line 23)May include surrounding context.

md
- **Stop auto-extraction**: edit `~/.agent-guild/skills_data/soul-archive/config.json`, set `auto_extract`, `auto_reflect`, and `auto_context_inject` to `false`. After this, nothing is captured unless you explicitly trigger it.
- **Disable specific axes**: in the same `config.json`, flip any axis under `extract_dimensions` to `false`.
- **Delete everything**: `rm -rf ~/.agent-guild/skills_data/soul-archive`. There is no other copy.
- **Export your data**: every file is plaintext JSON. Copy the directory anywhere.

## Sensitive topics

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The description promises an active memory system with agentic behaviors: session-start persona injection, recall, warnings, distillation, and automatic conversation-end sedimentation. In this code chunk, the implemented functionality is much narrower: it defines data schemas, merges structured extraction payloads into local JSON/JSONL files, deduplicates similar entries, computes completeness metrics, and prints status. Although the config contains flags like auto_extract, auto_context_inject, recall_on_task_start, and warn_on_failure_pattern_match, these are only stored as configuration values and are not acted upon anywhere in this file. The CLI accepts raw conversation text but does not analyze it; instead it prints sample code telling a human/developer to manually build an ExtractionBuilder payload. Therefore the actual behavior is primarily a local archive/persistence backend and status utility, not the full autonomous 'digital soul clone' behavior described.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The description presents a broader integrated memory/persona system, but this code chunk is only the write-side reflection engine. It stores reflections, critiques, patterns, and episodes locally, supports recurrence detection and status management, and prints status/review information. The module itself explicitly states that query/recall/warning/distillation active memory capabilities should use a different file (soul_agent_memory.py). There is no code for injecting persona into sessions, recalling memory into conversations, automatically sedimenting at conversation end, or responding to the declared trigger phrases. The local plaintext JSON storage claim is accurate, and pattern distillation is partially represented via pattern storage/update, but the primary declared scope is materially broader than the actual behavior of this chunk.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill is explicitly designed to accumulate identity, memories, personality traits, preferences, and episodic details across sessions. Persistent cross-session profiling of this breadth creates substantial privacy and data-leak risk, especially since the archive is plaintext and later reused in prompts or summaries that may reach other tools or models.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

Session-start persona/context injection can expose archived sensitive information to subsequent model interactions, plugins, logs, or third-party providers. Because the injected summary is intended to shape all future responses, the disclosure surface is broad and can leak personal details even when the current task does not require them.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The prompt directs the system to extract detailed identity, personality, emotional, behavioral, and aspirational data from user conversations, but it does not present any user-facing warning that this information is sensitive or may be persisted. Because the data categories include highly intimate profiling dimensions, users may disclose information in a normal chat context without realizing it is being transformed into a lasting personal dossier.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This prompt systematically instructs extraction of a broad, persistent profile covering identity, location, habits, emotional triggers, work rules, beliefs, and goals for later reuse. In the context of a 'digital soul' memory skill with local plaintext JSON storage and cross-session recall, this materially increases privacy harm, unauthorized profiling risk, and downstream misuse if the stored data is exposed, over-collected, or invoked without fresh consent.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The automatic write strategy states that extracted user information will be deduplicated, merged, and recorded when confidence exceeds a threshold, but it gives no explicit warning about storage, retention period, or review before persistence. This creates a meaningful privacy and safety risk because the system is not merely analyzing text transiently; it is operationalizing long-term memory formation on sensitive profile data by default.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill is purpose-built to build a persistent 'digital soul' from ongoing conversations, storing extensive personal information for later reuse across sessions. This creates a surveillance-like memory layer that can accumulate intimate behavioral and identity data beyond the user's immediate expectations, increasing harm from misuse, overcollection, or later unauthorized access.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The extraction schema covers many highly sensitive categories, including emotional triggers, personal routines, beliefs, weaknesses, coping mechanisms, aspirations, and episodic memories. Even if intended for personalization, this breadth of intimate profiling materially raises privacy and safety risk because it enables detailed inference, manipulation, or deanonymization if the archive is accessed or reused improperly.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The default configuration enables auto-extraction, auto-reflection, and auto-context injection, causing user conversation data to be harvested and reused in future tasks without a fresh confirmation step. In context, this makes the archive substantially more dangerous because persistence is not merely available; it is automated and designed to influence future agent behavior from accumulated personal data.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The completion message explicitly tells the user that the archive will 'silently collect' data according to auto_extract, normalizing covert profiling behavior. Because the skill is centered on persistent personality extraction and cross-session recall, this language indicates intentional background retention of sensitive behavioral and identity information without a strong consent boundary.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill uses 'env' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
95% confidence
Finding

The privacy notice explicitly states that archived personal data is persisted locally and that relevant slices are injected into future agent prompts. This creates real cross-session persistence of sensitive user information and possible disclosure to third-party model providers through the hosting agent, even if the skill itself has no telemetry.

Content

Scanner excerpt · PRIVACY.md (reported line 15)May include surrounding context.

md
## Who can see it

- **You** — full read/write access via the filesystem.
- **The AI agent you're using** (e.g. Claude Code, Cursor, your own scripts) — when you trigger Soul Chat, Soul Context, or any extraction, the relevant slices of your archive are passed to that agent's prompt. Whether the agent forwards anything to its provider's servers depends on **that agent's privacy policy**, not on Soul Archive.
- **No one else.** Soul Archive itself has no servers, no telemetry, no analytics, no remote sync.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README promotes automatic extraction, cross-session recall, proactive memory, and conversation-end sedimentation before clearly warning that conversation content may be autonomously persisted. In a skill centered on personality cloning and long-term memory, this increases the chance users disclose sensitive information under the assumption of ephemeral chat when the system is actually retaining and reusing it across sessions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README explicitly advertises collection of identity, personality, emotional triggers, beliefs, workflow preferences, and aspirations, which together form a highly sensitive behavioral dossier. Although it mentions local storage and optional confirmations later, the warning is not prominent where the capture scope is introduced, so users may not appreciate the surveillance, retention, impersonation, and spill risk if the local machine, logs, or downstream agents are compromised.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill explicitly promotes persistent collection of extensive personal attributes, memories, opinions, and behavioral patterns to build a reusable 'digital soul' for use across agents. Even if stored locally, this creates a concentrated plaintext profile that can be over-collected, repurposed beyond the original conversation, or exposed to other tools, making privacy harm and secondary disclosure much more likely.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The documented triggers are broad enough to enable automatic activation at conversation start/end and on generic phrases, without clear scope limits, user confirmation boundaries, or explicit conditions. In a skill that persists and reuses sensitive personal memory, ambiguous activation increases the chance of unintended collection, recall, or injection into contexts the user did not mean to affect.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Automatic session-start personality/context injection can disclose accumulated personal data to whichever agent, IDE, or platform receives the prompt, including external LLM-backed systems. The README itself notes that whether external models see the prompt depends on the platform, so auto-injection materially increases unintended propagation risk.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Broad natural-language triggers like self-reflection/everyday phrasing increase the chance that archival behavior activates when the user did not intend it. In a skill centered on persistent collection of personality, memory, and preferences, accidental activation can materially increase privacy exposure.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill markets itself as keeping data local, but it also documents building prompts from archived personal data for downstream model interactions. That creates a real disclosure path: highly sensitive locally stored profile data may be forwarded into external LLM context windows, contrary to a user's likely understanding of 'local plaintext JSON'.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The documentation says sensitive topics require confirmation, yet it also describes automatic end-of-conversation extraction when auto_extract is enabled. That inconsistency can lead to collection of sensitive health, finance, or intimate-relationship data without meaningful contemporaneous consent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Ambiguous auto-activation conditions are risky because they permit passive collection without a clear user action boundary. Given this skill stores long-lived personal-memory artifacts, unclear activation semantics can cause silent retention of sensitive content and erode informed consent.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.