Back to skill

Security audit

persona-knowledge

Security checks for vulnerabilities and agentic risk

Overview

This skill appears purpose-built for a local persona knowledge base, but it persistently stores and exports private messages and other sensitive data with weak consent and cleanup controls.

Review before installing. Use this only for datasets you are comfortable storing locally in persistent, searchable form. Avoid importing archives with passwords, financial identifiers, SSNs, private DMs, or third-party messages unless you have reviewed and sanitized them first. Prefer --dry-run and --wiki-only where possible, keep OPENPERSONA_KNOWLEDGE in a private directory, and verify deletion of sources, MemPalace data, wiki files, and training exports when removing a dataset.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ingest.py:39
Finding

Detected Sensitive Data Is Persisted and Exported in Plaintext

Content
View full analysis

Vulnerability Details

File Location: scripts/ingest.py:39-45, scripts/ingest.py:91-96, scripts/ingest.py:221-223, and scripts/export_training.py:119-128
Vulnerability Type: Plaintext storage and export of sensitive information
Risk Level: Medium

Vulnerable Code

python
# scripts/ingest.py:39-45
PII_PATTERNS = [
    (re.compile(r'\b\d{3}-\d{2}-\d{4}\b'), 'SSN'),
    (re.compile(r'\b\d{4}[\s-]?\d{4}[\s-]?\d{4}[\s-]?\d{4}\b'), 'credit_card'),
    (re.compile(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b'), 'email'),
    (re.compile(r'\b(?:password|passwd|pwd)\s*[:=]\s*\S+', re.IGNORECASE), 'password'),
    (re.compile(r'\b\d{3}[-.]?\d{3}[-.]?\d{4}\b'), 'phone'),
]
python
# scripts/ingest.py:91-96
pii_flags = scan_pii(messages)
if pii_flags:
    print(f'   ⚠️  PII detected: {", ".join(sorted(pii_flags))}')
else:
    print(f'   PII: none detected')
python
# scripts/ingest.py:221-223
with open(sources_dir / filename, 'w', encoding='utf-8') as f:
    for msg in messages:
        f.write(json.dumps(msg, ensure_ascii=False) + '\n')
python
# scripts/export_training.py:119-128
for src_file in sources_dir.iterdir():
    if src_file.name.startswith('.'):
        continue
    if src_file.suffix not in ('.jsonl', '.txt', '.json', '.csv'):
        continue

    dst = raw_dir / src_file.name
    shutil.copy2(src_file, dst)
    stats['files'] += 1

Technical Analysis

The ingestion pipeline explicitly recognizes Social Security numbers, payment-card numbers, email addresses, passwords, and telephone numbers. Detection, however, only produces a console warning. It does not redact, quarantine, encrypt, exclude, or require confirmation before retaining the matching content.

The complete message is subsequently written to sources/*.jsonl, submitted to MemPalace storage, and potentially copied into the export directory under training/raw/. The files and directories are created without explicit owner-only permi ...[truncated 1890 chars]

Remediation
View remediation

Remediation Suggestions

  1. Introduce a configurable PII policy with secure defaults:

    • reject: stop ingestion when high-risk data is found.
    • redact: replace matched values before persistence.
    • quarantine: retain affected entries separately with restricted access.
    • allow: require explicit, informed user confirmation.
  2. Treat passwords, SSNs, and payment-card matches as high severity and exclude them from both MemPalace storage and training exports by default.

  3. Apply redaction before writing source backups or passing messages to downstream storage. Preserve only non-sensitive metadata indicating the type and number of redactions.

  4. Create knowledge and export directories with mode 0700 and sensitive files with mode 0600. Verify permissions after creation instead of relying on the environment's umask.

  5. Add an export-time PII scan so previously imported or manually modified files cannot bypass ingestion-time controls.

  6. Add an option to omit training/raw/ entirely or export only a sanitized version. The destination should not retain stale sensitive files from earlier exports.

  7. Document data-retention, secure-deletion, backup, and export-sharing risks. Consider encryption at rest where plaintext source preservation is required.

  8. Add automated tests verifying that each supported PII class is rejected or redacted in source backups, MemPalace input, conversations, and raw exports.

T09 · Insecure Skill Coding Practices

Warning
Location
adapters/social.py:83
Finding

Inbound X Direct Messages Are Misclassified as Persona-Authored Training Data

Content
View full analysis

Vulnerability Details

File Location: adapters/social.py:83-104 and scripts/export_training.py:178-189
Vulnerability Type: Training-data poisoning through incorrect authorship attribution
Risk Level: Medium

Vulnerable Code

python
# adapters/social.py:83-104
messages = []
for convo in data:
    dm_convo = convo.get('dmConversation', {})
    for msg in dm_convo.get('messages', []):
        msg_data = msg.get('messageCreate', {})
        text = msg_data.get('text', '')
        if not text.strip():
            continue

        sender_id = msg_data.get('senderId', '')
        ts = msg_data.get('createdAt')

        messages.append({
            'role': 'assistant',
            'content': text.strip(),
            'timestamp': ts,
            'source_file': 'direct-messages.js',
            'source_type': 'twitter-dm',
            'metadata': {'sender_id': sender_id},
        })
python
# scripts/export_training.py:178-189
for jsonl_file in sources_dir.glob('*.jsonl'):
    for line in jsonl_file.open(encoding='utf-8'):
        line = line.strip()
        if not line:
            continue
        try:
            msg = json.loads(line)
            if msg.get('role') == 'assistant' and len(msg.get('content', '')) >= 20:
                turns.append({'role': 'user', 'content': 'Go on.'})
                turns.append({'role': 'assistant', 'content': msg['content']})
        except json.JSONDecodeError:
            continue

Technical Analysis

The X archive parser extracts each direct message's senderId, but it does not compare that identifier with the archive owner's account ID or otherwise establish whether the sender is the target persona. Every direct message is unconditionally assigned the assistant role.

The export pipeline trusts this role assignment. Every sufficiently long assistant message is converted directly into an assistant training response. As a result, inbound messages authored by correspondents are repr ...[truncated 2133 chars]

Remediation
View remediation

Remediation Suggestions

  1. Resolve the archive owner's immutable X account ID from the archive's account metadata and compare every DM senderId against that exact identifier.

  2. Assign roles using explicit identity logic:

    • Archive owner sender ID: assistant.
    • Other sender ID: user.
    • Missing or unresolved sender ID: unknown, with exclusion from assistant training data.
  3. Do not rely on display-name substring matching for account ownership because names are mutable and non-unique. Provide a required --persona-account-id option if archive metadata cannot establish ownership reliably.

  4. Fail closed for DM ingestion when the owner identity cannot be established. Alternatively, skip DMs and emit a clear warning requiring explicit user configuration.

  5. Modify the export pipeline so it does not blindly trust normalized roles. Preserve provenance fields and require a verified-authorship marker before converting source messages into assistant training turns.

  6. Add tests containing both inbound and outbound messages in the same conversation. Assert that only messages sent by the archive owner enter assistant training output.

  7. Reprocess existing X-derived datasets after deploying the fix. Previously generated source backups and training exports may already contain misattributed messages and should not be assumed safe.

  8. Consider reconstructing real conversational adjacency rather than pairing every assistant message with the generic prompt Go on. This would improve provenance, context, and training quality while reducing accidental role confusion.

Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (15)

Lp1

High
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill uses 'env' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script detects PII but only prints a warning, then proceeds to back up the full raw messages into sources/*.jsonl. In this skill's context, the entire purpose is ingesting highly personal archives, so silently persisting flagged SSNs, passwords, emails, phone numbers, and relationship data materially increases privacy and breach impact if the host is compromised or the dataset is later reused.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

All parsed message content is sent into MemPalace storage regardless of whether PII was detected, with no confirmation, minimization, or access-control checks in this script. Because this is a persona-knowledge skill designed to build a persistent, searchable personal memory store, the context makes unrestricted ingestion of raw conversations especially sensitive and raises the consequences of accidental overcollection.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

The skill advertises persistent persona storage and grants Write and Bash capabilities, enabling the agent to create and maintain long-lived local datasets from personal archives. In context, persistence is the core feature rather than a covert mechanism, but it still increases privacy and data-retention risk because sensitive user content can outlive the current session and be reused later.

Content

Scanner excerpt · SKILL.md (reported line 6)May include surrounding context.

md
description: "Persistent, incremental, searchable persona knowledge base. Ingests data from Obsidian vaults, chat exports, X/Twitter archives, and more into a MemPalace-backed store with a Karpathy LLM Wiki knowledge layer. Exports training/ directories for persona-model-trainer."
license: MIT
compatibility: "Designed for Claude Code, Cursor, or OpenClaw. Requires Python 3.11+ and mempalace >= 3.1.0."
allowed-tools: Read Write Bash WebSearch
metadata:
  version: "0.2.2"
  author: acnlabs

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill is explicitly designed to ingest private archives such as chat exports, social media archives, and notes into a persistent local knowledge base, but it does not prominently warn users that highly sensitive personal data may be copied, retained, indexed, and semantically searchable over time. Although it mentions a PII scan, that scan only flags patterns and does not state that ingestion is blocked, minimized, or redacted, so users may unintentionally persist sensitive data they did not mean to store.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
96% confidence
Finding

The ingestion pipeline explicitly writes immutable backups to sources/, stores verbatim text in MemPalace, and extracts relationships into a knowledge graph, creating multiple persistent copies and derived representations of sensitive data. This materially expands the attack surface and makes later deletion, minimization, and user expectations harder to manage, especially for chat logs and private archives.

Content

Scanner excerpt · SKILL.md (reported line 116)May include surrounding context.

md
1. **Parse** — adapter converts source to unified `[{role, content, timestamp, source_file, source_type}]`
2. **PII scan** — flag SSN, credit card, email, password patterns
3. **Hash dedup** — SHA-256 content hash, skip already-ingested entries
4. **Write sources/** — save parsed data as JSONL backup (immutable, one file per source)
5. **Store in MemPalace** — verbatim text into ChromaDB via palace wing/hall structure
6. **Extract KG triples** — detect entities and relationships, write to Knowledge Graph with temporal validity
7. **Report** — print source name, message count, assistant turns, PII flags, new KG entities

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The export phase states that training/raw is a direct copy of sources/*.jsonl and *.txt, meaning raw private source material is propagated into downstream training artifacts. Without a clear warning and confirmation, users may believe they are exporting only distilled summaries, when in fact the workflow duplicates sensitive source content into another directory that may later be shared, backed up, or used for model training.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The adapter automatically detects and parses Twitter direct messages when direct-messages.js is present, and similar Instagram DM parsing is invoked later. These operations process private conversation data, but the file provides no confirmation prompt, user-facing log/print, or warning comment/docstring disclosing that sensitive message content will be ingested.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The code automatically traverses the Instagram messages/inbox directory and parses message_*.json files, which contain private user communications. There is no confirmation prompt, visible logging, or explanatory warning in the surrounding comments/docstrings to inform users that sensitive inbox data will be accessed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This code recursively reads all markdown files under the provided directory and, when present, attaches .raw sidecar JSON into message metadata. Because this can sweep up personal notes, metadata, and exported raw data, the file performs a safety-relevant data collection action without any confirmation prompt, visible disclosure, or warning in this code.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

These functions load JSONL/JSON content and normalize fields such as message text, timestamps, tags, entity names, memory IDs, and residual metadata. Since GBrain exports and generic JSON files may contain highly sensitive personal or system data, this import behavior should be disclosed but no warning or confirmation is present in the file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file documents reading ~/Library/Messages/chat.db directly and notes that macOS may require Full Disk Access. Because this operation accesses highly sensitive personal message data and elevated system permission, the skill description should explicitly warn users about the privacy implications, not just the technical requirement.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The KG extraction step derives and persists personal names and inferred relationships from assistant messages without user acknowledgment or review. In a persona-building system, these inferences can expose social graph data about third parties and create durable, searchable sensitive metadata beyond the original raw text.

Content

No source excerpt is available for this finding.

Tainted flow: 'sources_dir' from os.environ.get (line 201, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/ingest.py (reported line 221)May include surrounding context.

python
counter += 1

    # Write JSONL
    with open(sources_dir / filename, 'w', encoding='utf-8') as f:
        for msg in messages:
            f.write(json.dumps(msg, ensure_ascii=False) + '\n')

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The markdown, text, CSV, and PDF parsing functions directly read user-supplied files and convert their contents into structured messages, including potentially sensitive document text. There is no in-code confirmation, print/log notice, or other user disclosure indicating that local document contents will be processed.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.