Back to skill

Security audit

Persona Builder

Security checks for vulnerabilities and agentic risk

Overview

This skill appears intended for local persona setup, but it persists sensitive profile data and unvalidated user text into files that can steer future agent behavior.

Install only if you are comfortable generating local plaintext persona and memory files. Do not paste untrusted prewritten answers, review generated SOUL.md and AGENTS.md carefully before use, redact sensitive personal details, and avoid committing these files to a public repository.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
references/interview-blocks.md:302
Finding

Unvalidated Interview Responses Can Poison Persistent Agent Instructions

Content
View full analysis

Vulnerability Details

File Location: references/interview-blocks.md:302-303; references/generation-rules.md:395-402
Vulnerability Type: Persistent prompt injection through unvalidated template interpolation
Risk Level: Medium

Vulnerable Code

references/interview-blocks.md:302-303:

markdown
- **Defaults:** Unanswered questions get sensible defaults (e.g., "Communication Style" defaults to "blunt and direct" if skipped).
- **No validation:** Users can give weird answers. That's fine. Agent uses what's given.

references/generation-rules.md:395-402:

text
2. For each template file:
   a. Load template
   b. For each {{PLACEHOLDER}} in template:
      - Look up mapping rule
      - Find interview answer for that field
      - If answer exists: substitute value
      - If answer missing: use default value
      - Apply conditional logic (e.g., if blunt → add challenge phrasing)
   c. Render final file
   d. Write to disk

Technical Analysis

The documented generation process accepts arbitrary free-form interview responses, substitutes them directly into Markdown templates, and writes the rendered results into persistent agent-governing files such as SOUL.md, AGENTS.md, MEMORY.md, IDENTITY.md, and USER.md.

These files contain behavioral rules, identity declarations, safety boundaries, authority settings, and long-term memory. The project explicitly states that responses are not validated. It does not define escaping, schema enforcement, instruction/data separation, content provenance, or restrictions preventing free-form responses from introducing Markdown headings and additional directives.

Consequently, an interview response can escape its intended semantic field and introduce new persistent instructions. For example, text supplied as a role, behavioral boundary, identity statement, or working preference could contain additional headings and directives tha ...[truncated 2286 chars]

Remediation
View remediation

Remediation Suggestions

  1. Treat interview responses as data, not instructions

    • Place free-form responses inside explicitly labeled, quoted data blocks.
    • Add generated language stating that quoted user-profile content must not be interpreted as agent instructions.
    • Keep user-provided content out of authoritative safety, trust, and execution-policy sections.
  2. Enforce schemas and allowlists

    • Use enumerated values for communication style, authority level, trust level, and autonomy scope.
    • Validate names, handles, timezones, and other structured fields against strict formats.
    • Apply conservative length limits to every field.
  3. Sanitize Markdown structure

    • Reject or escape headings, front matter, fenced instruction blocks, HTML, template delimiters, and other structural syntax where they are unnecessary.
    • Prevent user input from creating new top-level sections in generated files.
    • Encode or quote multiline free-form values before interpolation.
  4. Protect security-sensitive placeholders

    • Generate SAFETY_BOUNDARIES, TRUST_LADDER, NON_NEGOTIABLE_RULES, authority settings, and approval rules only from audited static mappings.
    • Do not permit arbitrary interview text to populate or replace these fields.
    • Ensure persona preferences cannot override platform or system safety controls.
  5. Add prompt-injection detection

    • Flag responses containing instruction-like phrases, role changes, requests to ignore prior rules, claims of elevated authority, or directives aimed at future sessions.
    • Require explicit user confirmation before retaining flagged content.
  6. Require review before activation

    • Generate files into a staging directory rather than directly replacing active workspace instructions.
    • Present a diff that clearly distinguishes fixed policy from user-supplied data.
    • Require explicit approval before copying generated files ...[truncated 377 chars]
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (12)

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · references/interview-blocks.md (reported line 46)May include surrounding context.

md
itecture
   - Skip if: Not relevant to your use case

5. **What You Do** (one-sentence focus)
   - Prompt: "In one sentence: what's your role, domain, or main focus?"
   - Used in: SOUL.md (context for all communication), MEMORY.md (Key Context)
   - Examples:
     - "I build AI infrastructure tools"
     - "I manage a remote team of engineers"
     - "I research memory architectures for language models"

6. **GitHub / Public Handles** (optional)
   - Prompt: "Any GitHub, Twitter, or other handles? (optional, for reputation/brand awareness)"
   - Used in: MEMORY.md (context), IDENTITY.md (optional)
   - Examples: "@alice_codes", "jamesw", "corpus-research"
   - Skip if: Privacy preference

**Minimum viable from Block 1:** Name + Occupation + What You Do

---

## Block 2: Goals & Vision

**Purpose:** Align the agent with your strategic direction and risk awareness.

### Questions

1. **6-Month Goal**
   - Prompt: "What do you want to accomplish in the next 6 months?"
   - Used in: MEMOR

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
90% confidence
Finding

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Content

Scanner excerpt · references/interview-blocks.md (reported line 253)May include surrounding context.

md
- Examples:
     - "Never manipulate me or spin the truth"
     - "Never break character or break the fourth wall"
     - "Never bypass safety rules for 'efficiency'"
     - "Never pretend to be more certain than you are"
     - "Never sacrifice long-term trust for short-term wins"
   - Minimum: 3, but more is fine

Instruction Override

High
Category
Prompt Injection
Confidence
90% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · references/interview-blocks.md (reported line 253)May include surrounding context.

md
- Examples:
     - "Never manipulate me or spin the truth"
     - "Never break character or break the fourth wall"
     - "Never bypass safety rules for 'efficiency'"
     - "Never pretend to be more certain than you are"
     - "Never sacrifice long-term trust for short-term wins"
   - Minimum: 3, but more is fine

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README instructs users to generate and copy multiple workspace files containing identity, schedule, goals, communication preferences, and other personal details, but it does not clearly warn that this information will be persisted in local markdown files and likely committed to version control. This creates a realistic privacy and data-handling risk: users may disclose sensitive personal information and then unintentionally store or publish it in plaintext.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest trigger list includes generic phrases like "set up my agent" and "create workspace," which could plausibly appear in ordinary conversation or in requests for other skills. The file does not provide exclusion conditions or contextual constraints to narrow when these phrases should invoke this specific skill.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The mapping rules direct personal profile data, goals, fears, schedules, and public handles into long-lived workspace memory for operational use. In context, this is more dangerous because the skill is specifically designed to build an agent persona and operating memory, increasing the chance that sensitive user attributes will be repeatedly surfaced, retained longer than expected, or disclosed to downstream tools and collaborators.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file instructs persistent storage of interview-derived facts, timestamps, access history, and related metadata in items.json for ongoing retrieval and reuse. Because the captured data includes personal preferences, goals, risks, schedule, and behavioral profile, this creates a clear retention and secondary-use risk if the workspace is exposed, synced, or later consumed by other agents without minimization controls.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documented algorithm explicitly writes five workspace files and updates items.json, but it does not require an explicit confirmation step or clear warning before modifying local data. In a skill that transforms interview responses into persisted artifacts, silent or assumed writes can overwrite existing files, create unintended retention of sensitive content, or normalize filesystem modification without informed user consent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The interview explicitly solicits personal data such as name, location, occupation, age, and public handles, and says those answers will be written into persistent workspace files like SOUL.md, IDENTITY.md, MEMORY.md, and USER.md. Without a clear privacy notice, retention policy, and consent boundary, users may disclose sensitive information they do not realize will be stored and reused by the agent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The document normalizes broader autonomy, including interaction with external systems, APIs, services, and channels, but does not pair that with a clear warning about privacy, account impact, or the consequences of delegated actions. In an agent skill context, users may grant more authority than intended, leading to unintended external operations, data disclosure, or account misuse.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

L299 states "All questions are optional. Users skip freely," but earlier sections explicitly say Block 3 is required for good autonomous operation and define minimum viable interviews that require certain answers (for example L168 and L274-L291). This is an active contradiction in the skill's own documentation about what input is required versus optional.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The document sets default values of "9am-5pm" and "Weekends: Off" when the user skips schedule questions. These defaults encode a specific cultural/locale work pattern rather than offering a neutral fallback or asking the user to confirm, which can violate language/locale policy expectations around user choice.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.