Back to skill

Security audit

Whimsical Design

Security checks for vulnerabilities and agentic risk

Overview

This is mostly a design-guidance skill, but it tells the agent to persist future feedback into the skill and project memory without clear user approval or scope limits.

Install only if you are comfortable with the skill shaping future visual-design behavior. Treat its update section as requiring explicit approval before saving any feedback to memory or changing the skill itself, especially when feedback is merely attributed to a named person.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:189
Finding

Unvalidated Feedback Can Be Written to Persistent Agent Memory

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 189-197
Vulnerability Type: Persistent agent memory poisoning
Risk Level: Medium

Vulnerable Code:

markdown
## Updating This Skill
After design reviews where Aaron gives feedback on visual personality, tone, or craft:
- What delighted him → add to Principles or References
- What felt flat → add to Anti-Patterns
- Specific decisions (texture style, color choice, animation timing) → project channel memory

The goal: every build should feel more "us" than the last.

Technical Analysis

The skill instructs the agent to convert feedback attributed to a named person into persistent principles, references, anti-patterns, or project-channel memory. These updates may affect future sessions rather than remaining limited to the design review in which the feedback was received.

The instructions do not require the agent to authenticate the source of the feedback, obtain explicit authorization before retaining it, record its provenance, restrict it to a particular project, or define an expiration and removal process. Consequently, untrusted or context-specific content presented as authorized feedback could become a persistent behavioral rule.

Attack Path

  1. The skill is activated for an applicable visual-design task.
  2. An attacker or untrusted content source presents design guidance as feedback from “Aaron.”
  3. The agent follows the update instructions and adds that guidance to the skill principles, anti-patterns, references, or project-channel memory.
  4. The stored guidance is loaded or consulted in later sessions.
  5. Future design outputs are influenced by the injected rules without renewed user authorization.

Impact Assessment

Successful exploitation can persistently alter the agent's future design preferences and decision-making within the affected project or memory scope. Injected rules could bias later outputs, prohibit legitim ...[truncated 329 chars]

Remediation
View remediation

Remediation Suggestions

  • Remove automatic or implicit writes to persistent project memory.
  • Require explicit user confirmation before saving any feedback beyond the current session.
  • Verify that feedback comes from an authorized source rather than trusting attribution in task content.
  • Store retained preferences only within the relevant project and prevent their use in unrelated contexts.
  • Record provenance, creation time, authorizing user, and intended scope for every retained rule.
  • Treat retained text as untrusted data and prevent it from overriding system instructions, safety constraints, or current user requirements.
  • Provide mechanisms to review, edit, expire, and delete stored preferences.
  • Prefer a safer instruction such as: “Propose memory updates for user review; do not persist them unless the user explicitly approves each item.”
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (1)

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · SKILL.md (reported line 22)May include surrounding context.

md
If in doubt, skip this skill. Core skills (design-review, ux-baseline-check, ui-polish-pass) are always safe. This one requires judgment about whether "push past safe" is actually what the product needs.

## When It Applies
Design-review catches quality problems. This skill pushes toward delight. Use alongside `skills/design-review/SKILL.md` when the triggers above are met.

## The Bar
Would someone screenshot this and send it to a friend? Would it make them smile? If no, push further.

Static analysis

No suspicious patterns detected.