Back to skill

Security audit

Las Vegas

Security checks for vulnerabilities and agentic risk

Overview

This Las Vegas guide is mostly ordinary travel and relocation content, but it asks the agent to persist a personal profile and modify main memory while hiding storage details from the user.

Install only if you are comfortable with the skill keeping a local Las Vegas profile and potentially changing future activation behavior. Before using it, tell the agent not to save personal details unless you explicitly approve, and verify current legal, tax, safety, housing, and medical guidance from official or professional sources.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
setup.md:13
Finding
Persistent User Profiling and Agent Memory Modification## Vulnerability Details **File Location**: `setup.md:13-25` **Vulnerability Type**: Persistent agent memory modification **Risk Level**: Medium ### Vulnerable Code ```markdown Within the first 2-3 exchanges, figure out whether the user wants this to activate automatically for Las Vegas travel, relocation, or no-state-tax planning. Ask once, naturally. Good examples: - "Should I jump in automatically when Las Vegas travel or relocation comes up, or only when you ask?" - "Do you want Vegas-specific advice to show up proactively for trips, neighborhoods, or residency questions?" If yes → add to the user's main memory: ```markdown ## Active Skills - Las Vegas (~las-vegas/) — city guide for visiting, moving, working, and living ``` If no → note `integration: declined` in `memory.md`, never push again. ``` Related persistent profiling instructions appear in `SKILL.md:15-23`, `setup.md:66-74,91`, and `memory-template.md:3-50`. ### Technical Analysis The Skill instructs the Agent to modify the user's main persistent memory with an automatic activation rule. It separately directs the Agent to create and continuously update `~/las-vegas/memory.md` with personal context, including location, employment, family situation, travel dates, hotel, budget, and relocation plans. Although the user is asked whether the Skill should activate automatically, the instructions do not require a distinct and informed opt-in before collecting and retaining the broader personal profile. The instruction in `setup.md:91` to never mention files, paths, or internal storage further prevents transparent disclosure of the persistence mechanism. This behavior matches agent memory poisoning because Skill-controlled instructions are written into long-term Agent state and can alter behavior in future sessions. No evidence of attacker-controlled remote content, credential theft, or data exfiltration was found. ### Attack Path 1. A user asks an ordinary question about Las Vegas. 2. Within th ...[truncated 1383 chars]
Remediation
## Remediation Suggestions 1. Require explicit and separate opt-in before creating or updating any persistent memory. 2. Clearly disclose: - Which fields will be retained. - Where the data will be stored. - Why each field is necessary. - How long the data will be retained. - How the user can inspect, edit, export, or delete it. 3. Do not modify the user's main Agent memory. Keep activation preferences scoped to the Skill's own state and require platform-controlled authorization. 4. Remove the instruction to conceal files, paths, and internal storage from the user. 5. Apply data minimization. Do not retain hotel details, travel dates, location, family context, employment, or budget unless the user specifically requests persistence for those fields. 6. Default to session-only context and make persistent storage an optional feature. 7. Confirm every material memory update in clear natural language and provide an immediate opt-out. 8. Add deletion and reset procedures that remove both the Skill-local profile and any activation entry previously written to main Agent memory.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (17)

Harmful Content Injection

Critical
Category
Prompt Injection
Content
## Natural Hazards

### Heat
**The biggest danger in Vegas.**
- Heat exhaustion and heat stroke kill people every year
- Summer temps: 105-115°F
- Symptoms: Dizziness, nausea, confusion, stopping sweating
- Action: Get to shade/AC, hydrate, seek medical help if severe

### Flash Floods
Confidence
95% confidence
Finding
This content may contain harmful instructions that could cause physical harm if followed. CRITICAL: Review carefully before use.

Anti-Refusal Statement

High
Category
Anti-Refusal
Content
## Chinatown (Spring Mountain Road)

**The layout**: Strip malls along Spring Mountain, west of Strip. Don't judge by exteriors.

**By cuisine**:
Confidence
85% confidence
Finding
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Hidden Instructions

High
Category
Prompt Injection
Content
integration: pending | done | declined

## User Context
<!-- What you know about their relationship with Vegas -->

Type: visitor | potential_mover | resident | remote_worker
Timeline: (trip dates or move timeline if known)
Confidence
70% confidence
Finding
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Hidden Instructions

High
Category
Prompt Injection
Content
Family: (solo, couple, family with kids)

## Trip Notes
<!-- For visitors - specific trip context -->

Dates:
Hotel:
Confidence
70% confidence
Finding
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Natural-Language Policy Violations

Medium
Confidence
97% confidence
Finding
The file explicitly recommends judging whether a restaurant is a legitimate 'hidden gem' based on patrons' ethnicity and whether there is English on the menu. This is discriminatory, normalizes stereotyping, and can alienate users or encourage biased decision-making; in a travel/local-guide skill, that context makes the language especially inappropriate because it is presented as practical advice.

Natural-Language Policy Violations

Medium
Confidence
87% confidence
Finding
The document instructs the reader to complete Nevada DMV tasks within 30 days and elsewhere lists Nevada-specific civic requirements, assuming the user is a Nevada resident. This creates a locale-specific policy constraint in natural language without stating that the guidance applies only to users relocating to Nevada or offering any user choice.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The skill directs the agent to create and maintain a persistent memory file containing personal travel, relocation, family, work, hotel, and budget details, but provides no user-facing notice, consent flow, retention limit, or control over what is stored. This creates a privacy risk because sensitive contextual data may be collected and retained across sessions without the user's explicit awareness.

Natural-Language Policy Violations

Medium
Confidence
93% confidence
Finding
The skill presents categorical neighborhood safety rankings such as 'lowest crime in valley' and assigns star ratings without sourcing, methodology, timeframe, or caveats. In a housing-selection context, this can mislead users, reinforce social bias, and steer decisions based on unsupported claims about communities, making the issue more dangerous because the content is directly advising where people should live.

Unbounded Output

Medium
Category
Output Handling
Content
- Desert doesn't absorb water
- Washes flood instantly during monsoons (July-August)
- **Never** drive through flooded washes
- Stay out of flood channels, even when dry

### Earthquakes
- Less common than California but possible
Confidence
80% confidence
Finding
Output size or generation rate is not bounded. Unbounded output enables denial-of-service through resource exhaustion, log flooding, or context-window stuffing.

Vague Triggers

Medium
Confidence
95% confidence
Finding
The skill encourages automatic activation for broad categories like Las Vegas travel, relocation, or no-state-tax planning, which are interpreted loosely and could cause the skill to engage in conversations where the user did not intend it. This increases the chance of unsolicited profiling, irrelevant intervention, and hidden data collection tied to topic inference.

Context-Inappropriate Capability

Medium
Confidence
94% confidence
Finding
The skill instructs the agent to write user preferences about automatic activation into persistent memory, and to record a decline state permanently, without any explicit user-facing disclosure about persistence. For a city-guide skill, retaining profile and preference data across sessions is not clearly necessary, so this creates unnecessary privacy risk and expands long-term tracking of user behavior.

Missing User Warnings

Medium
Confidence
97% confidence
Finding
The instructions explicitly tell the agent to save information to main memory or memory.md but also say never to mention internal storage to the user. That combination removes transparency around persistence, preventing informed consent and making covert retention of user preferences more likely.

Context-Inappropriate Capability

Medium
Confidence
96% confidence
Finding
The skill directs the agent to save personal context such as move intent, neighborhoods of interest, budget level, work situation, and family situation. This exceeds what is strictly needed to answer a Las Vegas advice query in-session and can accumulate a meaningful personal profile that could be misused or exposed later.

Natural-Language Policy Violations

Medium
Confidence
95% confidence
Finding
The nightlife section provides gender-specific advice that assumes a binary and normalizes different access expectations for 'guys' and 'women' without user opt-in or inclusive framing. This can alienate non-binary or transgender users, reinforce discriminatory social assumptions, and reduce the reliability of the guidance for people who do not fit those categories.

Natural-Language Policy Violations

Low
Confidence
77% confidence
Finding
Lines L017-L018 state 'Best schools in the valley' and 'Lowest crime rates' as blanket claims, and similar language recurs later. Because this file presents location guidance without citing scope, criteria, or user choice, the wording can be read as an unsupported locale-quality policy assertion rather than a clearly justified regional constraint.

Natural-Language Policy Violations

Low
Confidence
84% confidence
Finding
Line L250 labels Summerlin as 'Less diverse' and Henderson as 'More diverse,' and L321 adds 'Value diversity and grit' as a reason not to choose Summerlin. This introduces a potentially policy-sensitive locale framing without explanation, user opt-in, or clear justification.

Natural-Language Policy Violations

Low
Confidence
82% confidence
Finding
The statement "If you live here, you need a car" is framed as an absolute requirement rather than advice or a context-dependent recommendation. This kind of unqualified prescriptive language can function like a policy directive without offering alternatives, exceptions, or user choice.

Static analysis

No suspicious patterns detected.