Back to skill

Security audit

Self Verify

Security checks for vulnerabilities and agentic risk

Overview

This self-check skill is mostly coherent, but it automatically directs the agent to search persistent memory without clear consent, scoping, or redaction limits.

Review before installing. This skill may be useful for reducing confident mistakes, but it should only be used where automatic searches and persistent-memory checks are acceptable. Prefer adding explicit consent and redaction rules before allowing it to access long-term memory.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
SKILL.md:29
Finding

Unscoped Access to Persistent Agent Memory During Claim Verification

Content
View full analysis
" ``` ``` ### Technical Analysis The skill instructs the agent to consult persistent memory through `MEMORY.md` or the `memory_search` tool when verifying a claim. It does not limit searches to records relevant to the current user, session, or task, and it provides no consent, authorization, redaction, or data-minimization requirements. Persistent agent memory can contain personal information, prior conversation content, credentials, operational details, or information unrelated to the current request. Because the prescribed output format includes an `Evidence` field, retrieved memory content could be reproduced in an answer. The risk depends on the host environment granting this skill access to persistent memory; the file itself does not independently bypass access controls. ### Attack Path 1. The skill activates before the agent returns a consequential or high-confidence answer. 2. A claim is selected for verification. 3. Following lines 58–61, the agent submits that claim to `memory_search`; line 29 additionally directs it to check established memory. 4. An overly broad or adversarially constructed claim causes the search to return unrelated persistent records. 5. Sensitive content from those records is treated as supporting or contradictory evidence. 6. The agent may disclose that content in the skill's required `Evidence` output field. ### Impact Assessment If the host exposes sensitive persistent memory to the skill, the agent may read information beyond the minimum scope required for the current task. Potential impact includes disclosure of prior-session content, personal information ...[truncated 326 chars]
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (2)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill’s activation conditions are very broad and include common conversational phrases like 'verify', 'check this', 'are you sure', and any 'high-confidence claim'. This can cause the skill to trigger in many normal interactions, increasing the chance of unintended tool use, workflow hijacking, or unnecessary pre-delivery behavior across unrelated tasks.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 18)May include surrounding context.

md
### 1. Source Check
- [ ] Did I cite a source for this?
- [ ] Is the source credible and relevant?
- [ ] Or am I inferring from first principles without checking?

### 2. Uncertainty Check
- [ ] Have I labeled uncertainty correctly? (OBS/DER/INT/SPEC)

Static analysis

No suspicious patterns detected.