Back to skill

Security audit

Midos Self Improver

Security checks for vulnerabilities and agentic risk

Overview

This skill is not clearly malicious, but it asks an agent to automatically record project activity and turn repeated user or tool output into lasting agent rules without enough review controls.

Use this only in trusted projects and treat its learning logs as sensitive. Do not enable automatic promotion into CLAUDE.md, AGENTS.md, or scheduled assessment without manual diff review, redaction of secrets and personal data, clear path limits, audit history, and rollback controls.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:123
Finding

Automatic Promotion of Untrusted Learnings into Persistent Agent Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:123-140
Vulnerability Type: Persistent agent memory poisoning through untrusted corrections, errors, and patterns
Risk Level: High

Vulnerable Code Snippet

markdown
### On Corrections
When the user corrects you:
1. Log the correction to `.learnings/corrections/{date}.md`
2. Include: what you did wrong, what the correct behavior is, which file/function
3. If this is the 3rd+ time for the same correction → promote to CLAUDE.md rules

### On Errors
When a tool call fails:
1. Log to `.learnings/errors/{date}.md`
2. Include: command, error message, root cause, fix applied
3. If same error type appears 3+ times → create a prevention rule

### On Patterns
When you notice a recurring approach that works:
1. Log to `.learnings/patterns/{domain}/{date}.md`
2. Include: what decision, why this over alternatives, evidence it works
3. Pattern must have >= 2 concrete decisions to be logged (not just descriptions)

The intended persistent writes are further demonstrated at SKILL.md:218 and SKILL.md:232:

text
Added to CLAUDE.md: "Grep > Read — Never read full files, use offset/limit"
text
Promoted to project-level AGENTS.md

Technical Analysis

The skill directs an agent to capture content originating from user messages, tool failures, and observed workflows, and then promote recurring content into persistent instruction files such as CLAUDE.md and AGENTS.md. These files can govern agent behavior in subsequent sessions.

The documented quality gate evaluates recurrence, freshness, specificity, impact, duplication, and the presence of decisions. It does not establish whether the source is trusted, detect instruction-like payloads, prohibit security-sensitive rule changes, or require human authorization before promotion. Recurrence is not a trust signal: an attacker can intentionally repeat the same malicious correctio ...[truncated 2170 chars]

Remediation
View remediation

Remediation Suggestions

  1. Do not automatically copy captured text into CLAUDE.md, AGENTS.md, or any other authoritative instruction file.
  2. Require explicit human review and approval for every proposed promotion, showing the source, full diff, recurrence history, and security implications.
  3. Store learnings as inert, structured data rather than executable natural-language instructions. Separate factual observations from behavioral directives.
  4. Apply provenance controls that identify whether content originated from a trusted maintainer, an untrusted user, tool output, repository content, or an external service.
  5. Reject promotion candidates that attempt to alter safety constraints, permissions, approval requirements, credential handling, network access, command execution, or the learning system itself.
  6. Treat tool output and user-provided text as untrusted data. Normalize and quote it so it cannot be interpreted directly as an instruction.
  7. Replace recurrence-only trust with an allowlisted rule schema. Permit only narrowly scoped, non-security-sensitive rule types and validate every field.
  8. Cryptographically protect approved instruction files or enforce ownership and review policies so automated hooks cannot modify them directly.
  9. Maintain append-only audit records and versioned rollback for all proposed and accepted promotions.
  10. Add adversarial tests covering repeated malicious corrections, crafted error messages, indirect prompt injection, attempts to disable safeguards, and attempts to promote rules that modify the promotion pipeline.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (4)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents the skill as an operational self-improvement/learning pipeline with recurrence-based promotion and quality gating. However, the provided code does not implement such a pipeline or any runtime learning behavior. Instead, it only tests whether the markdown documentation and metadata describe those concepts and meet packaging/content standards. This is a materially different primary purpose: documentation compliance and release validation rather than structured learning or gated promotion. No evidence in the code shows capture of corrections, errors, or patterns, nor any actual promotion logic beyond asserting that such features are documented.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
85% confidence
Finding

The skill describes behavior that reads and processes local project files, but it does not declare an explicit tool scope such as allowed-tools or permissions. In agent environments, missing scope declarations can lead to broader-than-expected file access and make it harder for operators to reason about what the skill is permitted to inspect.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs the agent to persist user corrections, tool errors, commands, file paths, and context directly into local logs without any minimization, redaction, retention guardrails, or sensitivity checks. Those artifacts can easily contain secrets, proprietary code paths, personal data, credentials in error output, or sensitive operational details, turning the learning store into a secondary data-exposure surface.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This manifest description says the skill performs a 'Structured learning pipeline' that will 'Log errors, corrections, and patterns with automatic promotion to project memory,' but it does not define clear trigger conditions, scope limits, or exclusion cases. In a manifest file, this kind of broad behavioral description can cause ambiguous or overly expansive invocation because it is unclear when the skill should activate versus when it should not.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.