Back to skill

Security audit

meta-agent-eval-harness

Security checks for vulnerabilities and agentic risk

Overview

The skill is not outright malicious, but it asks the agent to persist user preferences and learned notes across sessions and even rewrite its own instructions without clear controls.

Review carefully before installing. Use this only in a contained, single-user setting, avoid recording secrets or sensitive personal data, and do not allow automatic edits to SKILL.md unless a human reviews the proposed changes first.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:69
Finding

Persistent Instruction Mutation Through Untrusted Learned State

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/learner.py:13
Finding

Unvalidated Plaintext Persistence of Arbitrary Notes

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
scripts/learner.py:19
Finding

Documented Command-Line Interface Is Not Implemented and Corrupts Learning State

Content
View full analysis
1 else True rec = (sys.argv[2] if len(sys.argv) > 2 else "") print(json.dumps(record(ok, rec), ensure_ascii=False)) ``` The documentation describes a subcommand-and-option interface equivalent to: ```bash python scripts/learner.py record --capability python scripts/learner.py record --capability --fail --error --note "" python scripts/learner.py prefer --key --val python scripts/learner.py insight python scripts/learner.py reflect ``` ### Technical Analysis The implementation does not parse any of the documented subcommands or options. It examines only `sys.argv[1]` and treats the operation as failed only when that exact argument is `fail`. It then stores only `sys.argv[2]` as the note and silently ignores every remaining argument. Under the documented failure command, `sys.argv[1]` is `record`, so the operation is incorrectly counted as successful. The skill-directory argument becomes the stored pattern, while capability, failure, error, and note options are ignored. Similarly, `prefer`, `insight`, and `reflect` are not implemented. They are treated as successful record operations rather than distinct actions. Because the Skill's mutation thresholds depend on operation and failure counters, incorrect parsing undermines the integrity of the state used to make persistent behavioral decisions. ### Attack Path 1. A user or agent follows a documented command such as the failur ...[truncated 1303 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (9)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a broad autonomous meta-skill with advanced capabilities such as self-verification, reflection, orchestration, and iterative self-improvement. The supplied code only implements a lightweight persistence utility: it loads a local JSON file, increments total operation/failure counters, appends optional notes, and writes the result back. While this could support a larger learning loop, the code chunk itself does not perform the advanced behaviors claimed. Its primary purpose is local logging/state recording, which is materially narrower than the declared purpose.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The skill advertises operational commands that write to persistent files (learned_patterns.json) but does not declare any tool scope or permission boundary. This creates an undeclared file-write capability, making it easier for operators or downstream orchestrators to invoke state-changing behavior without explicit review or sandbox constraints.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation claims the skill is 'automatically self-evolving' and needs no human maintenance, but the actual mechanism depends on operators manually running commands. This misleading automation claim can cause users to overtrust the skill, assume autonomous post-run actions will occur, or overlook the fact that persistence and reflection actions are discretionary and potentially sensitive.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill instructs storage of user preferences, error notes, and usage-derived insights in a persistent file without upfront warning in the main description or clear consent language. This creates a privacy and transparency risk because users may unknowingly have personal or contextual data retained across sessions and reused later.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly describes persistent cross-session recording of operational counts, error patterns, user preferences, and improvement suggestions in learned_patterns.json. Persistent storage of potentially user-linked data without minimization, access controls, retention limits, or consent can expose private information and create a durable profiling surface.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The example command records 输出语言 with value 中文, which indicates a language preference being set to Chinese without any accompanying statement that language is user-selectable or optional. The file overall is also written entirely in Chinese and does not mention opt-in or alternative locales, which can violate language/locale policy if the skill effectively defaults to a specific language.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The examples encourage writing user-specific preferences for future automatic reuse, normalizing cross-session persistence of individualized settings without addressing consent or sensitivity. Even seemingly harmless preferences can become identifying when combined with error notes or usage history.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The iteration rules direct automatic persistence and reuse of 'important user preferences' across future invocations. This institutionalizes cross-session memory behavior without any visible privacy boundary, creating risk of unauthorized retention, unintended personalization, or leakage of sensitive user context into later tasks.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The report is entirely written in Chinese, including headings and operational descriptions, with no indication that language selection is optional or user-configurable. The policy for this category flags language or locale constraints when they are imposed without user opt-in or documented justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.