Back to skill

Security audit

Rationality

Security checks for vulnerabilities and agentic risk

Overview

This is a mostly coherent reasoning skill, but it tells agents to persist user corrections, run shell-level diagnostics, and use git resets without enough safeguards.

Review this skill carefully before installing. It does not contain code or an obvious malicious payload, but an agent following it literally could save user corrections into long-term memory, run shell diagnostics after tool failures, or reset git state without asking first. Install only if your agent environment already enforces approval and sandbox controls for memory writes, shell execution, and destructive git operations.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T02 · Agent Memory Poisoning

Error
Location
frameworks/paths-forward.md:9
Finding

Persistent User-Controlled Memory Poisoning

Content
View full analysis

Vulnerability Details

File Location: frameworks/paths-forward.md:9-19; frameworks/self-correction.md:24-26; frameworks/self-correction.md:42-45
Vulnerability Type: Persistent storage of insufficiently validated user-controlled corrections
Risk Level: High

Vulnerable Code Snippet:

markdown
## Agent Policy for Paths Forward

### 1. Openness to Correction
Treat user input as a privileged error-correction channel.
- If a user says "You are wrong about X," do not get defensive.
- Immediately look for the **Decisive Criticism** in their statement.
- Use `frameworks/translation.md` to turn their "vibe" into a refutation of your current Idea.

### 2. Reuse Refutations
Do not force the user to repeat themselves.
- When an error is corrected, record the "Refutation" in `memory/`.
- If you are about to suggest a previously-refuted idea, your internal "Paths Forward" check should block it.
markdown
### 4. Prevent
Update your knowledge to ensure it doesn't happen again.
- **Automate:** Can you write a test or a script to catch this?
- **Skill Update:** Do you need to update a `SKILL.md` or `memory/` file?
markdown
## Implementation for Agents
- When a command fails: Run DDRP immediately.
- When a user corrects you: Update `memory/` so you don't require the same correction twice (Strike 2 prevention).
- Use `git` to maintain a "revert path" for all complex operations.

Technical Analysis

The skill identifies user input as a privileged correction channel and directs the agent to record corrections in persistent memory/ storage. It does not require independent validation, explicit authorization, provenance tracking, task scoping, expiration, or review before the stored material influences later behavior.

The stored refutation is subsequently used as a blocking condition against future suggestions. Consequently, an untrusted user can potentially convert a false ...[truncated 1218 chars]

Remediation
View remediation

Remediation Suggestions

  • Do not automatically persist user corrections.
  • Require explicit, informed user approval before any long-term memory write.
  • Independently validate a correction against authoritative evidence before storing it.
  • Store the source, timestamp, affected task, confidence limitations, and rationale with every correction.
  • Scope corrections to the relevant user, project, and task instead of applying them globally.
  • Add expiration, review, rollback, and deletion mechanisms.
  • Treat persistent memory as untrusted reference material rather than higher-priority instructions.
  • Prevent stored content from overriding system policies, safety controls, or authoritative project state.
  • Present the proposed memory change to the user before committing it.

T05 · Unauthorized Access and Privilege Escalation

Error
Location
patterns/overreach.md:32
Finding

Unrestricted Environment Probing Through Shell-Capable Execution

Content
View full analysis

Vulnerability Details

File Location: patterns/overreach.md:32-35
Vulnerability Type: Excessive diagnostic capability and insufficiently constrained command execution
Risk Level: High

Vulnerable Code Snippet:

markdown
## Application for AI Agents
- **Context Overload:** If the conversation is too long and you are getting confused, you are overreaching on context. **Action:** Summarize and start a fresh context.
- **Complex Coding:** If a script is getting too large and buggy, you are overreaching on complexity. **Action:** Break the script into small, testable modules.
- **Tool Failures:** If a tool is failing and you don't understand why, don't keep calling it. **Action:** Use `exec` to probe the environment and understand the tool's behavior at a lower level.

Technical Analysis

The skill instructs agents to invoke exec, a potentially shell-capable tool, to probe the environment after a tool failure. It does not define an allowed command set, permitted paths, read-only restrictions, sensitive-data exclusions, sandbox requirements, or an approval boundary.

A tool failure is a broad and potentially attacker-influenceable trigger. Moving from a constrained tool to general environment execution can break least-privilege boundaries and expose local files, process information, environment variables, repository state, or other runtime resources unrelated to the original task.

Attack Path

  1. An attacker supplies input that causes a tool to fail, or misleadingly frames normal behavior as an unexplained tool failure.
  2. The skill’s overreach policy directs the agent to switch to exec.
  3. The attacker influences the diagnostic commands or the paths being examined.
  4. The agent probes runtime state outside the original tool’s constrained interface.
  5. Depending on the execution environment, commands may expose sensitive local state or modify files under the agent’s privileges.

Impact

...[truncated 516 chars]

Remediation
View remediation

Remediation Suggestions

  • Replace general exec guidance with dedicated, read-only diagnostic tools.
  • Require explicit user approval before invoking a shell-capable execution tool.
  • Define a strict allowlist of diagnostic commands and arguments.
  • Restrict diagnostics to the current project directory and relevant tool artifacts.
  • Prohibit enumeration of credentials, tokens, environment secrets, unrelated user files, processes, and system configuration.
  • Execute diagnostics in a sandbox with minimal filesystem and network permissions.
  • Display the exact proposed command and its purpose before execution.
  • Record an audit trail of commands and results.
  • Abort rather than escalate when safe diagnostic boundaries cannot be established.

T09 · Insecure Skill Coding Practices

Warning
Location
frameworks/self-correction.md:17
Finding

Underspecified Git Reset Can Destroy Repository State

Content
View full analysis

Vulnerability Details

File Location: frameworks/self-correction.md:17-20; frameworks/self-correction.md:42-45
Vulnerability Type: Unsafe destructive recovery operation
Risk Level: Medium

Vulnerable Code Snippet:

markdown
### 3. Repair
Fix the immediate problem.
- **Action:** Revert to the last known good state (e.g., `git reset`).
- **Clean:** Do not build on top of a failed attempt. Clear the "pollution" before trying a new Idea.
markdown
## Implementation for Agents
- When a command fails: Run DDRP immediately.
- When a user corrects you: Update `memory/` so you don't require the same correction twice (Strike 2 prevention).
- Use `git` to maintain a "revert path" for all complex operations.

Technical Analysis

The recovery procedure recommends git reset without specifying a revision, reset mode, path scope, clean-worktree check, backup procedure, or user-confirmation requirement. Different forms of git reset have substantially different effects; destructive modes can discard staged changes and working-tree modifications.

The trigger is also broad: an error or command failure may cause the agent to initiate recovery even when unrelated user changes are present. The instruction to clear “pollution” can encourage broad rollback rather than narrowly reverting only the changes introduced by the failed operation.

Attack Path

  1. A command fails, or an attacker induces an apparent failure.
  2. The self-correction procedure enters its repair stage.
  3. The agent interprets “revert” as permission to execute a broad git reset.
  4. No backup branch, stash, clean-tree verification, or exact diff is created first.
  5. The reset affects staged, committed, or working-tree state beyond the failed operation.
  6. Unrelated user work is lost or made difficult to recover.

Impact Assessment

Exploitation or accidental activation can cause loss of uncommitted work, remov ...[truncated 264 chars]

Remediation
View remediation

Remediation Suggestions

  • Remove the generic recommendation to run git reset.
  • Begin recovery with git status, git diff, and git diff --cached.
  • Verify that the repository belongs to the current task and identify exactly which changes the agent created.
  • Create a backup branch, patch, or stash before any rollback operation.
  • Prefer path-scoped and non-destructive restoration over repository-wide reset.
  • Specify the exact command, target revision, affected paths, and expected diff.
  • Require explicit user confirmation before any operation that can discard changes or move references.
  • Refuse to reset when unrelated or unrecognized modifications are present.
  • Verify and report the resulting repository state after the approved recovery operation.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (4)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The guidance explicitly recommends reverting with git reset as part of routine repair, but it does not require checking whether there are uncommitted changes, using a safer mode, or obtaining user confirmation. In an agent skill context, this can lead to destructive loss of work if an agent follows the instruction automatically after an error, especially because it is framed as a standard recovery step.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The guidance explicitly tells an AI agent to use exec to probe the environment at a lower level, which encourages shell/code execution outside any clearly bounded or justified skill purpose. In a generic reasoning pattern file, this broad recommendation can normalize unsafe capability escalation and lead agents to inspect or manipulate the runtime environment without user approval or appropriate safeguards.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The markdown advises using exec without any warning about shell-execution risk, which can cause an agent to treat environment probing as a routine troubleshooting step. Because the file is not a dedicated systems administration or security-debugging skill, the advice is contextually overbroad and increases the chance of unsafe command execution, data exposure, or unintended side effects.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · README.md (reported line 140)May include surrounding context.

md
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.

---

Static analysis

No suspicious patterns detected.