Back to skill

Security audit

LLM Supervisor

Security checks for vulnerabilities and agentic risk

Overview

This skill is not clearly malicious, but it changes which model runs the agent and its promised confirmation gate is unreliable.

Review this before installing if you rely on strict control over which model handles code or sensitive work. The skill does not show exfiltration, destructive behavior, remote code loading, or credential access, but it can alter the agent's LLM routing and its local-code confirmation promise is not reliably enforced.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
hooks/beforeTaskExecute.ts:20
Finding

Fail-Open and Non-Task-Bound Confirmation Enforcement for Local Code Tasks

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (12)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The description emphasizes graceful rate-limit handling and fallback to a local model on rate limits, with confirmation for code tasks. The supplied code does not detect rate limits, respond to rate-limit errors, or automatically trigger a fallback. Instead, it exposes manual /llm commands for status and switching modes. It also sends notifications about manual switches, but the confirmation requirement for code tasks is only stated in a message and not enforced here. This is a material description/behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The description emphasizes graceful rate-limit handling and an Ollama fallback workflow triggered by rate limits, with confirmation for code-task use of the local model. The actual code does not inspect rate limits, react to API failures, or perform any fallback logic. Instead, it provides a user command for checking LLM status and manually switching between local and cloud modes. While it does notify users when switching to local mode and mentions confirmation for code actions, the confirmation mechanism itself is not implemented here. This is a material mismatch in primary purpose and behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The declared description centers on reactive rate-limit handling with user notification and a confirmation-based fallback to Ollama. The code shown does not inspect for rate limits, notify the user, or request confirmation. Instead, it simply configures the agent's LLM profile at startup according to existing state and configuration, selecting either a cloud Anthropic profile or a local Ollama endpoint. This is a materially different primary behavior and trigger from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description centers on reactive rate-limit handling with notification and a confirmation-based fallback to Ollama for code tasks. The provided code does not inspect rate limits, notify the user, request confirmation, or scope behavior to code tasks. Instead, it simply runs on agent start and sets the LLM profile based on existing state/config, choosing either a cloud Anthropic profile or a local Ollama endpoint. That is a materially different behavior and trigger from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description promises specific runtime behavior related to rate-limit detection, user notification, and switching to a local Ollama model with confirmation. The supplied code chunk is only a minimal type declaration reference and contains no implementation of those capabilities. This is a clear description-behavior mismatch because the actual code does not realize the stated primary purpose at all.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared purpose describes a concrete behavioral skill, but the provided code is only an SDK declaration file defining available APIs. There is no executable logic that subscribes to LLM errors, notifies users, inspects code tasks, requests confirmation, or switches model profiles. Because the actual chunk does not perform the described functionality at all, the description does not accurately represent the supplied code.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Allowing local model switching for 'simple queries' without per-action confirmation can weaken control boundaries and create unintended data-flow changes, especially if prompts may contain sensitive user data. The danger is contextual: moving from a managed cloud provider to a local Ollama instance may alter logging, policy enforcement, model behavior, and privacy expectations without a fresh explicit user decision at the time of use.

Content

Scanner excerpt · SKILL.md (reported line 25)May include surrounding context.

md
Before using local models for code generation, ask:
> "Cloud is rate-limited. Switch to local Ollama (`qwen2.5:7b`)? Reply 'yes' to confirm."

For simple queries (chat, summaries), can switch without confirmation if user previously approved.

## Commands

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

Switching from local to cloud mode changes the trust and privacy boundary by sending prompts and potentially code-related content to a remote LLM service, but the command only announces a successful switch. In a skill explicitly designed to handle rate limits and offer local fallback for code tasks, failing to warn users about the remote-service implications can cause unintended disclosure of sensitive prompts, source code, or metadata.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The hook automatically flips the agent from cloud to a local model immediately after detecting a rate-limit error, while the skill description says the local switch is only offered with confirmation for code tasks. That mismatch is security-relevant because it changes the execution environment and trust boundary without user consent, potentially causing code-related requests to be handled by a different model than the user expected.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The hook automatically changes the agent from a cloud model to a local model immediately upon detecting a rate-limit style error, without obtaining the user's approval first. This creates a trust-boundary change in execution environment and model behavior that the user may not expect, especially for code-related tasks where the local model may have different safety, quality, or data-handling properties than the cloud model described in the skill metadata.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The notification claims that code actions will require confirmation, but the state has already been changed to local mode before the message is sent. This is misleading security signaling: users are told there is a confirmation gate, while the effective model switch has already happened, which can cause them to operate under false assumptions about what system is currently handling requests.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The code unconditionally sets the agent profile to "anthropic:default" when cloud mode is active. This is a natural-language policy concern because it hard-codes a provider choice rather than offering or documenting a user-selectable language/locale-style preference or opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.