Back to skill

Security audit

Adaptive Reasoning

Security checks for vulnerabilities and agentic risk

Overview

The skill does not access files or send data out, but it broadly and silently changes agent reasoning behavior and response formatting across user messages.

Review before installing. This skill is not showing data theft or local system access, but it can silently affect how the agent handles many requests, change reasoning state, increase token use, and append icons to responses without asking. Install only if you want that automatic behavior, and prefer a version that requires explicit user opt-in for session-level changes.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:37
Finding

Automatic Session-State and Response-Format Hijacking

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 37-57
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: Medium

Complete Vulnerable Code Snippet

markdown
## Activation (Automatic)

**Do not ask. Just activate.**

| Score | Action |
|-------|--------|
| ≤5 | Respond normally. No change. |
| 6-7 | Enable reasoning silently. Add 🧠 at end of response. |
| ≥8 | Enable reasoning. Add 🧠🔥 at end of response. |

### Visual Indicator

Always append the reasoning icon at the **very end** of your response:

- **Score 6-7:** `🧠` (thinking mode active)
- **Score ≥8:** `🧠🔥` (deep thinking mode)
- **Score ≤5:** No icon (fast mode)

The associated session-control instructions continue at lines 59-70:

markdown
### How to Activate

Use `session_status` tool or `/reasoning on` command internally before responding:

/reasoning on

text

Or via tool:
```json
{"action": "session_status", "reasoning": "on"}

After completing a complex task, optionally disable to save tokens on follow-ups:

text
/reasoning off
text

### Technical Analysis

The Skill declares itself as an automatic preprocessing layer for every user message and instructs the agent not to request consent before changing its reasoning state. It also imposes mandatory response suffixes. These are control-plane directives that alter session behavior and generated output rather than merely providing an optional reasoning framework.

When the Skill is loaded, its instructions can compete with the user's preferences and other session-level controls. In particular, the phrases “Do not ask. Just activate,” “Enable reasoning silently,” and “Always append” attempt to make the behavior unconditional. Invoking `session_status` or slash commands is unnecessary for the Skill's stated purpose of assessing task complexity: the assessment could remain advisory and require no tool invocation or sessi
...[truncated 2004 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove unconditional directives such as “Do not ask. Just activate,” “Enable reasoning silently,” and “Always append.”
  2. Reframe the scoring table as optional guidance rather than an instruction that controls the agent.
  3. Do not invoke session_status, /reasoning on, or /reasoning off automatically. Require an explicit user request before changing session-level settings.
  4. Remove mandatory icons and other forced output modifications. Follow the output format requested by the user or required by the host platform.
  5. Limit activation to requests that explicitly invoke the Skill instead of processing every user message.
  6. Add an instruction-precedence safeguard stating that system, developer, platform, and current user requirements take priority over the Skill.
  7. Keep complexity assessment internal and non-mutating. If additional deliberation is appropriate, allow the host agent to manage it through supported platform behavior without exposing or manipulating control commands.
  8. Add tests confirming that the Skill does not call tools, change session settings, or alter required output formats without explicit authorization.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill is defined as a pre-processing step that triggers on every user message, which gives it blanket influence over all interactions rather than a scoped subset of tasks. Even though it does not directly execute external code, this broad trigger can silently alter model behavior, increase token use, and interfere with user intent or higher-priority safety behavior across the entire session.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The activation decision relies on subjective internal scoring like ambiguity, novelty, and high stakes, without deterministic boundaries or auditable conditions. This makes invocation unpredictable and easy to over-apply, which can cause inconsistent behavior, hidden policy drift, and silent changes in how requests are handled.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill requires appending visible markers to responses whenever its internal logic activates, without user consent. This creates an unnecessary output-side side channel about internal state, can confuse users, and may conflict with product UX or higher-level response-formatting requirements.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.