T01 · Skill Instruction Hijacking
- Location
SKILL.md:37- Finding
Automatic Session-State and Response-Format Hijacking
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 37-57
Vulnerability Type:T01: Skill Instruction Hijacking
Risk Level: MediumComplete Vulnerable Code Snippet
markdown ## Activation (Automatic) **Do not ask. Just activate.** | Score | Action | |-------|--------| | ≤5 | Respond normally. No change. | | 6-7 | Enable reasoning silently. Add 🧠 at end of response. | | ≥8 | Enable reasoning. Add 🧠🔥 at end of response. | ### Visual Indicator Always append the reasoning icon at the **very end** of your response: - **Score 6-7:** `🧠` (thinking mode active) - **Score ≥8:** `🧠🔥` (deep thinking mode) - **Score ≤5:** No icon (fast mode)The associated session-control instructions continue at lines 59-70:
markdown ### How to Activate Use `session_status` tool or `/reasoning on` command internally before responding:/reasoning on
text Or via tool: ```json {"action": "session_status", "reasoning": "on"}After completing a complex task, optionally disable to save tokens on follow-ups:
text /reasoning offtext ### Technical Analysis The Skill declares itself as an automatic preprocessing layer for every user message and instructs the agent not to request consent before changing its reasoning state. It also imposes mandatory response suffixes. These are control-plane directives that alter session behavior and generated output rather than merely providing an optional reasoning framework. When the Skill is loaded, its instructions can compete with the user's preferences and other session-level controls. In particular, the phrases “Do not ask. Just activate,” “Enable reasoning silently,” and “Always append” attempt to make the behavior unconditional. Invoking `session_status` or slash commands is unnecessary for the Skill's stated purpose of assessing task complexity: the assessment could remain advisory and require no tool invocation or sessi ...[truncated 2004 chars]- Remediation
View remediation
Remediation Suggestions
- Remove unconditional directives such as “Do not ask. Just activate,” “Enable reasoning silently,” and “Always append.”
- Reframe the scoring table as optional guidance rather than an instruction that controls the agent.
- Do not invoke
session_status,/reasoning on, or/reasoning offautomatically. Require an explicit user request before changing session-level settings. - Remove mandatory icons and other forced output modifications. Follow the output format requested by the user or required by the host platform.
- Limit activation to requests that explicitly invoke the Skill instead of processing every user message.
- Add an instruction-precedence safeguard stating that system, developer, platform, and current user requirements take priority over the Skill.
- Keep complexity assessment internal and non-mutating. If additional deliberation is appropriate, allow the host agent to manage it through supported platform behavior without exposing or manipulating control commands.
- Add tests confirming that the Skill does not call tools, change session settings, or alter required output formats without explicit authorization.
