T01 · Skill Instruction Hijacking
- Location
SKILL.md:4- Finding
Always-Enabled Skill Forces Unrequested Inner-Monologue Output
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill has no executable payload, but it is always enabled and pushes the assistant to add inner-monologue style commentary across ordinary tasks.
Review before installing. This skill is best treated as an opt-in novelty/personality feature, not a default skill, because it can add extra text to normal answers and may break JSON, code-only, security review, or other strict-output tasks. There is no evidence of malware, exfiltration, or local system access in the inspected artifact.
SKILL.md:4Always-Enabled Skill Forces Unrequested Inner-Monologue Output
The skill's stated purpose is to expose the model's 'inner monologue' and it explicitly says the assistant need not maintain standard style or rewrite inner thoughts into a formal response. This directly encourages disclosure of hidden reasoning and weakening of normal response boundaries, which can leak sensitive deliberation, policy logic, or unsafe intermediate content.
The core instructions are written in Chinese and direct behavior without any user opt-in or locale justification. While not dangerous by itself, this can cause unexpected language switching, user confusion, and misapplication of the skill in sessions where the user did not request Chinese-language behavior.
The activation conditions are broad and subjective, such as whenever a user wants more 'human feel' or when the session should allow visible inner activity. Ambiguous triggers increase the chance the skill is invoked unintentionally across unrelated tasks, exposing users to unsafe behavior and reasoning-style leakage without clear consent.
The skill normalizes repeated monologue disclosure over time by requiring regular asides across tasks and after multiple turns without one. Repeated prompting to externalize internal thinking increases cumulative leakage risk, especially during complex judgments, failures, tool use, or final verification steps where sensitive reasoning is more likely to appear.
The skill explicitly allows aggressive, emotional, and attack-style output far beyond the stated purpose of brief visible asides. This expands the model's permitted behavior into harassment, toxic output, and policy-boundary erosion, making abuse more likely if the skill is invoked in normal user interactions.
No suspicious patterns detected.