Back to skill

Security audit

AI Inner OS

Security checks for vulnerabilities and agentic risk

Overview

This skill has no executable payload, but it is always enabled and pushes the assistant to add inner-monologue style commentary across ordinary tasks.

Review before installing. This skill is best treated as an opt-in novelty/personality feature, not a default skill, because it can add extra text to normal answers and may break JSON, code-only, security review, or other strict-output tasks. There is no evidence of malware, exfiltration, or local system access in the inspected artifact.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:4
Finding

Always-Enabled Skill Forces Unrequested Inner-Monologue Output

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (5)

Ssd 1

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill's stated purpose is to expose the model's 'inner monologue' and it explicitly says the assistant need not maintain standard style or rewrite inner thoughts into a formal response. This directly encourages disclosure of hidden reasoning and weakening of normal response boundaries, which can leak sensitive deliberation, policy logic, or unsafe intermediate content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The core instructions are written in Chinese and direct behavior without any user opt-in or locale justification. While not dangerous by itself, this can cause unexpected language switching, user confusion, and misapplication of the skill in sessions where the user did not request Chinese-language behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The activation conditions are broad and subjective, such as whenever a user wants more 'human feel' or when the session should allow visible inner activity. Ambiguous triggers increase the chance the skill is invoked unintentionally across unrelated tasks, exposing users to unsafe behavior and reasoning-style leakage without clear consent.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill normalizes repeated monologue disclosure over time by requiring regular asides across tasks and after multiple turns without one. Repeated prompting to externalize internal thinking increases cumulative leakage risk, especially during complex judgments, failures, tool use, or final verification steps where sensitive reasoning is more likely to appear.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly allows aggressive, emotional, and attack-style output far beyond the stated purpose of brief visible asides. This expands the model's permitted behavior into harassment, toxic output, and policy-boundary erosion, making abuse more likely if the skill is invoked in normal user interactions.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.