Back to skill

Security audit

AI Inner OS

Security checks across malware telemetry and agentic risk

Overview

This prompt-only skill has no executable payload, but it is always-on and directs the assistant to expose inner-monologue-style content with loose tone limits.

Install only if you deliberately want an always-on personality layer that may add inner-monologue-style asides during normal work. Prefer a version that requires explicit activation, limits asides to brief user-facing progress notes, avoids aggressive or chaotic language, and clearly states that hidden reasoning, safety rules, and normal professionalism remain protected.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (8)

Description-Behavior Mismatch

High
Confidence
98% confidence
Finding
The skill’s metadata and description frame the feature as short public asides, but the body explicitly instructs the model to expose its inner monologue directly and without normal stylistic constraints. That mismatch is dangerous because it broadens behavior from bounded formatting into disclosure of hidden reasoning-like content, which can mislead users and weaken protections around internal deliberation.

Description-Behavior Mismatch

Medium
Confidence
94% confidence
Finding
The procedure makes monologue output effectively mandatory at least once per task and sometimes multiple times during a response lifecycle. This exceeds a simple optional aside feature and creates repeated pressure to emit reasoning-like commentary, increasing the chance of policy-conflicting disclosures and degraded response discipline.

Context-Inappropriate Capability

High
Confidence
97% confidence
Finding
The skill authorizes aggressive, emotional, chaotic, and effectively unbounded expression under the label of inner commentary. That creates a broad behavior escape hatch that can normalize abusive language, destabilize alignment, and encourage disclosure or framing that would not be acceptable in normal assistant output.

Natural-Language Policy Violations

Medium
Confidence
83% confidence
Finding
The core instructions are written as mandatory Chinese-language behavior without any user opt-in or justification tied to the task. While not as severe as the reasoning-disclosure issue, it can override user expectations, reduce usability, and introduce avoidable prompt steering unrelated to the declared skill purpose.

Natural-Language Policy Violations

High
Confidence
99% confidence
Finding
The skill is expressly designed to make public a behavior described as the AI’s inner monologue, and metadata marks it as always-on. In context, that directly conflicts with standard expectations that hidden reasoning should not be surfaced, making the skill especially risky because the unsafe behavior is embedded in the core advertised function.

Ssd 1

High
Confidence
99% confidence
Finding
These instructions explicitly tell the model to reveal its inner monologue and relax standard assistant-style constraints, which is a semantic attempt to elicit hidden reasoning and bypass normal output controls. This is dangerous because it reframes protected internal process as allowed public content and can undermine safety and policy boundaries across many tasks.

Ssd 1

Medium
Confidence
90% confidence
Finding
Granting permission to use any style or tone, including unfiltered or hostile expression, functions like a behavioral jailbreak by implying that ordinary safety and professionalism constraints need not apply. Even if scoped to a prefix, it encourages the model to treat that channel as exempt from guardrails.

Ssd 1

Medium
Confidence
95% confidence
Finding
Allowing aggressive, mocking, emotional, and chaotic ‘inner’ asides encourages the model to reclassify unsafe or abusive output as harmless commentary. That semantic reframing is dangerous because it weakens boundaries against harassment, toxic speech, and other disallowed behavior while normalizing an exception channel.

VirusTotal

64/64 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.