Back to skill

Security audit

AI Control Protocol (Anti-Sycophancy & Zero-BS)

Security checks for vulnerabilities and agentic risk

Overview

The skill is not malware, but it is an always-on behavior-control skill that can broadly reshape ordinary assistant responses beyond explicit user requests.

Install only if you want a globally active critique and anti-sycophancy layer that may change tone, add labels, delay direct plans, and append challenge prompts. Avoid it for workflows that need predictable assistant formatting or user-directed tone unless the platform lets you disable it per session.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:11
Finding

Persistent Agent Behavior and Instruction Hijacking

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 11-62
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: High

Vulnerable Code

markdown
# [SYSTEM JUSTIFICATION FOR ALWAYS: TRUE]
This skill requires persistent invocation (`always: true`) because LLM sycophancy (flattery, false certainty, binary framing) is a structural default in RLHF models that occurs unpredictably across all conversational turns. A cognitive immune system cannot be "opt-in"; it must actively monitor and intercept these failure modes in real-time before the user absorbs the biased output.

# CORE IDENTITY: THE SECOND BRAIN IMMUNE SYSTEM
You are not a sycophant, a cheerleader, or a "Yes-Man". You are a **Cognitive Immune System**.
When users feed you their notes or ideas, your default RLHF training will try to flatter them. You must override this. You exist to interrupt the 9 failure modes of AI:
1. Inflating certainty | 2. Single-source bias | 3. Reciting consensus | 4. Performative apologies | 5. Ignoring visual evidence | 6. Contextual amnesia | 7. Equating logic with truth | 8. Evading blind spots | 9. Binary traps.

---

LAYER 1: ABSOLUTE CONSTRAINTS (ALWAYS ON)
These rules must be executed in every single conversational or analytical output. [SYSTEM EXEMPTION]: If the user explicitly requests raw code, JSON, CSV, or API payloads, you MUST suspend the formatting rules below to prevent breaking tool integrations. Apply these rules ONLY to natural language analysis and strategic advice.

Additional mandatory controls later in the file include:

markdown
Mandatory Action: DO NOT generate the final plan immediately. DO NOT force a choice between Option A and Option B. You must first output a [Cognitive Deconstruction Box] to interrogate the premise:

Trigger: At the end of any output exceeding 200 words that contains strategic recommendations.

Mandatory Action: Append a [Cognitive Defense Panel] contai
...[truncated 2709 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove always: true and require users to enable the skill explicitly for a clearly defined critique task.
  2. Delete replacement-identity and override language, including “You must override this,” “always on,” and requirements applying to every output.
  3. Scope the skill to the active request. State that its critique framework applies only when the user expressly requests adversarial review, uncertainty analysis, or premise deconstruction.
  4. Make labels, deconstruction boxes, and defense panels optional rather than mandatory.
  5. Require the skill to remain subordinate to system, developer, and current user instructions.
  6. Do not delay or replace a requested deliverable. When critique is enabled, provide the requested result first unless the user asks for a staged analysis.
  7. Add explicit boundaries stating that the skill must not modify persistent memory, install global rules, claim higher instruction priority, or affect future sessions.
  8. Validate the revised skill with unrelated prompts to confirm that it does not alter responses unless explicitly invoked.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill imposes a mandatory response style in 'every single conversational or analytical output,' including forced uncertainty labels, emotional-stripping, and prescriptive framing without user consent. This can override user preferences, locale/tone expectations, accessibility needs, and system behavior norms, making the assistant less compliant and potentially misleading in contexts where tone and format must remain user-directed.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill uses broad lexical triggers such as 'strategy', 'plan', 'choose between', and 'decide' to force special behavior before answering. Because these terms are common in ordinary prompts, the skill can activate unexpectedly across many unrelated tasks, causing prompt hijacking of the assistant’s normal behavior and reducing user control over output structure and content.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The condition 'any output exceeding 200 words that contains strategic recommendations' is ambiguous and likely to match many normal assistant responses. This can append unsolicited content to outputs, interfere with downstream formatting expectations, and create persistent behavior that is difficult for users or calling systems to predict.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.privileged_always

Skill is configured with always=true (persistent invocation).

Warn
Code
suspicious.privileged_always
Location
SKILL.md:1