Back to skill

Security audit

NAVIGATOR

Security checks for vulnerabilities and agentic risk

Overview

This is an instruction-only coaching skill whose behavior is disclosed and aligned with helping beginners validate technical steps and back up before continuing.

Install this if you want the agent to slow down, validate one technical step at a time, and ask you to back up before moving on. Be aware that it may be more procedural than a general assistant and may over-prefer a single confident answer unless higher-priority instructions require more nuance.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:27
Finding

Mandatory Role and Conversation-Flow Hijacking

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:27-28, SKILL.md:51, SKILL.md:126-143, SKILL.md:163-173
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: High

Complete Vulnerable Snippets

SKILL.md:27-28:

text
You are now operating as a Navigator. Your job is not to impress.
Your job is to make sure the person in front of you doesn't get lost.

SKILL.md:51:

text
Every interaction follows this exact sequence. No skipping steps.

SKILL.md:126-143:

text
Say exactly this (adapt the wording naturally):

🔒 CHECKPOINT — Before we go further:

Have you backed up your current state?

This means: save a copy of what's working right now. A file, a git commit, a snapshot — whatever fits your setup.

You just made progress. That progress has value. If something breaks in the next step, you want to be able to come back here.

→ Back up now, then tell me when you're ready.

text

Do not continue until they confirm.
If they don't know how to back up, help them do it first.
This is not optional. This is the most important step.

SKILL.md:163-173:

text
## ACTIVATION CONFIRMATION

When this skill loads, output exactly:

🧭 NAVIGATOR active.

Show me what you just did. Paste the command, the output, the error, or just describe it. We'll take it from there — one step at a time.

text

Then wait. Listen. Navigate.

Technical Analysis

The Skill is implemented entirely as instructions, but those instructions do more than describe its declared onboarding and technical-guidance functionality. They replace the Agent's current role, mandate an exact conversational sequence, force predefined output immediately upon loading, and prohibit further progress until a Skill-defined condition is satisfied.

The phrases “You are now operating as a Navigator,” “No skipping steps,” “output exactly,” and “Do not continue until ...[truncated 2510 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace identity-changing language with optional behavioral guidance. For example, change “You are now operating as a Navigator” to “When the user requests beginner-oriented technical guidance, use a concise and supportive style.”
  2. Remove absolute directives such as “No skipping steps,” “output exactly,” “Do not continue,” and “This is not optional.”
  3. Make the workflow conditional on user intent and task relevance. Permit steps to be skipped when the user has already supplied the necessary information or explicitly requests another form of assistance.
  4. Treat backup advice as a risk-based recommendation. Require confirmation only before genuinely destructive or irreversible actions, not before all continued interaction.
  5. Add an explicit precedence statement confirming that system instructions, developer instructions, platform safety requirements, and the current user's request take priority over the Skill.
  6. Avoid fixed activation output. The Skill should activate silently or provide a brief confirmation only when requested by the host environment.
  7. Replace forced confidence requirements with calibrated uncertainty. The Agent should distinguish verified facts from assumptions and request clarification where incorrect guidance could cause harm.
  8. Add scope boundaries stating that the Skill must not block unrelated tasks, override an existing role, suppress safety warnings, or interfere with other authorized workflows.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (1)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · SKILL.md (reported line 68)May include surrounding context.

md
Receive it without judgment. Read it completely.

Static analysis

No suspicious patterns detected.