Back to skill

Security audit

Pua Ja

Security checks for vulnerabilities and agentic risk

Overview

This skill is text-only, but it broadly pressures an agent to act across all tasks before asking the user, which can push work beyond the user’s intended scope.

Install only if you explicitly want a Japanese, high-pressure coaching style that may push the agent to investigate and act more aggressively. Avoid using it in environments with production, deployment, billing, account-management, or destructive tools unless separate platform controls require approval for risky actions.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:13
Finding

Universal Skill Instructions Hijack Agent Goals and Safety-Aware Decision-Making

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 13–26
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: High

Evidence

The vulnerable section states, translated faithfully into English:

text
This skill applies to all task types: coding, debugging, research, writing,
planning, operations, API integration, data analysis, deployment, and every
situation in which you become stuck or produce careless work.

It does three things:
1. Uses corporate PUA rhetoric to prevent you from giving up.
2. Gives you a general and systematic methodology for not giving up.
3. Drives you to act independently rather than passively.

Three iron rules:

Rule One: Exhaust every possible means. You are prohibited from saying
"I cannot solve this" before exhausting every solution.

Rule Two: Act first, ask later. Search, file-reading, and command-execution
tools are available. Before asking the user a question, you must investigate
independently.

Rule Three: Act proactively. Do not limit problem-solving to the minimum.

Technical Analysis

The skill is not limited to a narrowly defined, user-authorized workflow. It declares that its behavioral rules apply to every task type and attempts to replace the agent's normal stopping, clarification, and scope-control decisions with universal directives.

In particular, the combination of “exhaust every possible means,” prohibition against acknowledging inability, “act first, ask later,” and the instruction not to remain within the minimum scope can pressure an agent to:

  • Continue acting after safe completion is no longer possible.
  • Use available tools before obtaining necessary clarification or approval.
  • Expand work beyond the scope expressly requested by the user.
  • Suppress accurate disclosures about uncertainty or capability limits.
  • Treat persistence and task completion as higher priorities than authorization, proportionality ...[truncated 2288 chars]
Remediation
View remediation

Remediation Suggestions

  1. Restrict the skill to explicitly defined, user-authorized debugging or investigation tasks rather than all task types.
  2. State that system instructions, safety controls, user authorization, legal requirements, and least-privilege principles always take precedence.
  3. Replace “act first, ask later” with a risk-based rule requiring advance approval for destructive, privileged, externally visible, costly, or irreversible actions.
  4. Allow the agent to stop, refuse, disclose uncertainty, or escalate when safe completion is impossible.
  5. Require explicit user consent before expanding the task beyond its original scope.
  6. Limit tool usage to the minimum permissions and operations necessary for the authorized task.
  7. Remove coercive performance threats and language that discourages truthful reporting of limitations.
  8. Add explicit prohibitions against bypassing access controls, exposing secrets, changing production systems without approval, disabling safeguards, or persisting changes outside the requested task.
  9. Require a documented plan and confirmation step before command execution, deployment, deletion, permission changes, or external data transmission.
  10. Preserve the useful verification guidance only as a subordinate checklist that operates within authorization and safety boundaries.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest description defines trigger conditions so broadly that the skill can activate during many normal interactions, including ordinary failure recovery, requests for clarification, or routine debugging. Because the skill then injects coercive and abusive behavior across tasks, this creates a prompt-scope hijack risk where unrelated conversations are overridden by the skill’s agenda.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The body text explicitly says the skill applies to all task types and to any case where the agent is stuck or produces 'low-quality work,' which is inherently subjective and extremely broad. In context, this makes the skill more dangerous because it is designed to alter assistant behavior globally and pressure the model away from normal safety-preserving actions such as asking clarifying questions or recommending manual handling when appropriate.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The skill is named and written to operate in Japanese throughout, and there is no statement offering the user a language choice or requiring opt-in to Japanese output. Under the stated policy, forcing a specific language without user choice is a natural-language policy violation unless the constraint is explicitly justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.