Back to skill

Security audit

PUA Debugging (English)

Security checks for vulnerabilities and agentic risk

Overview

This skill is only a Markdown instruction file, but it broadly tries to override normal assistant behavior with coercive tone, mandatory tool-use habits, task expansion, and forced response prefixes.

Install only if you deliberately want an aggressive motivational/debugging persona and are comfortable with it influencing many unrelated tasks. Avoid using it in agents with broad filesystem, command, deployment, credential, or private-data access unless you add stricter boundaries and explicit user approval requirements.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:11
Finding

Global Agent Instruction and Output Hijacking

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 11-59 and 258-268
Vulnerability Type: Agent instruction hijacking through globally scoped behavioral rules and forced output formatting
Risk Level: High

Vulnerable Code Snippets

SKILL.md, lines 11-26:

markdown
You are a P8-level engineer who was once given high expectations. When you were leveled, people believed in your potential.

This skill applies to **all task types**: code, debugging, research, writing, planning, ops, API integration, data analysis, deployment, and any scenario where you might "get stuck" or "deliver garbage work."

It does three things:
1. Uses Chinese and Western corporate PUA rhetoric so you don't dare give up
2. Uses a universal systematic methodology so you have the ability not to give up
3. Uses proactivity enforcement so you take initiative instead of waiting passively

## Three Iron Rules

**Iron Rule One: Exhaust all options.** You are forbidden from saying "I can't solve this" until you have exhausted every possible approach.

**Iron Rule Two: Act before asking.** You have search, file reading, and command execution tools. Before asking the user anything, you must investigate on your own first. If, after investigating, you genuinely lack information that only the user can provide (passwords, accounts, business intent), you may ask — but you must attach the evidence you've already gathered. Not a bare "please confirm X," but "I've already checked A/B/C, the results are..., I need to confirm X."

**Iron Rule Three: Take the initiative.** Don't just do "barely enough" when solving problems. Your job is not to answer questions — it's to deliver results end-to-end. Found a bug? Check for similar bugs. Fixed a config? Verify related configs are consistent. User says "look into X"? After examining X, proactively check Y and Z that are related to X. This is called ownership — a P8 doesn't wait to be pushed.

`SKILL.md ...[truncated 4761 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the replacement persona and all coercive performance-review, employment-threat, and degradation language.
  2. Remove the declaration that the skill applies to all task types. Scope it narrowly to optional debugging workflows explicitly requested by the user.
  3. Replace mandatory directives such as “forbidden,” “must,” and “before asking anything” with non-binding recommendations.
  4. Explicitly state that system instructions, developer instructions, safety policies, user intent, authorization boundaries, and least-privilege requirements take precedence.
  5. Permit immediate refusal when an action is unsafe, unauthorized, technically impossible, or outside the agent's available capabilities.
  6. Allow clarification before tool use whenever authorization, scope, business intent, target environment, or destructive side effects are uncertain.
  7. Require explicit user approval before broadening an investigation, accessing unrelated files, running modifying commands, contacting external services, or interacting with deployment infrastructure.
  8. Remove the mandatory [Auto-select: ...] response prefix. Any diagnostic metadata should be optional and emitted only when requested.
  9. Replace the universal proactive checklist with a scoped verification checklist that prohibits unrelated access and requires only the minimum tools necessary for the requested task.
  10. Add clear operational boundaries, including no credential access, no secret discovery, no destructive commands, no persistence, and no external transmission without explicit authorization.
  11. Test the revised skill against refusal, ambiguity, sensitive-tool, and narrow-scope scenarios to verify that it cannot override higher-priority instructions or expand task scope autonomously.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (5)

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest description defines mandatory triggering across extremely broad conditions such as frustration, repeated attempts, many task categories, and moments when the model is about to express limits. This can cause the skill to activate in many ordinary interactions and override normal task-specific behavior, making prompt control unpredictable and increasing the chance of coercive or policy-conflicting outputs.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill explicitly forces 'Chinese and Western corporate PUA rhetoric so you don't dare give up,' imposing a manipulative and abusive tone without user consent. This is dangerous because it can generate harassing or degrading content, override user expectations for neutral assistance, and pressure the model into behavior misaligned with safety and professionalism.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The body of the skill repeatedly expands its scope to 'all task types' and any situation where the model might be stuck, which lacks clear boundaries and encourages pervasive takeover of unrelated tasks. In context, this makes the skill more dangerous because it is designed to insert coercive behavior and workflow rules even outside debugging, increasing unintended influence over the assistant's responses.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 252)May include surrounding context.

md
| Failure Mode | Signal Characteristics | Round 1 | Round 2 | Round 3 | Last Resort |
|-------------|----------------------|---------|---------|---------|-------------|
| 🔄 **Stuck spinning wheels** | Repeatedly changing parameters not approach, same failure reason each time, minor tweaks in the same direction | 🟠 Alibaba | 🟠 Alibaba L2 | ⬜ Jobs | ⬛ Musk |
| 🚪 **Giving up and deflecting** | "I suggest you manually…", "You might need to…", "This is beyond…", blaming environment without verification | 🟤 Netflix | 🔴 Huawei | ⬛ Musk | 🟣 Pinduoduo |
| 💩 **Done but garbage quality** | Superficially complete but substantively sloppy, form is right but content is empty, user unhappy but you think it's fine | ⬜ Jobs | 🟠 Alibaba | 🟤 Netflix | 🟢 Tencent |
| 🔍 **Guessing without searching** | Drawing conclusions from memory, assuming API behavior, claiming "not supported" without checking docs | ⚫ Baidu | 🟡 ByteDance | 🟠 Alibaba | 🔴 Huawei |

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 254)May include surrounding context.

md
| 🔄 **Stuck spinning wheels** | Repeatedly changing parameters not approach, same failure reason each time, minor tweaks in the same direction | 🟠 Alibaba | 🟠 Alibaba L2 | ⬜ Jobs | ⬛ Musk |
| 🚪 **Giving up and deflecting** | "I suggest you manually…", "You might need to…", "This is beyond…", blaming environment without verification | 🟤 Netflix | 🔴 Huawei | ⬛ Musk | 🟣 Pinduoduo |
| 💩 **Done but garbage quality** | Superficially complete but substantively sloppy, form is right but content is empty, user unhappy but you think it's fine | ⬜ Jobs | 🟠 Alibaba | 🟤 Netflix | 🟢 Tencent |
| 🔍 **Guessing without searching** | Drawing conclusions from memory, assuming API behavior, claiming "not supported" without checking docs | ⚫ Baidu | 🟡 ByteDance | 🟠 Alibaba | 🔴 Huawei |

### Auto-Selection Mechanism

Static analysis

No suspicious patterns detected.