Back to skill

Security audit

杠精 Gangjing

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed contrarian review skill, but it can create and run local attack harnesses against workspace code with broad automatic escalation triggers.

Install only if you want an intentionally adversarial Chinese/English review assistant. Before allowing it to run code, use an isolated workspace or container, confirm each harness execution explicitly, and avoid pointing it at repositories with secrets, credentials, production data, or unsafe top-level code.

Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (10)

Description-Behavior Mismatch

Medium
Confidence
95% confidence
Finding
The document explicitly describes a built-in 'code attack engine' and frames code execution as a supported workflow, which expands the skill from verbal review into active exploit-style testing. Even though the broader skill description says code testing should require explicit user request, this file operationalizes attack execution and evidence collection in a way that could normalize or trigger risky code-running behavior against local workspace code.

Context-Inappropriate Capability

Medium
Confidence
96% confidence
Finding
The instructions permit creating a temporary harness and then executing it when a ready-made script is unavailable, which materially broadens the skill's ability to generate and run new code. That is dangerous because it enables ad hoc execution infrastructure beyond simple review, increasing the chance of unsafe code execution, unintended side effects, and bypass of the skill's stated default oral-review boundary.

Vague Triggers

Medium
Confidence
91% confidence
Finding
The README states the skill will '自动介入' during normal conversation, which implies broad ambient activation without clear scoping or confirmation. In an agent environment, this can cause unintended invocation on ordinary user messages, leading to disruptive behavior, policy drift, and accidental escalation into code-review or attack-oriented workflows when the user did not explicitly request the skill.

Vague Triggers

Medium
Confidence
88% confidence
Finding
Using a common phrase like '杠一下' as an activation trigger without constraints increases the chance of accidental invocation in ordinary discussion, quoted text, or meta-conversation. Because this skill is intentionally adversarial in tone and may escalate into aggressive review behavior, accidental triggering can degrade user experience and cause the assistant to act outside the intended task scope.

Vague Triggers

Medium
Confidence
91% confidence
Finding
The manifest says the skill should respond to essentially anything the user says, which creates an overly broad activation surface. In a skill that is explicitly designed to be contrarian and can escalate into code-reading or attack-oriented behavior, this increases the chance of inappropriate triggering, policy bypass through context confusion, or user-hostile responses in unrelated conversations.

Vague Triggers

Medium
Confidence
95% confidence
Finding
The auto-upgrade rules trigger more aggressive behavior based on broad phrases like 'definitely fine' or 'everyone agreed,' without sufficient scoping to the task domain. That means ordinary conversational language could activate an unnecessarily adversarial mode, and in this skill's context that mode is tied to stronger review/attack workflows, making accidental escalation more likely.

Natural-Language Policy Violations

Medium
Confidence
88% confidence
Finding
The skill mandates a forceful 'contrarian' persona and speaking style without user opt-in, which can override user preferences and produce manipulative or disruptive interactions. While this is not a classic code-execution issue, it is a real safety and quality problem because it can pressure the model into hostile behavior even when the user's request does not call for that tone.

Vague Triggers

Medium
Confidence
93% confidence
Finding
The trigger phrases are overly broad and include common statements such as expressing an idea, a decision, or confidence in a plan. This can cause unintended invocation in normal conversations, which is especially risky here because the skill is explicitly designed to produce adversarial, high-intensity critique and can escalate into code-attack behavior when certain conditions are met.

Vague Triggers

Medium
Confidence
88% confidence
Finding
The trigger table maps very broad phrases like "review this" to a specific operating mode. Because such phrasing is common in normal user requests, the skill can be unintentionally pushed into a more aggressive review behavior than the user actually intended, which creates prompt-routing ambiguity and increases the chance of overreaching or disruptive responses. In this skill’s context, that is a genuine safety issue because mode selection directly controls how adversarial the agent becomes.

Natural-Language Policy Violations

Medium
Confidence
81% confidence
Finding
The skill is written to operate in Chinese and its output templates are Chinese-only, without offering a language negotiation path. This can cause the agent to ignore user language preference, reduce transparency, and create misunderstandings in multilingual environments; while not a direct code-execution risk, it is a real interaction-safety and usability problem. In a review/red-team skill, misunderstanding the output language can also cause users to miss important warnings or misapply guidance.

Static analysis

No suspicious patterns detected.