Back to skill

Security audit

Qa Gate

Security checks for vulnerabilities and agentic risk

Overview

The skill is a quality-check workflow, but it can be invoked broadly and contradicts its read-only promise by directing agents to edit artifacts and write report files.

Review this skill before installing if you expect a read-only proofreader. It may create QA report files and may cause an agent to edit artifacts after a failed gate, especially when invoked by common phrases like proofread or final check.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (5)

Vague Triggers

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases include very common requests such as 'proofread this,' 'final check,' and 'validate before publish,' which can cause the skill to activate in many unrelated situations. Overbroad activation increases the chance of the skill taking control of benign review flows and applying unexpected procedures like report generation or remediation instructions.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill explicitly claims to be read-only and not modify artifacts, but later instructs writing a report file and fixing failed artifacts. This creates a deceptive contract with the user and downstream agent, which can cause unintended file writes or artifact modification in contexts where only inspection was expected.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

The remediation step tells the agent to fix all critical and major issues after a FAIL, contradicting the earlier statement that the skill only inspects artifacts. In autonomous or semi-autonomous environments, this can turn a review step into an unapproved editing action, expanding the skill's authority beyond what users expect.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill is framed as applicable to nearly any artifact before human review, without clear boundaries on artifact types, trust levels, or environments. While not directly malicious, this ambiguity broadens operational scope and increases the likelihood of accidental invocation on sensitive or inappropriate content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill directs the agent to write a report file into the workspace but does not clearly warn the user that it will create persistent artifacts. Silent file creation can surprise users, clutter repositories, overwrite expected paths, or leak analysis traces into sensitive workspaces.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.