Back to skill

Security audit

LLM as Judge

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent review-helper, but it may send sensitive code or plans to other model providers without requiring user approval or redaction guidance.

Install only if you are comfortable with a judge model seeing the material being reviewed. For private repositories, customer data, secrets, incident details, payment/auth code, or trading systems, use only approved providers and redact sensitive content before invoking this workflow.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · SKILL.md (reported line 70)May include surrounding context.

md
- **Vague judge output is useless** — If the judge says "looks good" without specifics, the prompt is too weak. Always require the judge to produce scored dimensions + specific actionable items, even if approving.
- **Judge scope creep** — Judges sometimes rewrite the entire plan instead of reviewing it. Constrain the verdict to APPROVE / REVISE / REJECT with specific feedback, not a replacement solution.
- **Approval rate drift** — If the judge approves >80% of submissions, the model pairing is too similar or the prompts are too lenient. Target 60-70% approval rate.
- **Don't judge trivial tasks** — A 50-line CSS fix doesn't need cross-model review. Use the activation criteria in this skill strictly.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill explicitly encourages sending plans, code, architecture, and security-sensitive material to a different model/provider, but it provides no warning about third-party data disclosure, retention, logging, or contractual/privacy constraints. In practice, this can cause users to exfiltrate proprietary source code, credentials embedded in configs, incident details, or regulated data to external services during routine use of the skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The model-pairing section operationalizes use of named alternate providers and presents it as best practice without any safeguards around data sharing. Because this section is actionable and specific, it increases the likelihood that users will paste sensitive code or architecture directly into external systems that may have different security, retention, or training policies.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.