Back to skill

Security audit

Gauntlet

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed adversarial-review workflow with an optional local ledger; users should treat install provenance and stored review notes with normal care.

Install from a pinned and trusted source when possible, especially for user-wide installation. Use the skill for deep reviews rather than routine edits, keep any SQLite ledger in a private directory outside version control, and review stored notes before including sensitive code, secrets, customer data, or production details.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (27)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a general-purpose adversarial review and repair skill for analyzing code, systems, tests, and workflows. The supplied code instead implements a concrete offline validator for a specific package format. It checks for required files, validates SKILL.md metadata formatting and size limits, verifies local markdown links, parses Python and JSON files for syntax, and rejects symlinks. These behaviors are not supporting details of a security review tool; they define a distinct primary purpose: package conformance checking. There is no sign of adversarial analysis, defect triage, iterative repair orchestration, sub-agent workflow management, or broader security auditing. Therefore the code chunk materially differs from the declared purpose.

Content

No source excerpt is available for this finding.

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · scripts/gauntlet.py (reported line 139)May include surrounding context.

python
def validate_state(state: Any) -> dict[str, Any]:
    """Reject corrupted/foreign state before any workflow or write operation."""
    if not isinstance(state, dict) or type(state.get("schema")) is not int or state["schema"] != SCHEMA:
        raise GauntletError("Unsupported or corrupt state schema; restore a trusted ledger")
    for key in ("scope", "revision", "created_at", "updated_at", "version"):
        text(state.get(key), key)
    for key, options in (("mode", {"review", "repair"}), ("tier", set(TIERS)),

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · scripts/gauntlet.py (reported line 148)May include surrounding context.

python
for key in ("round", "max_rounds", "required_clean"):
        integer(state.get(key), key, 1, 20)
    if state["round"] > state["max_rounds"]:
        raise GauntletError("Corrupt state: round exceeds budget")
    for key in ("closed", "changed_in_round"):
        if type(state.get(key)) is not bool:
            raise GauntletError(f"Corrupt state: {key} is not boolean")

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · scripts/gauntlet.py (reported line 151)May include surrounding context.

python
for key in ("round", "max_rounds", "required_clean"):
        integer(state.get(key), key, 1, 20)
    if state["round"] > state["max_rounds"]:
        raise GauntletError("Corrupt state: round exceeds budget")
    for key in ("closed", "changed_in_round"):
        if type(state.get(key)) is not bool:
            raise GauntletError(f"Corrupt state: {key} is not boolean")

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · scripts/gauntlet.py (reported line 154)May include surrounding context.

python
for key in ("round", "max_rounds", "required_clean"):
        integer(state.get(key), key, 1, 20)
    if state["round"] > state["max_rounds"]:
        raise GauntletError("Corrupt state: round exceeds budget")
    for key in ("closed", "changed_in_round"):
        if type(state.get(key)) is not bool:
            raise GauntletError(f"Corrupt state: {key} is not boolean")

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · scripts/gauntlet.py (reported line 156)May include surrounding context.

python
for key in ("round", "max_rounds", "required_clean"):
        integer(state.get(key), key, 1, 20)
    if state["round"] > state["max_rounds"]:
        raise GauntletError("Corrupt state: round exceeds budget")
    for key in ("closed", "changed_in_round"):
        if type(state.get(key)) is not bool:
            raise GauntletError(f"Corrupt state: {key} is not boolean")

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · scripts/gauntlet.py (reported line 163)May include surrounding context.

python
for key in ("round", "max_rounds", "required_clean"):
        integer(state.get(key), key, 1, 20)
    if state["round"] > state["max_rounds"]:
        raise GauntletError("Corrupt state: round exceeds budget")
    for key in ("closed", "changed_in_round"):
        if type(state.get(key)) is not bool:
            raise GauntletError(f"Corrupt state: {key} is not boolean")

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · scripts/gauntlet.py (reported line 165)May include surrounding context.

python
for key in ("round", "max_rounds", "required_clean"):
        integer(state.get(key), key, 1, 20)
    if state["round"] > state["max_rounds"]:
        raise GauntletError("Corrupt state: round exceeds budget")
    for key in ("closed", "changed_in_round"):
        if type(state.get(key)) is not bool:
            raise GauntletError(f"Corrupt state: {key} is not boolean")

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

The README instructs users to run npx skills add ... without pinning an exact package version. npx resolves and may fetch the latest published package, so a compromised upstream release or unexpected breaking change could cause users to install or execute unreviewed code. In a security-focused skill, this is more concerning because users may place additional trust in the project and follow install instructions directly.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

This command again relies on unpinned npx skills, which can pull a newer or maliciously altered package at execution time. That creates a supply-chain risk where readers following the README may run code different from what the repository authors reviewed. The risk is amplified slightly because this variant performs a global installation, increasing persistence of any compromise.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

The host-specific Codex install example also uses unpinned npx skills, preserving the same supply-chain exposure. Users may assume host-targeted instructions are curated and safe, but the command still executes whatever version npm resolves at that moment.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

The Claude Code example repeats the unpinned npx skills pattern, leaving installation dependent on mutable upstream state. If an attacker controls or compromises the package or a dependency, users could execute arbitrary code during installation. The README context makes this a real operational risk rather than a purely theoretical issue because it is an explicit copy-paste command.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill enables implicit invocation (allow_implicit_invocation: true) without any visible trigger constraints, scope limits, or approval gates. Because this skill is explicitly designed for adversarial review and may steer workflows toward deep inspection or iterative fixing, unbounded auto-invocation can cause unintended activation, surprise actions, or prompt-surface expansion in contexts where the user did not clearly request this behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger set uses broad phrases such as 'review', 'red-team', 'production-readiness', and 'find and fix' without defining sharper activation boundaries or negative constraints. This can cause the skill to trigger on general diagnostic or editing requests that merely resemble adversarial review, leading to inappropriate invocation, scope escalation, or unexpected use of deeper review workflows.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger is broad enough to activate on a wide range of ordinary AI-related tasks, which can cause unintended invocation of an adversarial review skill outside the user's actual intent. In an agent ecosystem, overly broad activation can redirect workflows, consume budget, and introduce inappropriate security-review behavior into benign requests, especially because this skill is explicitly designed for deep review and red-team activity.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger lists broad categories like 'implementation logic', 'libraries', and 'schema/API changes' without clearly constraining when this skill should activate or when it should not. Because there is no explicit scope boundary or negative examples, the invocation conditions are ambiguous and could match a wide range of ordinary engineering discussions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger lists broad concepts like "browser/app interaction," "UI state," and "client-side changes" without clear boundaries or exclusions. This makes it unclear when the skill should activate versus when a general frontend discussion should not invoke it.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger line lists very broad concepts like "large inputs," "caches," "queues," and "production-readiness claims" without defining whether all must be present or what context qualifies. This can cause unintended invocation during ordinary engineering conversations because the activation boundary is not narrowly specified.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The activation line lists broad nouns like "plans," "requirements," and "prose claims" without clear boundaries, exclusions, or concrete invocation phrases. This makes it unclear when the skill should activate versus when a general writing or review request should not invoke it.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_tracker.py (reported line 348)May include surrounding context.

python
with self.assertRaises(g.GauntletError): g.db_path(str(link))

    def test_concurrent_writers_do_not_lose_updates(self):
        processes = [subprocess.Popen([sys.executable, str(SCRIPT), "record", "--db", str(self.db),
                                       "--event", json.dumps({"type": "note", "text": f"writer-{i}"})],
                                      stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True) for i in range(6)]
        for process in processes:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_tracker.py (reported line 357)May include surrounding context.

python
self.assertEqual(len(g.read_state(self.db)["notes"]), 6)

    def test_cli_status_and_blocked_exit(self):
        p = subprocess.run([sys.executable, str(SCRIPT), "gate", "--db", str(self.db)], capture_output=True, text=True, timeout=10)
        self.assertEqual(p.returncode, 3)
        self.assertEqual(json.loads(p.stdout)["outcome"], "BLOCKED")

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_tracker.py (reported line 362)May include surrounding context.

python
self.assertEqual(json.loads(p.stdout)["outcome"], "BLOCKED")

    def test_cli_bad_input_has_no_traceback(self):
        p = subprocess.run([sys.executable, str(SCRIPT), "record", "--db", str(self.db), "--event", "[]"], capture_output=True, text=True, timeout=10)
        self.assertEqual(p.returncode, 2)
        self.assertNotIn("Traceback", p.stderr)

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_tracker.py (reported line 367)May include surrounding context.

python
self.assertNotIn("Traceback", p.stderr)

    def test_stdin_event(self):
        p = subprocess.run([sys.executable, str(SCRIPT), "record", "--db", str(self.db), "--input", "-"], input='{"type":"note","text":"stdin"}', capture_output=True, text=True, timeout=10)
        self.assertEqual(p.returncode, 0, p.stderr)
        self.assertEqual(g.read_state(self.db)["notes"][0]["text"], "stdin")

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · tests/test_tracker.py (reported line 374)May include surrounding context.

python
def test_oversized_input_rejected(self):
        path = self.root / "huge.json"
        path.write_bytes(b" " * (g.MAX_INPUT + 1))
        p = subprocess.run([sys.executable, str(SCRIPT), "record", "--db", str(self.db), "--input", str(path)], capture_output=True, text=True, timeout=10)
        self.assertEqual(p.returncode, 2)
        self.assertIn("exceeds", p.stderr)

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_tracker.py:19