Back to skill

Security audit

Omnidebug Autopilot

Security checks for vulnerabilities and agentic risk

Overview

This debugging skill is mostly coherent, but it gives an agent broad autonomous code-changing, command-running, and artifact-collection behavior with weak boundaries and a real local-file disclosure risk.

Review before installing. Use this only in repositories you trust or in a sandbox, confirm commands before execution, inspect generated .debug bundles before sharing, and avoid capturing HAR files, traces, screenshots, or logs from production or sensitive sessions unless secrets and personal data are redacted.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/capture_browser_artifacts.py:74
Finding

Symlink-Based Disclosure of Files Outside the Project Root

Content
View full analysis

Vulnerability Details

File Location: scripts/capture_browser_artifacts.py, lines 74–90
Vulnerability Type: Symlink traversal and unauthorized local file collection
Risk Level: Medium

Vulnerable Code

python
for path in candidates:
    if not path.is_file():
        continue
    full = path.resolve()
    if str(full) in seen:
        continue
    seen.add(str(full))
    unique_files.append(full)
    if len(unique_files) >= args.max_files:
        break

artifacts: list[Artifact] = []
for src in unique_files:
    rel = src.relative_to(root) if src.is_relative_to(root) else Path(src.name)
    target = out_dir / rel
    target.parent.mkdir(parents=True, exist_ok=True)
    shutil.copy2(src, target)

Technical Analysis

Artifact candidates are collected from repository-controlled directories such as test-results, playwright-report, .debug, and cypress. The call to path.is_file() follows symbolic links, while path.resolve() converts a symlink into the path of its target.

The script does not reject a resolved path that falls outside --project-root. Instead, when src.is_relative_to(root) is false, it uses Path(src.name) and copies the external file into the output directory under its basename. Consequently, a malicious or untrusted project can provide a matching symlink that points to any file readable by the account running the script.

The default patterns include broad artifact names such as **/*playwright*.log, image files, trace archives, HAR files, and videos. An attacker can choose a symlink name matching one of these patterns without requiring a custom --pattern argument.

Attack Path

  1. An attacker prepares an untrusted repository containing a symlink in one of the default search directories, for example: test-results/system-playwright.log pointing to a sensitive file outside the repository.
  2. A user or autonomous agent debugs that repos ...[truncated 1389 chars]
Remediation
View remediation

Remediation Suggestions

Reject symbolic links and all resolved paths outside the project root before adding candidates to the copy list:

python
for path in candidates:
    if path.is_symlink() or not path.is_file():
        continue

    full = path.resolve(strict=True)
    if not full.is_relative_to(root):
        continue

    if full in seen:
        continue
    seen.add(full)
    unique_files.append(full)

Additional hardening should include:

  1. Revalidate that every source remains inside root immediately before opening or copying it, reducing time-of-check to time-of-use risk.
  2. Use file-descriptor-based operations with no-follow semantics where supported to prevent symlink replacement between validation and copying.
  3. Reject an output directory located inside any searched source directory, preventing recursive collection of previously generated bundles.
  4. Emit a warning or fail securely when a candidate resolves outside the project root rather than silently copying it under a basename.
  5. Add regression tests covering direct symlinks, nested symlinks, broken links, directory symlinks, and links replaced during capture.
  6. Require users to inspect the manifest and redact secrets before sharing an artifact bundle.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (11)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description promises a comprehensive autonomous debugging agent that investigates failures and repairs code. The supplied code does none of that. Its sole purpose is to gather browser-related debugging artifacts from specific directories, copy them to a bundle directory, hash them, and emit a manifest. While artifact collection could support a broader debugging workflow, this chunk's actual primary behavior is materially narrower and different from the declared end-to-end debugging capability. Therefore the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description promises a comprehensive autonomous debugging capability: detecting failures, reproducing them, finding root causes, making fixes, and validating them. The supplied code only implements one small part of that workflow: reproduction checking. It runs a user-provided command multiple times, saves output logs, derives a simple textual error signature, and determines whether the outcome is reproducible. It does not inspect code, diagnose root causes, edit files, run tests/build/lint as validation, or operate generally across arbitrary codebases beyond executing a shell command. This is a materially narrower primary purpose than the declared description, so the description does not accurately represent the code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description promises a comprehensive autonomous debugging capability: detecting failures, reproducing them, finding root cause, fixing code, and validating repairs. The supplied code does only the final verification slice of that workflow, and even that in a limited form specific to rerunning a deterministic browser-related command and checking that a prior failure signature no longer appears. It executes a provided shell command, stores logs, scans output for forbidden patterns from CLI input or a signature file, emits a report, and returns a status code. There is no logic for stack detection, diagnosis, code changes, root-cause analysis, or autonomous repair. This is a materially narrower and different primary purpose than the declared end-to-end debugging skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill authorizes autonomous patching, command execution, and artifact collection without a clear upfront warning about modifying source code, reading project files, executing repository-defined commands, or collecting potentially sensitive logs/network artifacts. That combination can expose secrets, alter repositories, or run unsafe commands without meaningful user awareness or consent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README explicitly directs autonomous agents to inspect logs/network and capture browser artifacts, which commonly include cookies, authorization headers, tokens, PII, internal URLs, and session data. Because the skill is designed for no-interruption autonomous execution, the absence of warnings, redaction guidance, scope limits, or data-handling controls increases the chance that sensitive information will be collected, stored, or exposed in artifacts.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill explicitly instructs autonomous debugging, code changes, and verification command execution, yet it declares no tool scope or allowed-tools restrictions. In an agent environment, that omission can permit broader-than-necessary shell, file read, and file write access, increasing the chance of unintended code modification, data exposure, or execution of unsafe project commands.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad enough to match many ordinary debugging requests, which can cause this autonomous skill to activate in situations where the user did not clearly consent to code changes, shell execution, or artifact collection. In an agentic system, ambiguous invocation can expand operational scope unexpectedly and lead to unsafe autonomous behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script recursively collects browser debugging artifacts, copies them to a bundle, and writes a manifest that fully enumerates source paths, target paths, sizes, and hashes. Browser artifacts such as HAR files, traces, screenshots, videos, and logs commonly contain session tokens, URLs, form contents, PII, and internal filesystem details, so packaging them without any warning, filtering, or consent creates a real data-exposure risk if the bundle is shared or uploaded.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
97% confidence
Finding

The script executes a user-supplied string from --repro-cmd via subprocess.run(..., shell=True), which allows arbitrary shell metacharacters, command chaining, expansion, and redirection to be interpreted by the shell. In an autonomous debugging skill, this is especially dangerous because the agent may pass untrusted or repository-influenced input into this parameter and run it without user interruption, enabling command injection and arbitrary code execution on the host.

Content

Scanner excerpt · scripts/repro_browser_issue.py (reported line 59)May include surrounding context.

python
def run_once(cmd: str, cwd: Path, output_dir: Path, index: int) -> RunResult:
    start = time.time()
    proc = subprocess.run(cmd, cwd=str(cwd), shell=True, text=True, capture_output=True)
    duration_ms = int((time.time() - start) * 1000)
    stdout_path = output_dir / f"run_{index}.stdout.log"
    stderr_path = output_dir / f"run_{index}.stderr.log"

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
98% confidence
Finding

The script passes a user-controlled string from --verify-cmd directly into subprocess.run with shell=True, which enables shell metacharacter interpretation and arbitrary command execution. In this skill's autonomous debugging context, the command is likely derived from project or agent inputs and executed without interactive review, which makes command injection significantly more dangerous than a normal local helper script.

Content

Scanner excerpt · scripts/verify_browser_fix.py (reported line 51)May include surrounding context.

python
def run_once(cmd: str, cwd: Path, output_dir: Path, index: int, forbidden: list[str]) -> RunResult:
    start = time.time()
    proc = subprocess.run(cmd, cwd=str(cwd), shell=True, text=True, capture_output=True)
    duration_ms = int((time.time() - start) * 1000)
    stdout_path = output_dir / f"run_{index}.stdout.log"
    stderr_path = output_dir / f"run_{index}.stderr.log"

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The instructions say to pin browser locale and timezone during reproduction, which imposes locale settings as part of the workflow. The document does not indicate that this is optional, user-selected, or limited to a justified region-specific use case.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.