Back to skill

Security audit

E2e Test Orchestrator

Security checks for vulnerabilities and agentic risk

Overview

The skill has a coherent E2E testing purpose, but its Docker runner exposes a real command-injection and supply-chain risk that merits Review before installation.

Review before installing or running. Use this only in repositories where E2E test execution and artifact capture are acceptable, avoid passing untrusted filter strings, and prefer fixing the Docker script to pass arguments safely, enforce locked dependencies, remove npm i fallback, pin container images by digest, and limit writable mounts and network access. Treat captured screenshots, traces, videos, and logs as potentially sensitive.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/run-playwright-docker.sh:14
Finding

Shell Command Injection Through the Playwright Test Filter

Content
View full analysis

Vulnerability Details

File Location: scripts/run-playwright-docker.sh, lines 14–16 and 27
Vulnerability Type: OS command injection through unsafe shell-string construction
Risk Level: High

Vulnerable Code

bash
GREP_ARG=""
if [[ -n "$FILTER" ]]; then
  GREP_ARG="--grep '$FILTER'"
fi

echo "[run][docker] 镜像: $IMAGE"
echo "[run][docker] 输出目录: $OUT_DIR"

docker run --rm -t \
  -v "$PWD":/work \
  -w /work \
  "$IMAGE" \
  bash -lc "npm ci || npm i; npx playwright test --reporter=line,html --output='$OUT_DIR' $GREP_ARG" | tee "$OUT_DIR/run.log"

Technical Analysis

The user-controlled FILTER argument is interpolated into GREP_ARG and subsequently inserted into a command string executed by bash -lc. Wrapping the value in single quotes does not provide protection because an attacker can include another single quote in the input, terminate the intended quoted context, and append additional shell commands.

The local runner correctly uses a Bash array, but the Docker runner converts the filter back into shell syntax. The Docker container also mounts the current project directory read-write at /work, so an injected command can modify or delete repository content and generated artifacts.

Attack Path

  1. An attacker obtains the ability to control the test filter supplied to the Docker runner, directly or through run-playwright-auto.sh.

  2. The attacker provides a filter containing shell metacharacters, for example:

    bash
    ./scripts/run-playwright-docker.sh "'; touch /work/pwned; #"
    
  3. The script constructs a command equivalent to:

    bash
    npx playwright test ... --grep ''; touch /work/pwned; #'
    
  4. bash -lc parses the injected semicolon as a command separator.

  5. The injected command executes inside the Playwright container with write access to the mounted project directory.

More destructive payloads could overwrite tests, source ...[truncated 665 chars]

Remediation
View remediation

Remediation Suggestions

  • Do not concatenate the filter into a command interpreted by bash -lc.

  • Pass the filter as a positional argument to a fixed shell program, preserving it as data:

    bash
    docker run --rm -t \
      -v "$PWD":/work \
      -w /work \
      "$IMAGE" \
      bash -c 'npm ci && npx playwright test --reporter=line,html --output="$1" --grep "$2"' \
      _ "$OUT_DIR" "$FILTER"
    
  • Handle the empty-filter case separately so that --grep is omitted rather than passed an empty value.

  • Prefer an entrypoint script that constructs and executes a command array without additional shell evaluation.

  • Mount project source read-only where practical and mount only the results directory as writable.

  • Disable container network access with --network=none when tests do not require it.

  • Add regression tests containing single quotes, semicolons, command substitutions, spaces, and newline characters in the filter.

T08 · Insecure Dependencies

Warning
Location
scripts/run-playwright-docker.sh:10
Finding

Unpinned and Implicitly Retrieved Executable Dependencies

Content
View full analysis

Vulnerability Details

File Location: scripts/run-playwright.sh, line 13; scripts/run-playwright-docker.sh, lines 10 and 27; references/playwright-重点与docker兜底.md, lines 33–34
Vulnerability Type: Insecure dependency resolution and mutable container image selection
Risk Level: Medium

Vulnerable Code

From scripts/run-playwright.sh:

bash
CMD=(npx playwright test --reporter=line,html --output="$OUT_DIR")

From scripts/run-playwright-docker.sh:

bash
IMAGE="${PW_DOCKER_IMAGE:-mcr.microsoft.com/playwright:v1.58.2-jammy}"
bash
docker run --rm -t \
  -v "$PWD":/work \
  -w /work \
  "$IMAGE" \
  bash -lc "npm ci || npm i; npx playwright test --reporter=line,html --output='$OUT_DIR' $GREP_ARG" | tee "$OUT_DIR/run.log"

From references/playwright-重点与docker兜底.md:

bash
npm i -D playwright
npx playwright install chromium

Technical Analysis

The execution flow does not consistently require preinstalled, lockfile-verified dependencies:

  • npx playwright may retrieve and execute a package when the expected local binary is absent.
  • The documented installation command does not pin Playwright to a specific version.
  • The Docker runner falls back from deterministic npm ci to npm i, which can resolve or update dependencies differently from the committed lockfile state.
  • npm installation can execute package lifecycle scripts.
  • The Docker image is selected through the unrestricted PW_DOCKER_IMAGE environment variable and is referenced by a mutable tag rather than an immutable digest.
  • The selected container receives a read-write mount of the repository.

These behaviors weaken reproducibility and expand the supply-chain execution surface. A compromised package release, unsafe dependency resolution, compromised image tag, or attacker-controlled image setting could introduce executable code during a test run.

Attack Path

A package-bas ...[truncated 1522 chars]

Remediation
View remediation

Remediation Suggestions

  • Require a committed and reviewed lockfile and use npm ci without falling back to npm i.

  • Fail closed when lockfile installation fails rather than performing mutable dependency resolution.

  • Pin Playwright and all direct dependencies to reviewed versions.

  • Invoke only the installed local binary, for example:

    bash
    npx --no-install playwright test
    

    or:

    bash
    ./node_modules/.bin/playwright test
    
  • Pin the Playwright container image by immutable digest rather than only by tag.

  • Remove unrestricted image overrides or validate PW_DOCKER_IMAGE against an explicit allowlist of approved digest-pinned images.

  • Mount source files read-only and provide a separate writable mount for test artifacts.

  • Use --network=none when network connectivity is unnecessary.

  • Restrict credentials passed to test processes and containers to the minimum required scope.

  • Add dependency integrity, provenance, vulnerability, and container-image scanning to CI before execution.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (16)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill explicitly requires preserving screenshots, videos, traces, and logs as test evidence, but it does not warn that these artifacts can capture credentials, session tokens, PII, internal URLs, or other sensitive system data. In an E2E testing context this is materially risky because the workflow also emphasizes structured reporting and artifact retention, increasing the chance of broad storage or sharing of sensitive evidence.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The entire skill guidance is written in Chinese and does not indicate that users may choose another language or that the skill is intentionally limited to a Chinese-speaking context. Under the policy, language constraints should be opt-in or clearly justified when a specific locale is enforced.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
85% confidence
Finding

The documentation invokes npx playwright install chromium without pinning the package version. In environments where Playwright is not already installed locally as a dev dependency or where resolution falls back unexpectedly, npx may fetch a newer package version than intended, reducing reproducibility and introducing supply-chain risk if an unexpected release is pulled at execution time.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

The Docker fallback command runs npx playwright test without version pinning inside the container. Although the image tag is pinned, npx still executes whatever Playwright package is present in the mounted project or downloads one if missing, which can lead to non-reproducible behavior and exposure to unreviewed package versions during test execution.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

This command again uses npx playwright test without pinning the package version. In an automation skill that encourages execution in Docker or CI, unpinned runtime tool resolution increases the chance of pulling a changed or compromised package version and makes auditability of test runs weaker.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
83% confidence
Finding

The guidance references npx playwright install-deps as an operational path without explicit version pinning. Even in instructional content, recommending unpinned npx usage can cause users or agents to execute an unexpected package version, which is a supply-chain and reproducibility concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script's natural-language comments and runtime output are entirely in Chinese, which imposes a specific language on users without any opt-in or indication that the skill is region-specific. This matches the language/locale policy violation category because the file provides no alternative language option or justification for the constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script prints operational messages such as image, output directory, and completion status exclusively in Chinese. Because there is no user choice or documented reason for the fixed language, this is a natural-language policy violation under the locale/language rule.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The script executes npx playwright inside a container after npm ci || npm i, which can resolve and run tooling from the project dependency graph or fetch packages if the environment is not fully locked down. In an E2E orchestration skill that may run against varied repositories, this increases supply-chain risk because an unexpected or compromised Playwright package version could be executed during test runs.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

The script invokes npx playwright without pinning an explicit package version, so npx may resolve and execute whatever Playwright version is available from the local project or, in some environments, fetch a package dynamically. That creates a supply-chain and reproducibility risk: a compromised dependency, unexpected major-version change, or registry-level issue could alter test execution or run untrusted code during a CI job.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language content of the skill is presented entirely in Chinese and does not indicate that users may choose another language or that the skill is intentionally limited to a Chinese-speaking context. Per the policy, forcing a specific language without opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This markdown template forces a specific language/locale for all users, and the file does not indicate that Chinese is optional or that the template is intended only for a Chinese-speaking or region-specific context. Under the policy, fixed language requirements without user opt-in are natural-language policy violations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

This shell script uses user-facing comments and console output entirely in Chinese, including usage instructions and status/error messages. Because the file provides no opt-in, fallback language, or documented region-specific justification, it violates the language/locale policy criterion for natural-language content.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.