Back to skill

Security audit

Playwright 网页自动化工具

Security checks for vulnerabilities and agentic risk

Overview

This browser-testing skill is broadly coherent, but it needs Review because raw user text is passed through a shell command and a helper script exposes broad local command execution.

Install only if you trust the skill runner sandbox and the users who can invoke it. Review or fix the agent entrypoint so user input is passed as an argument without shell evaluation, restrict `with_server.py` to approved commands, and avoid using file:// targets or console capture on sensitive local content.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
agents/main/agent.yaml:18
Finding

Shell Command Injection Through Raw Agent Input Interpolation

Content
View full analysis

Vulnerability Details

File Location: agents/main/agent.yaml, lines 18–21
Vulnerability Type: Shell command injection
Risk Level: High

Vulnerable Code

yaml
steps:
  - name: run_task
    run: |
      python3 scripts/run_task.py "{{input}}"

Technical Analysis

The agent entrypoint directly interpolates attacker-controlled {{input}} into a shell command. Wrapping the value in double quotes does not neutralize shell metacharacters that remain active inside double-quoted strings, particularly command substitution expressions such as $(...) and backticks.

If the agent platform renders this template and executes the resulting run block through a shell, malicious input can cause the shell to execute an injected command before scripts/run_task.py starts. The Python script's URL parsing does not mitigate the flaw because shell evaluation occurs before Python receives its arguments.

For example, an input containing a valid task followed by a command substitution expression could cause the embedded command to execute during shell expansion.

Attack Path

  1. An attacker submits a crafted Skill request containing shell command substitution syntax, such as $(attacker_command).
  2. The platform substitutes the complete request into {{input}}.
  3. The rendered run block is passed to a shell.
  4. The shell evaluates the command substitution despite the surrounding double quotes.
  5. The injected command runs with the privileges and environment of the agent runner.
  6. The shell then invokes scripts/run_task.py with the expanded command output as part of its argument.

Impact Assessment

Successful exploitation provides arbitrary command execution under the operating-system identity used by the Skill runner. Depending on the runner's permissions and isolation controls, an attacker could:

  • Read files and environment variables accessible to the runner.
  • Modify or delete accessible files.
  • Execute installed programs and scri ...[truncated 448 chars]
Remediation
View remediation

Remediation Suggestions

  • Do not place raw user input into shell command text.
  • Configure the agent platform to invoke the executable with an argument array, passing {{input}} as one opaque argument without a shell.
  • If direct argument-array execution is unavailable, pass the request through standard input or a dedicated environment variable and read it from Python.
  • Prefer a structured request format with explicit fields such as mode and target rather than forwarding unrestricted natural-language input to a command template.
  • Apply strict validation to structured fields after safe transport. Restrict modes to an allowlist and validate target schemes and locations according to the intended threat model.
  • Do not rely on double quotes or custom character replacement as a substitute for avoiding shell interpretation.
  • Add a regression test using command-substitution characters and verify that no injected command is executed.

A safer conceptual invocation is an argv-based execution equivalent to:

python
subprocess.run(
    ["python3", "scripts/run_task.py", user_input],
    shell=False,
    check=True,
)

The platform configuration must preserve user_input as one argument rather than first composing and evaluating a shell command.

Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (16)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The documented purpose is browser automation and testing, but the detected behavior includes starting local services, monitoring ports, executing arbitrary commands, and terminating processes, which are materially broader and more dangerous than described. This mismatch can mislead operators into approving a skill that effectively has general local code-execution and process-control behavior.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
98% confidence
Finding

This duplicate finding points to the same dangerous sink: shell-based process creation from an externally supplied command string. In the context of an automation skill, this is especially risky because the helper script is designed to run before tests and may be invoked automatically, increasing the chance of abuse if parameters are not tightly controlled.

Content

Scanner excerpt · scripts/with_server.py (reported line 47)May include surrounding context.

python
try:
        for i, (cmd, port) in enumerate(zip(args.servers, args.ports), start=1):
            print(f'正在启动服务 {i}/{len(args.servers)}:{cmd}')
            process = subprocess.Popen(cmd, shell=True)
            server_processes.append(process)

            print(f'等待端口 {port} 启动完成...')

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
98% confidence
Finding

This duplicate finding points to the same dangerous sink: shell-based process creation from an externally supplied command string. In the context of an automation skill, this is especially risky because the helper script is designed to run before tests and may be invoked automatically, increasing the chance of abuse if parameters are not tightly controlled.

Content

Scanner excerpt · scripts/with_server.py (reported line 47)May include surrounding context.

python
try:
        for i, (cmd, port) in enumerate(zip(args.servers, args.ports), start=1):
            print(f'正在启动服务 {i}/{len(args.servers)}:{cmd}')
            process = subprocess.Popen(cmd, shell=True)
            server_processes.append(process)

            print(f'等待端口 {port} 启动完成...')

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The README explicitly demonstrates opening local files via file:///tmp/demo.html but does not warn users that the skill can access local filesystem content through file:// URLs. In an automation skill that drives a browser, this expands the reachable attack surface beyond web pages to local artifacts, which can lead to unintended disclosure or testing against sensitive local content if users do not understand the capability.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
84% confidence
Finding

The skill advertises executable behavior that implies shell access, environment interaction, and file-writing capability, but it does not declare any explicit tool scope or permissions boundary. In an agent ecosystem, this increases the risk of over-privileged execution and makes it harder for reviewers and runtime policy engines to constrain what the skill can do.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description and the full markdown content are written in Chinese, but there is no indication that the skill is intended only for Chinese-speaking users or that language selection is optional. This creates a natural-language policy concern because it forces a locale/language choice without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description and instructions are written entirely in Chinese and present the assistant role in Chinese without offering a language choice. This can violate language/locale policy when users have not opted into Chinese or when the skill is expected to operate in the user's preferred language.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill uses broad keyword-based routing such as matching mentions of screenshots, elements, console logs, or file:// paths. This can cause the automation flow to trigger unintentionally on ambiguous user input, which is risky in a browser automation skill because it may visit pages, load local content, or perform actions the user did not clearly request.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This code persists collected console logs to /tmp/console.log, which is a file write operation that may store application output or sensitive debugging data. The file write has no confirmation prompt, warning comment/docstring, or other disclosure explaining that captured browser console data will be saved locally.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code creates an output directory and later saves page screenshots and collected console logs there. Although it prints the file paths after writing, there is no prior confirmation prompt, warning comment, or user-facing disclosure that browsing artifacts from the target page will be persisted locally.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This Python file contains user-facing natural-language text exclusively in Chinese, including the module docstring and command-line help/output strings. That imposes a language choice on users without offering an opt-in, alternative locale, or documented justification for a Chinese-only interface.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This helper script accepts arbitrary --server commands and then executes them, which means it can be used to run any local shell command rather than being constrained to Playwright-related service startup. In a skill context, that broad execution capability is dangerous because untrusted inputs or downstream agent misuse could lead to arbitrary code execution on the host running the skill.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The printed status messages are hardcoded in Chinese, which imposes a specific language on users without opt-in or explanation. This is a natural-language locale policy concern because the skill does not offer language selection or document that it is intended only for a Chinese-speaking context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The script prints status and element labels in Chinese (e.g. '共发现', '[隐藏]', '[无法获取]') with no option for the user to select another language. This is a natural-language locale policy concern because the skill imposes a specific language by default rather than offering opt-in or documenting a justified regional scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The printed completion message is only in Chinese, which imposes a specific language on users without any opt-in or indication that this skill is region-specific. This is a natural-language policy concern because the file provides no alternative language handling or justification for the locale restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

All user-visible status, error, and usage messages in this script are written in Chinese, with no mechanism for locale selection or opt-in. This can violate language/locale policy where skills should not force a specific language on users by default.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.