Back to skill

Security audit

Web Qa Bot

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent web QA tool, but its implementation contains unsafe command execution paths that could let crafted test inputs run local system commands.

Review before installing. Only run this on trusted test suites and staging or test accounts until the shell command construction is fixed, because malicious URLs, YAML/JSON steps, report paths, or company names could execute local commands. Treat generated screenshots, console logs, and reports as sensitive, and avoid PDF output unless ai-pdf-builder is reviewed, declared, and pinned.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
src/browser.ts:46
Finding

OS Command Injection in Browser Automation Command Construction

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
src/reporter.ts:164
Finding

OS Command Injection in PDF Report Generation

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
src/reporter.ts:173
Finding

Runtime Execution of an Undeclared and Unpinned Package Through npx

Content
View full analysis
Remediation
View remediation
", "yaml": "^2.3.4" } } ``` 2. Commit the resulting lockfile and verify the package integrity metadata. 3. Review the selected package version, its install scripts, executable entry point, maintainers, publication history, and transitive dependencies before adoption. 4. Invoke the installed local binary directly rather than allowing `npx` to download missing packages. If `npx` remains necessary, use a mode that refuses installation and fails when the local executable is absent. 5. Prefer a programmatic API where available, while still pinning and reviewing the dependency. 6. Use automated dependency scanning and periodic lockfile review. 7. In CI, install dependencies using lockfile-enforcing commands and run report generation with minimal filesystem, secret, and network access. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (24)

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The supplied code does not perform automated QA tasks such as smoke tests, accessibility checks, or visual regression analysis. Instead, it implements helper utilities for waiting, retrying, and determining page stability/loading state. While such utilities could support a QA system, this chunk’s actual behavior is materially narrower and different from the declared purpose. Therefore, the description does not accurately represent what this specific code chunk actually does.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: ws==8.19.0 — 2 advisory(ies): CVE-2026-45736 (ws: Uninitialized memory disclosure); CVE-2026-48779 (ws: Memory exhaustion DoS from tiny fragments and data chunks)

High
Category
Supply Chain
Confidence
97% confidence
Finding

The lockfile pins ws to 8.19.0, and the reported advisories describe memory disclosure and memory-exhaustion denial of service in the WebSocket library. Because this skill depends on agent-browser, which in turn uses ws for browser/automation connectivity, the vulnerable package is plausibly reachable during normal operation and could allow a malicious or malformed peer/server to crash the process or expose memory.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The plan includes capturing console logs, screenshots, and embedding them in reports without warning that these artifacts may contain sensitive application data, credentials, tokens, or personal information. In QA contexts, recorded artifacts are often shared or stored, which increases the chance of unintended data exposure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The documented browser testing workflow describes clicks, typing, navigation, and interactive mode without warning that it may perform real state-changing actions on live sites. In this skill context, that can cause unintended form submissions, data modification, purchases, account actions, or test contamination if run against production systems.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

The plan instructs users to invoke the tool via npx web-qa-bot without pinning a specific version. This allows execution of whatever package version is currently resolved from the registry, which can lead to unexpected or malicious code execution if a compromised or breaking release is published.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

Using npx web-qa-bot without a pinned version creates a supply-chain risk because npx may fetch and execute the latest package version at runtime. If the package is hijacked, typo-squatted, or a malicious update is published, users could run attacker-controlled code on their systems.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

The command example relies on an unpinned package resolution path, which undermines reproducibility and permits execution drift over time. In a security-sensitive automation tool that interacts with browsers and local files, a malicious update could abuse those permissions for code execution or data access.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

An interactive command run via unpinned npx is especially risky because it typically grants broad local access and user trust during exploratory sessions. A compromised package version could exfiltrate credentials, inspect browser state, or alter local files under the guise of QA tooling.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

The success criteria reinforce unpinned npx execution, normalizing insecure installation and runtime behavior across adopters. Because this is a browser automation skill, a malicious package update could quickly affect many users with local code execution and access to captured test artifacts.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The exec() helper constructs a shell command by concatenating untrusted arguments into a single string and passes it to execSync(), which invokes a shell. Inputs such as URL, ref, key, script, and screenshot path can contain shell metacharacters, enabling command injection and host-level arbitrary command execution if attacker-controlled data reaches these methods.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest describes a web QA skill focused on automated testing activities, but this file also provides a general-purpose evaluate method that executes caller-supplied JavaScript in the page. While browser automation is expected for QA, arbitrary script execution materially expands the capability beyond the specifically stated testing functions and enables actions unrelated to those declared purposes.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The evaluate() method forwards arbitrary script content into the browser automation subprocess through the same string-built shell execution path. Because only double quotes are escaped, shell metacharacters like $(...), backticks, semicolons, or single quotes may still break out of the intended argument and achieve host command injection, making this more dangerous than mere page-context script evaluation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The reporter writes assembled markdown output directly to the user-specified path, and that output includes URLs, test errors, console errors, and optional screenshot references. While the code comments describe report generation, there is no user-facing warning, confirmation, or disclosure that potentially sensitive QA data will be persisted to disk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The code builds a shell command with interpolated mdPath, options.output, and company values and passes it to execSync, which invokes a shell. If any of those fields are attacker-controlled, shell metacharacters can break out of quoting and achieve arbitrary command execution; this is especially dangerous in an automation tool that may consume external input.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
97% confidence
Finding

The PDF generation path invokes npx ai-pdf-builder without pinning a version, so execution depends on whatever package version is resolved at runtime. In a tool that processes user-controlled report content, this creates a real supply-chain risk and can lead to unexpected code execution if a malicious or compromised package version is installed or fetched.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The JSON report writer serializes the full this.results structure and writes it to disk, which may include detailed test metadata, URLs, console events, errors, and screenshot paths. There is no user-facing disclosure that this potentially sensitive execution data will be stored in a file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This markdown file includes commands that generate report artifacts such as report.md and report.pdf, but it does not explicitly warn users that running these commands writes files to disk and may overwrite existing paths. For markdown files, SQP-2 applies when user-facing descriptions omit warnings about behaviors that can affect user data or system state.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: yaml==2.8.2 — 1 advisory(ies): CVE-2026-33532 (yaml is vulnerable to Stack Overflow via deeply nested YAML collections)

Low
Category
Supply Chain
Confidence
90% confidence
Finding

The lockfile includes yaml 2.8.2, which is reported vulnerable to stack overflow from deeply nested YAML collections. Since this skill directly depends on yaml and likely consumes configuration or test definitions, untrusted or malformed YAML input could crash the process, though the impact is primarily denial of service rather than code execution.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This is a manifest file, so vague-trigger checks apply. The description says the skill is a "Drop-in for manual QA" and "Works with Cursor, Claude, ChatGPT, Copilot" but does not define any explicit trigger phrases, invocation boundaries, or exclusion conditions, which could make activation scope ambiguous in systems that infer routing from manifest text.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.

Content

Scanner excerpt · package.json (reported line 64)May include surrounding context.

json
"node": ">=18.0.0"
  },
  "dependencies": {
    "yaml": "^2.3.4"
  },
  "devDependencies": {
    "@types/node": "^20.10.0",

Known Vulnerable Dependency: yaml==2.8.2 — 1 advisory(ies): CVE-2026-33532 (yaml is vulnerable to Stack Overflow via deeply nested YAML collections)

Low
Category
Supply Chain
Confidence
92% confidence
Finding

The package depends on yaml with a version range that may resolve to a release affected by a stack overflow denial-of-service issue in deeply nested YAML parsing. For an AI/automation skill that may ingest user-controlled configuration or test inputs, this can let an attacker crash the process or disrupt automated workflows if malicious YAML is supplied.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.

Content

Scanner excerpt · package.json (reported line 67)May include surrounding context.

json
"yaml": "^2.3.4"
  },
  "devDependencies": {
    "@types/node": "^20.10.0",
    "typescript": "^5.3.0"
  },
  "peerDependencies": {

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.

Content

Scanner excerpt · package.json (reported line 68)May include surrounding context.

json
},
  "devDependencies": {
    "@types/node": "^20.10.0",
    "typescript": "^5.3.0"
  },
  "peerDependencies": {
    "agent-browser": ">=0.7.0"

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The screenshot method creates directories and writes screenshot files to disk, which affects the local filesystem. While the code comments describe the implementation, there is no user-visible prompt, log, or warning indicating that files will be created under the configured screenshot directory.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/browser.ts:46

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/reporter.ts:173