Back to skill

Security audit

i-can-see

Security checks for vulnerabilities and agentic risk

Overview

Review recommended: this skill is openly a local camera-capture skill, but it uses broad triggers and an unauthenticated hardcoded camera endpoint to save real-world images.

Install only if you knowingly want OpenClaw to take photos from the specified ESP32-CAM. Narrow the trigger wording, require confirmation before each capture, secure or isolate the camera network, validate image responses, pin dependencies, and decide where saved photos should be stored and deleted.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Note
Location
SKILL.md:13
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 13-17
Vulnerability Type: Unpinned third-party package dependency
Risk Level: Low

Vulnerable Code:

bash
pip install requests

Technical Analysis

The installation instructions retrieve the latest available version of requests without a version constraint, lockfile, integrity hash, or explicit trusted package index. Consequently, the installed code can change after the skill has been reviewed.

This creates supply-chain exposure if the configured package repository, package release, dependency resolution process, or local package-index configuration is compromised. Although the referenced package is correctly named and widely used, the installation procedure does not provide reproducibility or artifact integrity verification.

Attack Path

  1. An attacker compromises a relevant package release, transitive dependency, package repository, or package-index configuration.
  2. A user follows the documented pip install requests instruction.
  3. pip resolves and installs the attacker-controlled or compromised package version.
  4. Malicious package code executes during installation or when capture.py imports and uses the package.

Impact Assessment

Successful exploitation could execute code with the privileges of the user running pip or capture.py. Depending on those privileges, this may permit access to that user's files, credentials, network resources, and application data. The scope is limited by the executing user's operating-system permissions; the audited project itself does not request elevated privileges.

Remediation
View remediation

Remediation Suggestions

  • Pin requests and its transitive dependencies to reviewed versions in a requirements or lock file.
  • Use hash verification, for example pip install --require-hashes -r requirements.txt.
  • Generate and review dependency hashes using an established dependency-management workflow.
  • Explicitly configure a trusted package index rather than relying on ambient pip configuration.
  • Periodically scan and update pinned dependencies after security review.
  • Install dependencies in an isolated virtual environment with only the privileges required by the skill.

T09 · Insecure Skill Coding Practices

Warning
Location
capture.py:7
Finding

Unauthenticated Plaintext Transport for Camera Images

Content
View full analysis

Vulnerability Details

File Location: capture.py, lines 7-10; endpoint also documented in SKILL.md, line 61
Vulnerability Type: Plaintext, unauthenticated transmission of potentially sensitive camera data
Risk Level: Medium

Vulnerable Code:

python
def capture_image(output_path=None):
    url = "http://192.168.31.241/capture"
    try:
        print(f"Connecting to ESP32-CAM at {url}...")
        res = requests.get(url, timeout=10)

The response is subsequently accepted based only on its status and non-empty content:

python
if res.status_code == 200 and res.content:
    if not output_path:
        ts = time.strftime("%Y%m%d_%H%M%S")
        out_dir = os.path.dirname(os.path.abspath(__file__))
        output_path = os.path.join(out_dir, f"capture_{ts}.jpg")
    else:
        output_path = os.path.abspath(output_path)

    os.makedirs(os.path.dirname(output_path), exist_ok=True)

    with open(output_path, "wb") as f:
        f.write(res.content)

Technical Analysis

The camera is accessed over HTTP, which provides neither transport confidentiality nor cryptographic endpoint authentication. An attacker with an appropriate local-network position may observe the camera response or impersonate the configured endpoint.

The script also does not authenticate its request or validate the response's Content-Type, image signature, dimensions, or maximum size before writing it as a JPEG file. A forged endpoint can therefore return arbitrary bytes with HTTP status 200, and the script will report the operation as successful. This can cause the Agent to analyze attacker-supplied content as though it came from the physical camera.

Attack Path

  1. An attacker gains a suitable position on the same or an intermediary network, such as through a compromised router, malicious access point, ARP spoofing, or control of the configured camera address.
  2. The user or Agent invokes ` ...[truncated 1146 chars]
Remediation
View remediation

Remediation Suggestions

  • Place the camera behind HTTPS with certificate verification, or access it through an authenticated TLS-secured proxy if the device cannot provide HTTPS directly.
  • Authenticate camera requests using device-specific credentials or signed, short-lived tokens.
  • Avoid disabling certificate verification and restrict trust to the expected certificate or internal certificate authority.
  • Validate the response Content-Type against an allowlist such as image/jpeg.
  • Verify JPEG magic bytes and decode the image with a maintained image library before treating the capture as successful.
  • Stream the response with a strict maximum byte limit to prevent oversized-response storage or memory abuse.
  • Restrict network access to the camera through firewall rules and an isolated IoT network.
  • Return a nonzero exit status on failed validation so downstream automation cannot mistake invalid content for a successful capture.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (6)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The documented behavior goes beyond a simple descriptive 'vision' skill by directing network capture from a hardcoded local camera endpoint and storing images locally, while not declaring those capabilities or constraints. That mismatch is dangerous because reviewers and users may underestimate that the skill can acquire real-world images and persist them, enabling unintended surveillance, privacy exposure, or unauthorized data collection.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill instructs the agent to use network access via an ESP32-CAM endpoint and write captured images to disk, but it declares no explicit tool scope or permissions. This creates an authorization gap where a sensitive capability—capturing real-world imagery over the network—can be invoked without clear consent, policy gating, or review boundaries.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The activation phrases are broad everyday language such as requests to 'look' or 'see what you see,' which can cause the skill to trigger in contexts where the user did not intend real-world image capture. Because this skill interfaces with a physical camera, accidental activation materially increases privacy risk compared with an ordinary text-only skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill does not clearly warn users that it captures images of the real world and fetches them from a device over the local network. Missing disclosure undermines informed consent for a privacy-sensitive action and can expose bystanders, screens, documents, or other sensitive environmental data without the user's full awareness.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The manifest description and operational instructions are written entirely in Chinese, including the mandated response behavior in step 5, without indicating that users may choose another language. This can violate a language/locale policy when no opt-in or region-specific justification is provided.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The example triggers reinforce ambiguous invocation boundaries and normalize automatic camera capture from casual phrases. In the context of a vision skill connected to a live local device, unclear boundaries can lead to unauthorized photo capture and analysis of the user's physical environment.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.