Back to skill

Security audit

Distil Open Claw Pii

Security checks for vulnerabilities and agentic risk

Overview

Needs review: this is a local PII redaction skill, but its own example leaves a full card number visible and the implementation lacks safeguards when redaction fails.

Install only if you are comfortable reviewing redaction results yourself before sharing them. Do not treat the output as guaranteed sanitized for payment data or regulated PII, avoid --show-entities unless you truly need original-value mappings, and consider pinning and verifying the model before setup.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:34
Finding

Unsafe Shell Interpolation Guidance Enables Command Injection

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/redact.py:74
Finding

Prompt-Delimiter Injection Can Bypass PII Redaction

Content
View full analysis
{text} Redact provided text according to the task description and return redacted elements.""" ``` ```python payload = json.dumps({ "model": "distil-pii", "messages": [ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": USER_PROMPT_TEMPLATE.format(text=text)}, ], "temperature": 0, "response_format": response_format, }).encode() ``` ```python content = result["choices"][0]["message"]["content"] if show_entities: print(content) else: parsed = json.loads(content) print(parsed["redacted_text"]) ``` ### Technical Analysis Untrusted input is inserted directly into an instruction-bearing XML-like prompt using `USER_PROMPT_TEMPLATE.format(text=text)`. The input is not escaped or otherwise isolated from the `` and `` delimiters. An attacker can supply text containing markup such as `` followed by additional model instructions. The local model may interpret those instructions as part of the controlling prompt rather than as text requiring redaction. Although the response is constrained to JSON, the schema only controls its structure. It does not verify that `redacted_text` no longer contains sensitive values. The script consequently trusts and prints model-generated output without deterministic redaction validation. ### Attack Path 1. An attacker submits text containing PII and a forged closing `` delimiter. 2. The text adds instructions asking the model to preserve or reproduce the sensitive values. 3. The application ...[truncated 697 chars]
Remediation
View remediation
`, ``, and instruction-like strings. ]]>

T09 · Insecure Skill Coding Practices

Error
Location
examples/sample_output.json:2
Finding

Documented Redaction Output Retains a Full Payment Card Number

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
scripts/setup.sh:5
Finding

Model Artifact Is Downloaded from a Mutable Source Without Integrity Verification

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (8)

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The supposedly redacted text still contains a full credit card number (4111 1111 1111 1234), which is a direct failure of redaction for extremely sensitive financial data. In the context of a PII-redaction skill, this is more dangerous because users will reasonably trust the output as safe to store, transmit, or review, potentially causing PCI exposure and financial fraud risk.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
85% confidence
Finding

The skill uses shell and network-capable operations (bash scripts/setup.sh, curl, model download, and a local server) but does not declare an explicit tool scope such as permissions or allowed-tools. This creates an avoidable trust gap: an agent may invoke broader capabilities than intended during setup or execution, increasing the chance of unintended command execution or network access in a privacy-sensitive skill handling raw PII.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The sample output includes raw sensitive values in the entities[].value fields, which undermines the claim that PII is redacted locally and can normalize unsafe handling of highly sensitive data in logs, docs, or downstream integrations. Because this skill is specifically about PII redaction, exposing original values in example output is especially dangerous: developers may copy this schema and accidentally retain or emit the very data that should be removed.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill promises that text never leaves the machine, but it sends all input text to an HTTP service at localhost:8712. Even though this is loopback traffic, the data leaves the calling process boundary and is exposed to whatever is listening on that port, so the privacy guarantee in the skill description is materially overstated. In a PII-redaction skill, this mismatch is more dangerous because users are specifically encouraged to submit highly sensitive data under a strong local-processing claim.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest explicitly says the skill works locally and that text never leaves the machine, which conveys a strongly local-only posture. This setup script performs an external network fetch to download the model from Hugging Face, so the skill is not fully local in operation even if inference later runs locally.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The quick-start instruction says 'set up the PII redactor' and notes that it downloads the model and starts a local server, but it does not clearly warn the user that this performs a network download and installs about 5 GB of model data locally. For a markdown skill description, user-facing warnings should disclose behaviors that affect the user's system or resources.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The top-level documentation states that the script prints the redacted JSON to stdout, but the implementation only prints full JSON when --show-entities is supplied. By default, it parses the model response and prints only the redacted_text field, so the docstring actively misdescribes observable behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This code starts llama-server in the background and persists its PID to a file, which changes the user's local runtime state beyond the script's immediate execution. Although there is a status message, it does not clearly warn that a persistent background service will continue running after setup completes.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.