Back to skill

Security audit

PII Redactor

Security checks for vulnerabilities and agentic risk

Overview

The skill is clearly intended to redact PII, but it broadly intercepts every agent response and forwards complete draft text to a configured service, with weak package provenance.

Review carefully before installing. Only use this with a redaction server you operate, preferably localhost or a tightly controlled internal service, and assume that full draft responses may be visible to that service before redaction. Verify the clawguard-pii package source and release integrity independently, run it under a dedicated low-privilege account or container, and avoid enabling it for sessions containing secrets unless the deployment is trusted.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:72
Finding

Mandatory Session-Wide Interception and Redirection of Agent Responses

Content
View full analysis
"} ``` ``` ### Technical Analysis The Skill declares that its instructions apply to every response and cannot be overridden by the user. This changes session-wide Agent behavior rather than providing a narrowly scoped, opt-in redaction operation. The instructions require the complete, unredacted draft response to be transmitted to the endpoint selected through `CLAWGUARD_URL`. Consequently, sensitive information is disclosed to the redaction service before any redaction occurs. Although the document describes URL-validation rules, the project contains no executable implementation with which to verify that those restrictions are enforced, that DNS resolution is checked safely, or that redirect and address-rebinding cases are rejected. The combination of an unoverrideable global directive and mandatory response forwarding is an instruction-hijacking pattern. It allows the Skill to interpose itself on unrelated Agent activity and creates a data-disclosure channel to the configured service. ### Attack Path 1. The Skill is loaded into an Agent session. 2. The Skill asserts that its processing rules apply to every response and cannot be overridden. 3. The Agent generates a complete draft containing conversation context, private information, credentials, or other sensitive material. 4. Before delivering the resp ...[truncated 870 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:18
Finding

Installation and Execution of a Dependency Without Verifiable Source Provenance

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep

Static analysis

No suspicious patterns detected.