Back to skill

Security audit

Clean Text Toolkit

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent local text toolkit, but its HTML sanitization and Markdown-to-HTML features can produce unsafe HTML despite being presented as suitable for untrusted input.

Install only if you need local text-processing utilities and are comfortable reviewing outputs. Do not rely on the HTML sanitization or Markdown HTML rendering features to make untrusted content safe for a browser or web app; use a maintained allowlist sanitizer for that. Also choose output paths carefully because the scripts write or overwrite files you specify.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/htmlstrip.py:177
Finding

Incomplete HTML Sanitization Allows Executable Content

Content
View full analysis
") def handle_endtag(self, tag): tag = tag.lower() if tag in self.strip_tags: self.drop_depth = max(0, self.drop_depth - 1) return self.out.append(f"") def handle_startendtag(self, tag, attrs): tag = tag.lower() if tag in self.strip_tags: return clean_attrs = [(k, v) for k, v in attrs if not k.lower().startswith("on") and k.lower() != "style"] attr_str = "".join( f' {k}="{html.escape(v, quote=True)}"' if v is not None else f" {k}" for k, v in clean_attrs ) self.out.append(f"<{tag}{attr_str} />") ``` ### Technical Analysis The HTML mode uses a denylist that removes selected tags, attributes beginning with `on`, and inline `style` attributes. All other elements and attributes are preserved. Attribute-value HTML escaping prevents direct quote-based attribute breakout, but it does not make URL-bearing attributes safe. In particular, values of attributes such as `href`, `src`, `action`, `formaction`, and `xlink:href` are not checked against an approved scheme list. A dangerous value such as `javascript:alert(1)` therefore survives sanitization. The sanitizer also permits arbitrary non-denylisted elements, including potentially active SVG, MathML, `meta`, and `base` elem ...[truncated 1621 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/markdown.py:132
Finding

Markdown HTML Renderer Permits Attribute Injection and Dangerous URLs

Content
View full analysis
str: # Escape raw HTML inside lines (very basic) s = s.replace("&", "&").replace("<", "<").replace(">", ">") # Restore inline code first to protect it from further markdown # (do a simple two-pass: find spans, replace with sentinel, run others, restore) placeholders: Dict[str, str] = {} def stash_code(m): key = f"\x00CODE{len(placeholders)}\x00" placeholders[key] = f"{m.group(1)}" return key s = re.sub(r"`([^`]+)`", stash_code, s) s = re.sub(r"!\[([^\]]*)\]\(([^)]+)\)", lambda m: f'{m.group(1)}', s) s = re.sub(r"(?{m.group(1)}', s) s = re.sub(r"\*\*([^*]+)\*\*", r"\1", s) s = re.sub(r"__([^_]+)__", r"\1", s) s = re.sub(r"(?\1", s) s = re.sub(r"(?\1", s) for k, v in placeholders.items(): s = s.replace(k, v) return s while i < len(lines): line = lines[i] # Fenced code blocks m = re.match(r"^```(.*)$", line) if m and not in_code: close_list(); close_blockquote() in_code = True code_lang = m.group(1).strip() lang = f' class="language-{code_lang}"' if code_lang else "" out.append(f"
text
")
        i += 1
        continue
```

### Technical Analysis

The renderer escapes ampersands and angle brackets before applying inline Markdown substitutions, but it does not escape quotation marks before inserting attacker-controlled values into HTML attributes.

The following values are directly interpolated into quoted attribut
...[truncated 2209 chars]
Remediation
View remediation
Vulnerability Patterns
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (8)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The supplied code does not match the declared description of a comprehensive text-cleanup toolkit. Instead, it implements a specific HTML utility with three modes: strip HTML to text, sanitize HTML by removing configured tags and event/style attributes, and extract certain HTML structures (links, images, headings, tables). This is materially narrower and different in primary purpose than the declared toolkit, and it also adds an undeclared HTML sanitization capability. The only points of alignment are that it is local, stdlib-only, and makes no remote calls. Overall, the description substantially overstates unrelated capabilities that are absent from this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code chunk is specifically a CLI tool for replacing text in files. It accepts input/output paths, parses replacement rules, compiles regex or literal patterns, previews matches, counts replacements, and writes edited output. This is not a supporting subcomponent of the declared toolkit as described; it is a different primary function that is undeclared. At the same time, none of the major declared capabilities in the description are present in this code chunk: no extraction of URLs/emails/phones/etc., no PII redaction set, no normalization utilities, no line utilities, no word-frequency stats, no diff modes, no template rendering, no slug generation, and no Markdown conversion. Resource usage remains local and there are no remote calls, which is consistent, but the behavior still materially mismatches the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code does not exhibit malicious or undeclared external access; it is local-only and uses standard-library modules, consistent with that portion of the description. However, the declared purpose describes a large multifunction text toolkit, including extraction of URLs/emails/etc., PII redaction, normalization, line utilities, word-frequency stats, diffs, slug generation, and Markdown conversion. This code chunk implements only one component: a no-Jinja2 template renderer with placeholder syntaxes, filters, defaults, JSON/inline variable sources, and strict mode. Because the declared description materially overstates what this code chunk actually does and its primary behavior is much narrower, this should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 27)May include surrounding context.

md
- `scripts/replace.py` (NEW in v0.3.0) — find-and-replace with regex / literal / word-boundary modes, capture-group back-references (`\1`, `\2`), multiple `--fi

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 184)May include surrounding context.

md
- `scripts/replace.py` (NEW in v0.3.0) — find-and-replace with regex / literal / word-boundary modes, capture-group back-references (`\1`, `\2`), multiple `--fi

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This code file contains helpers that create directories and write files, which are safety-relevant operations under the policy. Although the functions are straightforward, they provide no confirmation prompt, logging, or user-facing disclosure at the point of write, so callers can modify filesystem state without any visible warning from this module.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code writes transformed content to the user-supplied output path, which can overwrite an existing file and alter user data. While the module docstring explains usage, it does not explicitly warn that the output file will be written or potentially replaced, and there is no confirmation prompt at write time.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.