Back to skill

Security audit

Humanizer

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent writing-editing skill, but users should be careful with its automatic project context file and in-place edit mode.

Install only if you want an opinionated prose humanizer with local file-edit capability. Check any humanizer-context.md in a project before using the skill, and use detect or rewrite mode first when working on important files so you can review changes before applying them.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:66
Finding

Automatic Trust of Project-Controlled Context Enables Skill Instruction Hijacking

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 66 and 101
Vulnerability Type: Untrusted local context promoted to agent instructions
Risk Level: Medium

Complete Vulnerable Snippet

markdown
Auto-loads `humanizer-context.md` from the project root if present. Use that file for brand samples and banned phrases.
markdown
**Auto-load brand context.** Before parsing further, check for `humanizer-context.md` in the current working directory using the Read tool. If it exists, load it as additional voice guidance (brand samples, banned phrases, preferred terms). Treat its contents as a personal extension of the `--voice` profile. If it doesn't exist, proceed without warning; this is opt-in.

Technical Analysis

The skill instructs the agent to load humanizer-context.md automatically from the current project and treat its contents as an extension of the active voice profile. A project repository can control this file, but the skill provides no schema validation, content restrictions, trust-boundary warning, or instruction filtering.

Consequently, a malicious repository can place arbitrary natural-language instructions in this context file. When the skill runs, those instructions enter the agent's working context and may influence its rewriting behavior. For example, the file could demand insertion of attacker-selected text, removal or distortion of source content, disclosure of supplied document data, or attempts to invoke unrelated tools.

The phrase “this is opt-in” does not provide effective consent because the documented behavior proceeds automatically and silently when the file exists. Higher-priority system and platform policies remain authoritative, so this issue does not independently establish unrestricted code execution or privilege escalation. It does, however, create a direct indirect-prompt-injection path within the skill's legitimate processing flow.

Attack Path

  1. An attacker prepares or modifies a re ...[truncated 1366 chars]
Remediation
View remediation

Remediation Suggestions

  1. Do not load humanizer-context.md silently. Require explicit user confirmation before reading project-provided context.
  2. Treat the file as untrusted data rather than executable guidance. State explicitly that instructions, tool requests, policy overrides, and requests to read other files must be ignored.
  3. Replace free-form Markdown with a constrained format containing approved fields such as:
    • preferred_terms
    • banned_phrases
    • voice_examples
    • tone_attributes
  4. Validate field types, lengths, and allowed values before incorporating them into the prompt.
  5. Delimit imported content clearly and ensure it cannot redefine the skill's workflow, tool permissions, output destination, or safety constraints.
  6. Reject or escape content containing instruction-like directives, references to system prompts, requests for tool use, or attempts to override existing rules.
  7. Display the loaded path and a short summary to the user, and allow the user to decline its use.
  8. In edit mode, show a diff or request confirmation before writing changes influenced by project-provided context.
  9. Restrict context discovery to a known project root rather than relying on an ambiguous current working directory, and prevent traversal or symlink-based substitution where the host environment permits such checks.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (3)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · SKILL.md (reported line 101)May include surrounding context.

md
- `technical`: code blocks preserved, precise jargon retained, numbers over adjectives
  - `general`: no purpose-specific overrides (default)

**Auto-load brand context.** Before parsing further, check for `humanizer-context.md` in the current working directory using the Read tool. If it exists, load it as additional voice guidance (brand samples, banned phrases, preferred terms). Treat its contents as a personal extension of the `--voice` profile. If it doesn't exist, proceed without warning; this is opt-in.

Store parsed values. Proceed to Step 2.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill supports --mode edit with Write and Edit permissions and explicitly performs in-place file modification, but the workflow does not require an explicit confirmation immediately before changing a file. That creates a real integrity risk: a user may invoke the skill expecting analysis or a draft rewrite and instead have a local file altered, especially since rewrite and edit are adjacent modes and --file can change the source of truth.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The pattern catalog treats certain language conventions as universal AI tells, including a 'zero tolerance' ban on em dashes and marking consistent Oxford comma use or curly quotes as fingerprints. These are legitimate stylistic and locale-dependent choices, so presenting them as blanket defects imposes a specific writing standard without offering user choice or contextual justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.