Back to skill

Security audit

Humanizer Rewrite

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed AI-text humanizer, but it explicitly optimizes for hiding AI authorship signals and can overwrite user files without confirmation.

Install only if you intentionally want a Chinese text-rewriting skill that changes style toward human-like prose. Avoid using it to misrepresent authorship, and do not give it file paths unless you are comfortable with the original file being replaced; a safer version should return rewritten text or require explicit confirmation and backups before writing.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (6)

Ssd 2

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill is explicitly designed to make AI-generated text appear human-written by removing detectable traits and injecting deceptive stylistic signals. In context, this is dangerous because it facilitates evasion of authorship disclosure, academic or editorial integrity checks, and downstream trust mechanisms that rely on honest provenance.

Content

No source excerpt is available for this finding.

Ssd 4

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The multi-round rewrite loop systematically scores and intensifies techniques for human-authorship mimicry, including deliberate inconsistency, emotional injection, and anti-pattern suppression. This iterative optimization makes the deception more effective than ordinary style editing and increases the likelihood of misuse for fraud, plagiarism concealment, or policy evasion.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill tells the agent to read a user-supplied path and overwrite the same file without warning or confirmation. This is high risk because a single ambiguous prompt could lead to silent destructive changes to local content, and the user is not given an approval point to verify the target or the resulting edits.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instruction '全程简体中文' mandates a specific language for all interactions. This is a locale/language restriction without any user opt-in or documented justification, which matches the policy-violation category for forced language behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest describes a rewriting skill, but the body extends behavior to reading files and overwriting the original path in place. This capability expansion is dangerous because users invoking a text-editing skill may not expect destructive file operations, increasing the chance of accidental data loss or misuse through path-based prompts.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instruction to overwrite the provided file lacks a clear necessity boundary, safety checks, or backup flow. Even without malicious intent, direct overwrite creates a real risk of irrecoverable content corruption, especially when the skill's task can be completed by returning transformed text instead.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.