Back to skill

Security audit

ai-fingerprint-desensitizer

Security checks for vulnerabilities and agentic risk

Overview

This skill is locally scoped and disclosed, but its main workflow is to reduce AI-detection signals for platform publication, so users should review it carefully before installing.

Install only if you are comfortable with a tool whose primary function is to rewrite content to reduce AI-detection signals. Use it only on your own or authorized material, preserve required AI disclosures, and do not use it to hide AI authorship, bypass platform rules, or fabricate personal details, data, interviews, or experiences.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (13)

Ssd 2

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This section operationalizes evasion through a measurable feedback loop using a fingerprint comparison tool and specific score-drop targets. That materially increases misuse potential by helping users iteratively tune content until automated detection is less effective.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The self-check defines success as lowering a fingerprint score by a target amount and iterating until detectability decreases. This is a direct operational loop for evasion, enabling systematic optimization against AI-detection mechanisms despite the stated claim that the tool is not for bypassing checks.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger list contains broad phrases such as '改写', 'AI生成内容', and '内容原创度' that can match many ordinary writing-assistance requests, causing the skill to activate in contexts beyond narrowly scoped fingerprint analysis. In this skill, that overbroad activation is more concerning because the core workflow explicitly helps users reduce detectable AI signals and '过检测', which can facilitate evasion of platform disclosure or moderation mechanisms even if the text also includes compliance language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The entire rule set is framed specifically around Chinese writing characteristics, using statements such as 'AI 写中文句子' and prescribing what qualifies as acceptable '人味' in Chinese prose. There is no indication that this language-specific standard is optional, user-selected, or limited to a clearly documented Chinese-only context, which creates a potential language/locale policy issue.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The document operationalizes detection and rewrite tactics to make AI-generated text appear human by manipulating semantic and stylistic signals used by platforms. Even though it includes compliance framing, the concrete guidance focuses on reducing detectability ('去AI味','过检测','降AI率') and iterating until the content falls below suspicion thresholds, which materially enables policy evasion and deceptive provenance masking.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown file presents all instructions and policy guidance exclusively in Chinese, and there is no indication that the skill is region-specific or that users can opt into another language. That creates a natural-language policy concern under the language/locale rule because the file effectively forces a specific language without documented justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The title and all instructional content are written as prescriptive guidance entirely in Chinese, with no indication that language selection is optional or limited to a China-specific workflow. Under the policy, forcing a specific language or locale without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The playbook explicitly frames rewriting as a way to defeat AI-style fingerprint scoring, not merely improve clarity or originality. Even though it uses compliance language elsewhere, the operational guidance is detector-oriented evasion and can be used to conceal AI-generated or derivative content from platform review systems.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The document semantically teaches users how to alter structure and wording to avoid AI-generated-text detection, rather than simply edit for better communication. In context, this increases the capability to disguise synthetic or repurposed content and undermines platform integrity controls.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The advice to add numbers, scenes, regret, and errors as 'human signals' is detector-targeted and can encourage fabricated details that make generated text appear authentic. Even with a later warning not to invent facts, the framing teaches an evasion strategy that can be misused to create deceptive content.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

These lines instruct users to insert cues described as proof of being human, such as regret, mistakes, and hindsight, to manipulate classifier signals. That goes beyond normal revision advice and teaches users how to fabricate authenticity markers that may mislead moderators, platforms, or readers.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code file contains natural-language instructions and outputs entirely in Chinese, including the docstring usage text and report labels, with no indication that the user can select another language. Under the policy for natural-language violations, forcing a specific language without user opt-in is a reportable issue unless the locale restriction is explicitly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

This markdown file contains user-facing instructions exclusively in Chinese, which can amount to a language/locale policy issue when no opt-in, alternative language, or justification is provided. The policy specifically calls for flagging skills that force a specific language without user choice.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.