Back to skill

Security audit

anti-ai-foolish

Security checks for vulnerabilities and agentic risk

Overview

The skill is not malware, but it is explicitly designed to help Chinese articles pass AI-detection and platform review, so it belongs in Review before installation.

Install only if you intentionally want a Chinese-language pre-publication detector-evasion workflow. Do not use it to disguise AI-generated work from platforms, publishers, readers, or reviewers that require disclosure. If used for legitimate editing, keep it limited to transparency, factual accuracy, and style quality, and avoid the external detector pass/fail loop.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (52)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a user-facing content gatekeeper that checks an article and removes or reduces AI-style signals before publishing. The supplied code does not do that. It does not transform article text, provide final article checks, or operate on user requests like 去AI味/终检. Instead, it is a maintenance/analysis script for validating the effectiveness of existing regex rules on a preloaded Zhuque-labeled corpus of 52+ fragments, computing confidence intervals and effect sizes, then persisting validation statuses back into the rule base and a results file. While this supports the broader system described, the code chunk’s primary purpose is materially different from the declared skill behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The code substantially aligns with the detection/audit portion of the description: it scans Chinese article text using a ruleset, identifies AI-like patterns, tracks 'human flavor' hits, computes a score, and produces a report for review before publishing. However, the description overstates behavior by saying it 'detects and removes' AI-writing tells. This implementation does not edit the article; it only reports matches and fix suggestions. The trigger language and platform-flagged use cases are also not implemented in the code itself, though that could be orchestration outside this file. Finally, the specific validation and rule-count claims are not verifiable from this chunk. The main material mismatch is the claimed removal capability versus actual scan/report-only behavior.

Content

No source excerpt is available for this finding.

Ssd 4

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The workflow gives iterative, step-by-step guidance to modify content, retest it, and continue until it passes an external detector threshold. This materially increases misuse potential by providing a repeatable evasion playbook, not just general writing advice, making the skill more dangerous in context than a single isolated suggestion would be.

Content

No source excerpt is available for this finding.

Ssd 2

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

These instructions prescribe concrete detector-evasion tactics such as removing signals, altering structure, and optimizing toward a detector score target. Because the advice is semantically tuned to suppress detection features rather than to improve truthfulness or readability, it facilitates circumvention of safety and authenticity controls.

Content

No source excerpt is available for this finding.

Ssd 2

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The introductory framing uses editorial and quality-assurance language to mask an anti-detection purpose, including repeated references to '去AI味', '降AI率', WeChat rejection, and detector-specific calibration. This euphemistic packaging is dangerous because it normalizes evasion as ordinary copy-editing and can cause downstream systems or reviewers to under-classify a capability that is clearly oriented toward defeating AI-origin detection.

Content

No source excerpt is available for this finding.

Ssd 4

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This section gives detailed, procedural guidance for systematically rewriting text to evade AI-content detectors and to alter publication behavior so platform enforcement is less likely to detect synthetic or assisted content. It goes beyond benign style editing by prescribing chunk-based testing, signal removal, 'humanizing' insertions, detector thresholds, and release-discipline tactics, which directly facilitate deceptive bypass of trust/safety systems.

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/scanner.py (reported line 17)May include surrounding context.

python
for f in sorted(os.listdir(RULES_DIR)):
        if f.startswith('R_') and f.endswith('.json'):
            rules.extend(json.load(io.open(os.path.join(RULES_DIR, f), encoding='utf-8')))
    return rules

def strip_body(text):
    for mark in ('**信源备注**', '**处理说明**', '\n---\n'):

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The README states a general-purpose workflow in English, but line L073 describes the core finding in terms of Zhuque scoring behavior and line L101 later limits the corpus to Chinese self-media commentary. This can mislead users into applying a Chinese-locale writing policy by default rather than through explicit opt-in or clear locale scoping.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest says to use the skill not only for specific phrases like 去AI味 or 微信打回, but also 'even if they only say "这篇文章帮我看看"'. That phrase is generic everyday speech and could overlap with many unrelated article-review requests, making activation scope too broad.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill is explicitly designed to help users alter text so it will evade AI-content detectors before publication. That is dangerous because it operationalizes deceptive evasion tactics against platform integrity systems, enabling users to bypass moderation or authenticity checks rather than improve writing quality in a neutral way.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The document states 'Do not use' for English text because the corpus is Chinese commentary, which imposes a language restriction. While there is a technical rationale, the skill does not present this as a user choice or opt-in and broadly forces a single language context.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The entire roadmap is written as a Chinese-only artifact and repeatedly describes the skill packaging and workflow without offering any language or locale choice. Under the policy rule for natural-language constraints, this can indicate a language/locale restriction that is not documented as optional or region-specific.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger list includes broad phrases such as '终检' and especially '这篇文章帮我看看' by implication in the skill metadata, making accidental invocation likely outside a narrowly consented anti-detection workflow. Because the skill provides detector-evasion guidance, overbroad triggering increases the chance the system applies deceptive publication tactics in ordinary editing contexts where the user did not explicitly request that behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The document is written entirely around a Chinese-language, WeChat-specific workflow and repeatedly assumes a single language/locale context without offering user choice or opt-in. While region specificity can be acceptable, this file does not clearly present the locale restriction as an explicit documented constraint for users outside that context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This JSON rule set is entirely built around Chinese lexical patterns and labels, such as Chinese trigger terms, quantity words, and sentence-ending phrases. Because the file provides no natural-language indication that the skill is limited to Chinese-language content or that users can choose locale, it appears to enforce a specific language/locale implicitly.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

This JSON file defines numerous Chinese-specific patterns and labels such as Chinese rule names and regexes matching Chinese text, but it provides no natural-language indication that the behavior is limited to Chinese content or requires user opt-in for that locale. A skill that enforces one language's stylistic policy by default can violate language/locale choice expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This JSON rule set is entirely expressed in Chinese across user-facing fields such as "name", "pattern", and "fix", and there is no accompanying natural-language indication that the skill is intended only for Chinese-language users or content. Under the language/locale policy, forcing a specific language without opt-in or explicit justified scope is a policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This JSON manifest defines all rule names, patterns, fixes, and exemption phrases in Chinese only, such as "赋能", "说白了", and other Chinese-language matching terms. Because the file provides no indication of multilingual support, locale scoping, or user opt-in, it appears to enforce a specific language policy by default.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The rule triggers on the standalone token "仿佛" with mode "any" and threshold 0, so it will match many normal Chinese passages rather than a narrowly defined AI-writing artifact. In a skill explicitly designed to rewrite text to evade AI-content detection, broad lexical triggers increase the chance of unjustified rewrites and materially support detector evasion by overfitting on common language.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The rule flags the common word "犹如" unconditionally, making it far too broad to distinguish AI-generated phrasing from ordinary human writing. Because this skill's purpose is to reduce AI-detection signals before publication, such low-specificity triggers directly facilitate stylometric evasion while degrading reliability.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

"一丝" is an extremely common phrase fragment and will appear in many benign texts, so using it as an unconditional AI indicator creates a high false-positive trigger surface. In the context of an anti-detection editing skill, this broad match encourages unnecessary transformations that help users disguise generated content.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The standalone match on "微微" is too generic to serve as a safe detection rule and will capture ordinary descriptive writing. Within a system aimed at removing AI-writing tells, this overbroad trigger contributes to content laundering by rewriting normal prose based on weak signals.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.