Back to skill

Security audit

胡田 OPC导师 公众号运营

Security checks across malware telemetry and agentic risk

Overview

The skill is mostly a coherent WeChat public-account operations toolkit, but it includes explicit workflows for disguising AI-generated text to improve platform ranking and avoid AI-detection signals.

Install only if you are comfortable reviewing all generated public-facing content before use. Avoid using the AI-expression detector or 'human creation score' guidance to conceal AI authorship or bypass platform rules, and redact personal details from user comments before adding them to analytics or reports.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (17)

Vague Triggers

Medium
Confidence
93% confidence
Finding
The trigger phrase '帮我运营公众号' is broad and natural-language generic, making accidental invocation plausible during ordinary conversation about social media operations. Overbroad activation can cause the agent to switch into this skill unexpectedly, apply domain-specific instructions without clear user consent, and overshadow other safer or more appropriate skills.

Missing User Warnings

Medium
Confidence
88% confidence
Finding
The template explicitly asks operators to collect, categorize, and quote representative user comments and extract user needs, but provides no guidance on minimizing, anonymizing, or securing potentially sensitive personal data contained in those comments. In a real公众号运营 workflow, comments may include contact details, business needs, institutional affiliations, or other personal information, so encouraging structured reuse without privacy controls creates a meaningful risk of overcollection, unnecessary retention, and accidental disclosure.

Vague Triggers

Medium
Confidence
95% confidence
Finding
The activation phrases are very broad (for example, generic requests to add watermarking or generate branded content) and lack clear scoping conditions or confirmation requirements. This can cause the skill to trigger during ordinary editing or drafting tasks and silently alter article content by inserting branding, opinion blocks, or watermark text that the user did not explicitly authorize in that specific context.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The skill describes operations that directly modify article text and image files, but it does not clearly warn users that outputs will be altered or that image processing may overwrite or create changed files. In an agent workflow, this lack of disclosure increases the risk of unintended destructive or reputation-affecting edits, especially because the skill adds author-branded statements and opinion content rather than performing a neutral formatting pass.

Vague Triggers

Medium
Confidence
86% confidence
Finding
The trigger conditions are broad enough to match many ordinary writing or marketing requests, which can cause the skill to activate unexpectedly and override more appropriate workflows. In this skill, accidental activation is more concerning because it also pushes algorithm-gaming, AI-detection evasion, and image-generation behaviors the user may not have requested.

Vague Triggers

Medium
Confidence
88% confidence
Finding
The front-matter description is overly broad and frames the skill as applicable to general marketing copy and new-media content, increasing the chance of unintended routing. Because the skill contains behavior aimed at optimizing for platform algorithms and reducing detectable AI traits, overbroad invocation expands the blast radius of those problematic instructions.

Natural-Language Policy Violations

Medium
Confidence
90% confidence
Finding
The template hard-codes 'male editor' and 'female editor' personas as a primary workflow choice without indicating user consent, neutrality options, or accessibility safeguards. In a content-generation skill, this can lead to gender stereotyping, exclusion of users who do not want gendered framing, and biased outputs being produced at scale.

Natural-Language Policy Violations

Medium
Confidence
92% confidence
Finding
The guideline hard-codes gendered personas ('男编辑版' and '女编辑版') and maps them to stereotyped audience assumptions without requiring user choice or consent. This can cause the agent to generate biased, exclusionary, or inappropriate content for users and audiences, creating fairness, reputational, and potentially compliance risks in a content-generation workflow.

Natural-Language Policy Violations

High
Confidence
98% confidence
Finding
This section explicitly frames platform AI-detection as something to work around, describing detection mechanisms and the benefit of increasing '人工创作' scores and recommendation weight. That moves beyond normal style guidance into operational evasion of platform integrity controls, which can facilitate deceptive publication of AI-generated content at scale.

Natural-Language Policy Violations

High
Confidence
97% confidence
Finding
These examples and rewrite patterns systematically teach users how to transform AI-sounding text into content that appears human-authored specifically to reduce detection. In the context of a WeChat public-account operations skill, this directly enables deceptive content laundering and can be used to bypass platform moderation or provenance signals.

Natural-Language Policy Violations

High
Confidence
99% confidence
Finding
The inclusion of detection and checking tools to find '疑似AI表达' and then eliminate those markers operationalizes evasion into a repeatable workflow. This makes the risk more serious than passive advice because it provides a practical mechanism to test and optimize content against anti-abuse heuristics.

Vague Triggers

Medium
Confidence
90% confidence
Finding
The trigger condition is phrased very broadly ('when users need WeChat article typesetting, optimize reading experience, or improve传播力'), which overlaps with many ordinary writing or editing requests. In an agent system that auto-selects skills based on natural language, this can cause unintended activation and inappropriate routing, reducing reliability and potentially exposing user content to an unnecessary skill context.

Ssd 2

Medium
Confidence
95% confidence
Finding
The skill explicitly instructs the model to avoid AI-like traits and use 'humanized' expression to reduce detectable AI signals, which goes beyond normal stylistic editing into evasion guidance. In context, this is dangerous because it helps users disguise synthetic content for better platform performance and reduced scrutiny, facilitating deceptive content laundering.

Ssd 2

Medium
Confidence
96% confidence
Finding
The built-in quality checks operationalize AI-expression detection and then direct rewriting toward more human-like output, creating a feedback loop for evading AI-detection heuristics. This is more dangerous in a content-production skill because it can systematically scale deceptive synthetic content while preserving plausible deniability.

Ssd 2

Medium
Confidence
94% confidence
Finding
The resource guidance explicitly links reducing detectable AI traits with better platform performance, which encourages concealment as an optimization goal. That makes the issue more serious than harmless style advice, because it normalizes deceptive manipulation of trust and ranking systems.

Ssd 2

Medium
Confidence
91% confidence
Finding
The semantic behavior here is to coach paraphrasing for the purpose of making AI-generated text appear human and less detectable. Even though it is presented as content optimization, the stated goal is evasion of platform heuristics, which undermines trust and platform policy enforcement.

Ssd 4

Medium
Confidence
90% confidence
Finding
The document walks users through a staged process: understand detection signals, identify AI-like features, and then rewrite content to avoid those signals. That structured evasion workflow increases misuse potential because it lowers the skill threshold for bypassing platform content-authenticity checks.

VirusTotal

VirusTotal findings are pending for this skill version.

View on VirusTotal

Static analysis

No suspicious patterns detected.