Back to skill

Security audit

内容过滤工具专业版

Security checks for vulnerabilities and agentic risk

Overview

The skill appears intended for team content filtering, but its instructions are overbroad and under-specified for command/API use, multi-account changes, sensitive audit logs, and third-party data sharing.

Review this skill carefully before installing. Use it only for clearly authorized content-filtering workflows, give it least-privilege API keys, verify every webhook destination, redact or minimize sensitive content sent to external services, and define retention and approval rules for audit logs and multi-account rule changes.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Description-Behavior Mismatch

Medium
Confidence
93% confidence
Finding
The skill manifest presents a narrow content-filtering tool, but later documentation expands it into generic automation capabilities including file processing, API integration, and command execution. This scope mismatch can cause an agent or user to invoke the skill in broader contexts than intended, increasing the chance of over-privileged use and unsafe task routing.

Intent-Code Divergence

High
Confidence
97% confidence
Finding
The document simultaneously says the skill is not suitable for compliance auditing while also advertising compliance auditing as a core capability. This contradiction can mislead operators and automated systems about when the skill should be trusted, causing inappropriate use for governance-sensitive tasks without clear assurance boundaries.

Intent-Code Divergence

High
Confidence
98% confidence
Finding
The usage guidance tells users to use the skill for security detection, compliance audit, vulnerability scanning, and encryption protection, which materially exceeds its declared content-filtering scope. This can trigger unsafe reliance on the skill for security-critical workflows it does not implement correctly, creating a false sense of protection.

Vague Triggers

Medium
Confidence
86% confidence
Finding
The trigger language is broad and ambiguous, making the skill likely to activate for loosely related requests. Over-broad invocation conditions are dangerous because they can route unrelated tasks into a tool with Bash, file-write, and network-adjacent behaviors, increasing accidental misuse.

Vague Triggers

Medium
Confidence
90% confidence
Finding
The '使用时机' section gives ambiguous and contradictory activation guidance, which may cause the skill to be invoked for inappropriate high-risk tasks. In the context of a skill with Bash and external API usage, ambiguous routing materially raises the chance of unsafe execution paths.

Missing User Warnings

Medium
Confidence
91% confidence
Finding
The skill describes sending filtering events to a webhook, including potentially sensitive moderation outcomes, without a prominent warning about data disclosure, minimization, retention, and trust of the receiving endpoint. This creates a realistic risk of leaking sensitive content metadata or blocked-content details to external systems.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The documentation instructs users to configure API keys and use external AI APIs without clearly warning that user data, content samples, and moderation artifacts may be sent to third-party services. This is dangerous because users may unknowingly export regulated or sensitive data outside their trust boundary.

Static analysis

No suspicious patterns detected.