Back to skill

Security audit

AI Gen Guard -生成式 AI 服务合规护栏

Security checks for vulnerabilities and agentic risk

Overview

This is a local Chinese generative-AI compliance checker with no network or persistence; the main risk is over-relying on heuristic output as legal guidance.

Install only if you need a Chinese-language, PRC-focused generative-AI compliance triage helper. Confirm the service is in the relevant jurisdiction and consult qualified counsel for real compliance decisions; do not treat this tool's output as proof of compliance.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (8)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The trigger conditions are broad topical descriptions without clear inclusion/exclusion boundaries, so an agent may invoke this skill for loosely related AI or compliance queries that do not actually require this China-specific legal guardrail. In security terms this is a scope-control weakness: over-triggering can cause irrelevant or jurisdictionally incorrect compliance guidance to override better-suited tools, increasing policy and decision risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The package description is entirely in Chinese and frames the skill around PRC-specific regulations and standards without indicating any user language choice or opt-in. This can violate language/locale policy expectations when a skill is presented in a fixed language without documenting that it is region-specific or offering alternatives.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module presents itself as performing compliance assessment based on specific laws and a mandatory national standard, but the implementation is only keyword and substring matching with simple negation heuristics. In a compliance-guard context, this can mislead users into treating a shallow heuristic result as a substantive legal/regulatory assessment, causing false assurance or incorrect risk decisions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module description and all user-facing strings are written exclusively in Chinese and the implementation is tailored to Chinese-language text patterns. There is no indication of user opt-in, language selection, or justification that the skill is intentionally limited to a Chinese-only compliance context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module-level docstring and all rule labels, descriptions, questions, and indicators are written exclusively in Chinese, indicating a fixed language/locale assumption. The policy requires flagging language or locale constraints when they are imposed without explicit user opt-in or documented justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code file contains natural-language interface text and documentation that force a specific language/locale for usage, examples, prompts, and errors. The policy allows locale constraints only when justified or when the user is given a choice, neither of which is present here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The skill is hard-scoped to Chinese law and enterprises but does not require locale/jurisdiction confirmation before use, which can lead to mismatched legal guidance for users in other regions. This is less a direct exploit vector than a policy-safety issue, but it can still cause harmful misapplication of legal/compliance advice when the agent assumes the wrong jurisdiction.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The negative-signal detection only checks the first occurrence of each fail string with text.find(signal), despite comments implying robust exact matching. This can miss later occurrences, fail to evaluate multiple contexts, and produce inaccurate compliance conclusions when adversarial or complex text places a benign-looking instance first and a true negative instance later.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.