Back to skill

Security audit

NPC浓度测试

Security checks for vulnerabilities and agentic risk

Overview

This is a local Chinese-language self-reflection quiz with disclosed CLI and HTML tooling, and I found no evidence of hidden network access, credential use, persistence, or destructive behavior.

Install this only if you want a Chinese-language, sharp-toned self-reflection game. Do not treat its output as psychological, medical, or workplace advice, and avoid sharing sensitive personal details. Run developer commands such as --calibrate or HTML generation only when you intend to modify local files, and do not rely on the package's 'no logs' wording for guarantees about the host platform's own logging.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (19)

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared skill purpose is an interactive self-assessment, but the instructions embed operational behavior for local script execution, filesystem access, environment inspection, HTML build steps, and Playwright/browser smoke testing. This mismatch is dangerous because users and calling systems may authorize the skill as a simple quiz while it actually invokes materially broader execution and host-inspection behaviors.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill instructs the agent to run local Python scripts, generate HTML files, read project files, and perform browser-based smoke tests, but it does not declare any explicit tool scope or permissions boundaries. In an agent environment, this creates an overbroad capability surface where a seemingly harmless quiz skill can access the filesystem, shell, environment, and browser automation without user-visible restriction.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger phrases include broad, everyday expressions such as feeling repetitive, unheard, or wanting a reset, which can cause accidental invocation in unrelated or emotionally sensitive conversations. In context, this can steer users into a psychologically loaded self-assessment flow without clear intent, increasing the chance of inappropriate activation and mishandling of sensitive user states.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The description promotes a self-diagnostic style experience and social-sharing output but does not warn users that the content may feel personally sensitive or that results are designed for propagation and comparison. In this context, missing disclosure increases the risk of users being surprised by stigmatizing framing, emotional impact, or unintended sharing of reflective personal content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

This manifest contains user-facing title, zone names, opening text, exit text, banned words, and disclaimer entirely in Chinese, with no indication that users can opt into another language or that the skill is intentionally limited to a Chinese-speaking context. Under the policy, forcing a specific language without user choice can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file presents all instructions and content in a single language with no indication that users may choose another language or locale. Under the policy rule for natural-language violations, forcing a specific language without user opt-in is a reportable issue unless the locale restriction is documented and justified, which it is not here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The file enforces a single Chinese-language stylistic and worldbuilding presentation with no user choice of language or locale. While not directly a classic security flaw, this can cause misunderstanding of tone, consent text, and safety boundaries for users who are not comfortable in that locale, making the abrasive style harder to interpret safely.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The style guide explicitly directs the skill to use emotionally abrasive, cutting language ('嘴很毒', '准确 > 温柔', '被击中') without requiring an upfront warning, consent gate, or opt-out. In a self-reflection skill triggered by users expressing distress or alienation, this increases the risk of emotional harm, especially for vulnerable users who may not expect confrontational output.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/build_quiz_html.py (reported line 485)May include surrounding context.

python
tmp = Path(tempfile.gettempdir()) / "_npc_verify.js"
    tmp.write_text(js, encoding="utf-8")
    try:
        res = subprocess.run(["node", str(tmp)], capture_output=True, text=True,
                             encoding="utf-8", timeout=60)
    except FileNotFoundError:
        print("[skip] 未找到 node,跳过对拍")

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/build_quiz_html.py (reported line 535)May include surrounding context.

python
if not js.exists():
        print("[SKIP] 缺少 scripts/smoke.js")
        return 0
    r = subprocess.run([node, str(js), str(html_path)],
                       capture_output=True, text=True, encoding="utf-8")
    print((r.stdout or "").strip())
    if r.returncode != 0:

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This Python file contains the main user-facing instructions and generated output strings in Chinese only, including usage text, prompts, and results. That creates a language/locale restriction without any visible user opt-in or documented justification, which matches the natural-language policy violation criteria.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest describes an NPC-style self-diagnostic test that asks questions, classifies answers, and optionally produces perspective-writing output. However, the code also exposes maintenance-style commands such as --calibrate, which recomputes anchors and writes them back to meta.json, and --audit, which performs internal dataset validation; these are developer tooling behaviors rather than the described end-user test functionality.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The cmd_calibrate path recomputes calibration statistics and persists them by overwriting data/meta.json. For a skill described as generating questionnaire results and character cards, local file mutation is a meaningful extra behavior that is not implied by the manifest's stated purpose.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The metadata makes a categorical privacy/operational assurance ('不联网、不上传、不留档') that cannot be substantiated by this file alone. If the surrounding system logs, transmits, or stores user inputs, users may disclose sensitive emotional or personal information under a false sense of privacy, making this a deceptive trust claim.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This JSON file presents all user-facing note, question, and option text in Chinese, with no indication that users can select another language or opt in to this locale. Under the natural-language policy criteria, forcing a specific language without user choice can be a locale-policy violation unless clearly documented as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The page declares lang="zh-CN", and all visible UI strings are hard-coded in Chinese throughout the file. Under the policy rule, forcing a specific language without user opt-in can be a natural-language policy violation when no alternative locale or selection mechanism is provided.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The generated page sets lang="zh-CN", and the UI text throughout the template is written only in Chinese. Under the policy rule for language/locale, this is a forced locale choice without any visible opt-in or alternative.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This script’s natural-language content, including the header comments and all console messages, is entirely in Chinese, with no indication that the user can select another language. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.