Back to skill

Security audit

技能架构师 / Skill Architect

Security checks for vulnerabilities and agentic risk

Overview

The skill’s purpose is coherent, but it needs review because it can read prior conversation snapshots, modify local files, and run broad local test commands with limited containment.

Review before installing. Use this only on repositories and conversation exports you are comfortable letting it read and rewrite, avoid generic continuation triggers, provide explicit snapshot/output paths, inspect generated handoff zip contents before sharing, and do not pass untrusted text into --test-command.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
Findings (47)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the associated tooling performs privacy scanning and broad filesystem traversal while the skill is described as a document refactoring/conversation tool, users may authorize it under false assumptions. That mismatch can expose unrelated files or sensitive release artifacts to inspection without informed consent, which is a meaningful security and privacy concern.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the associated tooling performs privacy scanning and broad filesystem traversal while the skill is described as a document refactoring/conversation tool, users may authorize it under false assumptions. That mismatch can expose unrelated files or sensitive release artifacts to inspection without informed consent, which is a meaningful security and privacy concern.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

If the associated tooling performs privacy scanning and broad filesystem traversal while the skill is described as a document refactoring/conversation tool, users may authorize it under false assumptions. That mismatch can expose unrelated files or sensitive release artifacts to inspection without informed consent, which is a meaningful security and privacy concern.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the associated tooling performs privacy scanning and broad filesystem traversal while the skill is described as a document refactoring/conversation tool, users may authorize it under false assumptions. That mismatch can expose unrelated files or sensitive release artifacts to inspection without informed consent, which is a meaningful security and privacy concern.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

If the associated tooling performs privacy scanning and broad filesystem traversal while the skill is described as a document refactoring/conversation tool, users may authorize it under false assumptions. That mismatch can expose unrelated files or sensitive release artifacts to inspection without informed consent, which is a meaningful security and privacy concern.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the associated tooling performs privacy scanning and broad filesystem traversal while the skill is described as a document refactoring/conversation tool, users may authorize it under false assumptions. That mismatch can expose unrelated files or sensitive release artifacts to inspection without informed consent, which is a meaningful security and privacy concern.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger phrase '继续讨论 <问题>' is extremely broad and overlaps with ordinary conversational language, making accidental activation likely. In this skill, activation can lead to file discovery, reading snapshots, and execution of helper scripts, so an innocuous user message could unintentionally invoke sensitive operations or leak prior conversation context.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 16)May include surrounding context.

md
把过大的 `SKILL.md` 从“整本常驻”改成“目录 + 按需读取”。核心是结构外移,不是文字压缩。

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 37)May include surrounding context.

md
把过大的 `SKILL.md` 从“整本常驻”改成“目录 + 按需读取”。核心是结构外移,不是文字压缩。

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

This is a tool-parameter abuse issue: an attacker or untrusted caller can supply a malicious --test-command value that the script will execute verbatim through the shell. In agent or automation contexts, parameter injection is especially dangerous because the tool may be invoked indirectly from model-generated or externally influenced input.

Content

Scanner excerpt · scripts/validate.py (reported line 54)May include surrounding context.

python
ok = False

    if args.test_command:
        result = subprocess.run(args.test_command, shell=True, capture_output=True, text=True)
        if result.returncode == 0:
            print("test_command=OK")
        else:

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The product description is entirely in Chinese with no language choice, fallback, or stated locale constraint. For a general-purpose skill that performs audits, refactoring, snapshots, and transfer workflows, this can cause users or downstream agents to misunderstand capabilities, boundaries, and operational steps, increasing the chance of unsafe or incorrect use.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This README presents all user-facing instructions, triggers, and workflow descriptions exclusively in Chinese. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly documented and justified.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill instructs the agent to read, write, and execute local scripts, but it does not declare any explicit tool scope such as permissions or allowed-tools. That creates unnecessary ambiguity about what the skill may access or modify, increasing the chance of over-broad filesystem changes or shell execution in environments that rely on manifest-level gating.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

All listed trigger phrases are in Chinese, and the file does not state that the skill is intentionally limited to Chinese users or provide an opt-in language choice. This creates a natural-language policy concern because activation appears tied to a specific language without justification.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger phrase “帮我总结这个长对话” is overly broad and overlaps with ordinary summarization requests, so the skill may be invoked when the user only wants a simple summary rather than archival, extraction, or packaging behavior. In this skill’s context, accidental activation is more dangerous because it can cause unintended processing of large conversation history into snapshots or skill packs, increasing privacy exposure and surprising side effects.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest description describes a very broad set of capabilities: auditing and refactoring SKILL.md files, conversation snapshotting, reusable pack generation, and cross-agent continuation. Without clear activation constraints, a host system or user may invoke this skill in contexts involving sensitive conversation data or broader-than-expected repository changes, increasing the chance of overcollection, unintended execution, or misuse. The skill context makes this more dangerous because it explicitly handles long conversations, snapshots, and cross-agent transfer, which are operations often involving sensitive data and high trust.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger phrase “继续讨论 <问题>” is overly broad and overlaps with normal conversational language, which can cause the skill to activate when the user merely intends to continue a discussion. In this skill’s context, unintended activation is more dangerous because it may automatically locate and load prior conversation snapshots and skill packs, pulling in stale or unrelated context without an explicit handoff request.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The generated snapshot content uses fixed Chinese strings such as "整个当前对话框", "指定主题", and "自动候选,需复核". Combined with the Chinese-only fallback topic on L112 and other Chinese output strings later in the file, this forces a specific language/locale in generated artifacts without offering the user a language choice or documenting a justified locale restriction.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script trusts config-supplied topic output paths and writes generated files directly to those locations, only rebasing relative paths under the output directory while allowing absolute paths unchanged. If an attacker can influence the config, they can cause arbitrary file creation or overwrite anywhere writable by the process, which exceeds the intended scope of conversation snapshot generation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The generated README strings at L39-L44 are entirely in Chinese and provide no option for another language or indication that the skill is region-specific. This creates a natural-language locale policy issue because the exported artifact imposes a specific language on downstream users without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The default value for --replace-marker is hardcoded in Chinese ("已外移说明 / 内容已外移到独立文件") and will be written into output files unless the caller overrides it. This imposes a specific language choice in the skill's behavior without offering a user language preference or explaining a region-specific requirement.

Content

No source excerpt is available for this finding.

Tainted flow: 'extracted_text' from pathlib.Path.read_text (line 58, file read) → pathlib.Path.write_text (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/extract.py (reported line 48)May include surrounding context.

python
lines[start:end] = [marker]
            output = Path(item["output"])
            output.parent.mkdir(parents=True, exist_ok=True)
            output.write_text(extracted_text, encoding="utf-8")
            outputs += 1
    else:
        if not args.markers or not args.output:

Tainted flow: 'extracted_text' from pathlib.Path.read_text (line 58, file read) → pathlib.Path.write_text (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/extract.py (reported line 63)May include surrounding context.

python
lines[start:end] = [marker]
            output = Path(item["output"])
            output.parent.mkdir(parents=True, exist_ok=True)
            output.write_text(extracted_text, encoding="utf-8")
            outputs += 1
    else:
        if not args.markers or not args.output:

Tainted flow: 'lines' from pathlib.Path.read_text (line 33, file read) → pathlib.Path.write_text (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/extract.py (reported line 66)May include surrounding context.

python
output.write_text(extracted_text, encoding="utf-8")
        outputs = 1

    skill.write_text("".join(lines), encoding="utf-8")
    print(f"extracted_lines={extracted_lines}")
    print(f"outputs={outputs}")

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/self_test.py (reported line 16)May include surrounding context.

python
def run(*args, cwd):
    result = subprocess.run([PYTHON, *args], cwd=cwd, capture_output=True, text=True)
    if result.returncode != 0:
        raise SystemExit(f"FAILED: {args}\n{result.stdout}\n{result.stderr}")

Static analysis

No suspicious patterns detected.