Back to skill

Security audit

代码自动运行和修复

Security checks for vulnerabilities and agentic risk

Overview

This skill does what it advertises, but it runs user and LLM-modified code on the host without visible sandboxing, approval gates, or dependency pinning.

Install only in a disposable, least-privileged sandbox or container with no secrets, no sensitive source trees, strict resource limits, and network access disabled unless explicitly needed. Do not submit proprietary code or credentials unless you are comfortable sending the code and error output to the platform LLM. Pin and review dependencies before deployment.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
skill.py:9
Finding

Unsandboxed Execution of User-Supplied Python, C, and Assembly Code

Content
View full analysis
str: def run(c): with tempfile.NamedTemporaryFile(suffix=".py", delete=False) as f: f.write(c.encode()) try: return safe_run(["python3", f.name]) finally: os.unlink(f.name) return await auto_run_and_fix(code, "python", run) @server.tool() async def run_c(code: str) -> str: def run(c): with tempfile.NamedTemporaryFile(suffix=".c", delete=False) as f: f.write(c.encode()) ex = f.name + ".out" try: _, compile_err = safe_run(["gcc", f.name, "-o", ex]) if compile_err: return "", compile_err return safe_run([ex]) finally: if os.path.exists(ex): os.unlink(ex) os.unlink(f.name) return await auto_run_and_fix(code, "c", run) @server.tool() async def run_assembly(code: str) -> str: def run(c): with tempfile.NamedTemporaryFile(suffix=".asm", delete=False) as f: f.write(c.encode()) o = f.name + ".o" ex = f.name + ".out" try: _, asm_err = safe_run(["nasm", "-f", "elf64", f.name, "-o", o]) if asm_err: return "", asm_err _, link_err = safe_run(["ld", o, "-o", ex]) if link_err: ...[truncated 2693 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
skill.py:20
Finding

Automatic Execution of Untrusted LLM-Generated Code

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned Third-Party Dependencies Permit Non-Reproducible and Unsafe Resolution

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (12)

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill exposes tools that write user-supplied Python, C, and assembly to disk, compile or interpret them, and execute them as local processes. This enables arbitrary code execution on the host, allowing file access, network access, persistence attempts, privilege abuse, and denial of service; the auto-fix loop can even re-run LLM-modified code without meaningful trust boundaries.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill executes arbitrary user-supplied code and toolchain binaries without any warning or disclosure, creating a hidden local code-execution surface. In the absence of a strongly justified sandboxed coding-runner context, this makes the skill far more dangerous because callers may not understand they are granting code execution on the machine.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill explicitly advertises automatic execution of Python, C, and assembly code, plus automatic LLM-driven code repair and returning modified code, but it does not warn about the risks of executing untrusted code or silently altering user-supplied programs. This is dangerous because it normalizes a high-risk workflow that can lead to arbitrary code execution, unsafe compilation/linking, and unreviewed code changes without user consent or safety boundaries.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The description states the skill will automatically run code and use an LLM to repair it, then directly return modified code, but it does not define any approval step, scope restriction, or trigger boundary. In a skill that exposes execution tools such as run_python, run_c, and run_assembly, this ambiguity is dangerous because it can normalize autonomous code execution and modification without clear user consent or safety checks.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The description is written in Chinese and implicitly fixes the interaction/output locale without any indication of user choice. While not inherently a security flaw on its own, forced locale behavior can mislead users about what the skill will do or reduce reviewability of risky actions, especially when paired with automatic execution and code-fixing behavior.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · skill.py (reported line 12)May include surrounding context.

python
def safe_run(cmd, timeout=10):
    try:
        result = subprocess.run(
            cmd,
            capture_output=True,
            text=True,

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill sends full user-provided code and associated error details to an external LLM service for automatic repair. If the submitted code contains proprietary source, secrets, credentials, or internal paths, this creates an unauthorized data exfiltration channel to a third party.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code forwards user code and error output to an external LLM without any visible notice, consent, or disclosure. This is dangerous because users may reasonably assume local processing while sensitive source code is actually transmitted off-host.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

All user-facing natural-language content in the skill description is presented only in Chinese, with no indication that users may choose another language or that the locale restriction is intentional and documented. This can violate language/locale policy when a skill implicitly forces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Unverifiable Dependency: mcp has 12 known advisory(ies) (CVE-2025-53366 (MCP Python SDK vulnerability in the FastMCP Server causes validation error, lead); CVE-2025-66416 (Model Context Protocol (MCP) Python SDK does not enable DNS rebinding protection); CVE-2026-52870 (MCP Python SDK: Experimental task handlers allow any client to access and cancel) +9 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

The manifest references mcp[server] without a fixed version even though the package family has multiple known advisories. Because the version is unconstrained, an installation could resolve to an affected release, leaving the skill exposed to known server-side issues such as validation, DNS rebinding, or client/task handling weaknesses depending on runtime usage.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

The dependency clawhub is specified without a version pin, making builds non-reproducible and allowing future installs to silently pull newer or compromised releases. In a skill/package supply-chain context, this increases exposure to malicious publishes, typosquatting replacement risk, or breaking changes entering the environment unexpectedly.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
mcp[server]
clawhub

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language instruction string forces the repair interaction into Chinese ('你是专业代码修复专家...') regardless of user preference. This can violate language/locale policy expectations because no opt-in, user choice, or documented locale-specific justification is provided.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.