Back to skill

Security audit

灵枢·AI全栈构建师

Security checks for vulnerabilities and agentic risk

Overview

This skill is a broad full-stack development helper, but it includes executable helpers with powerful, under-scoped behavior that users should review before installing.

Review this skill before installing in a sensitive workspace. Do not run its code-execution helper on untrusted or model-generated input, restrict AI provider endpoints before using environment API keys, keep generated-code writes inside an approved project directory, and avoid enabling the backup or scheduled update scripts until their paths and shell usage are fixed.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
integrations/framework_enhancements.py:332
Finding

Unrestricted In-Process Python Code Execution

Content
View full analysis
Dict[str, Any]: result = { "agent": self.name, "code": code, "timestamp": datetime.now().isoformat(), "output": None, "error": None } try: local_vars = {} exec(code, {"__builtins__": __builtins__, **self.tools}, local_vars) ``` ### Technical Analysis The `execute_code` method passes caller-controlled Python source directly to `exec`. The execution context exposes the complete Python built-in namespace through `__builtins__`. Consequently, executed code can import modules, access the filesystem, inspect environment variables, create network connections, and start operating-system processes. Although the module describes this feature as code-agent or sandbox functionality, the implementation provides no process isolation, import restrictions, syscall filtering, filesystem boundary, network restriction, timeout, or resource limit. Exception handling only records errors after execution and does not constrain what the payload can do. This exceeds the minimum privilege required for generating or reviewing code. The dangerous method is a callable library entry point rather than an automatic startup path, but any untrusted, user-supplied, or model-generated input passed to it receives the full privileges of the host process. ### Attack Path 1. An attacker influences the `code` argument passed to `CodeAgent.execute_code`. 2. The supplied code imports modules such as `os`, `pathlib`, `socket`, or `subprocess`. 3. `exec` evaluates the payload with unrestricted built-ins. 4. The payload reads environment credentials or local files, modifies project or configuration files, launches commands, or opens outbound network connections. 5. ...[truncated 473 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
integrations/platform_adapters.py:42
Finding

API Credentials and Prompt Data Can Be Sent to Arbitrary Endpoints

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
code-generator/code_engine.py:188
Finding

Caller-Controlled Output Path Allows Arbitrary File Overwrite

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/auto_update_backup.py:22
Finding

Unsafe Shell Execution and Backup Writes Outside the Skill Boundary

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
integrations/platform_adapters.py:239
Finding

Remote Model Artifacts Are Loaded Without Immutable Version Pinning

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (172)

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
88% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
89% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

If the actual behavior is primarily a mock monitoring or dashboard UI for PRD/log/pattern display rather than the broad engineering functions advertised, users may misunderstand what data is collected, displayed, or stored. Misrepresentation is less severe than direct code execution, but it still expands privacy and integrity risk through incorrect trust assumptions.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
97% confidence
Finding

The documentation includes a sample .env block with a JWT secret placeholder and database URI, which can normalize storing secrets in plaintext files and may lead users to commit real credentials into source control. In a code-generation skill, examples are often copied verbatim, so this increases the chance of insecure secret handling in downstream projects.

Content

Scanner excerpt · code_snippets/nodejs_snippets.md (reported line 547)May include surrounding context.

text

```env
# .env
PORT=3000
NODE_ENV=development
MONGO_URI=mongodb://localhost:27017/myapp

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The module documentation advertises a 'code execution sandbox', but the implementation directly invokes exec() with full builtins, which is not a sandbox. This misleading safety claim can cause developers or operators to trust the feature and deploy it in higher-risk contexts, increasing exposure to arbitrary code execution and unsafe assumptions.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This component intentionally provides arbitrary Python execution capability inside a skill whose stated purpose is development guidance, templates, and framework integration, not trusted remote code execution. That scope mismatch increases the likelihood that unsafe execution paths are exposed where users and operators would not expect them, turning prompt/input content into code execution on the host.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

Dynamic code execution occurs without any user-facing warning, confirmation, or trust check, so untrusted or unexpected inputs may be executed silently. In agentic systems, this materially increases the chance of prompt-driven or workflow-driven unsafe execution because operators are not given a chance to review or deny dangerous actions.

Content

No source excerpt is available for this finding.

exec() call detected

High
Category
Dangerous Code Execution
Confidence
99% confidence
Finding

The code executes attacker-controlled Python via exec() with full builtins exposed and injected tools available, which enables arbitrary code execution, file access, process spawning, imports, and data exfiltration. In the context of a general full-stack development skill, this is especially dangerous because code execution is not narrowly constrained to a justified sandboxed runtime and could be reached by untrusted prompts or task content.

Content

Scanner excerpt · integrations/framework_enhancements.py (reported line 343)May include surrounding context.

python
try:
            local_vars = {}
            exec(code, {"__builtins__": __builtins__, **self.tools}, local_vars)
            result["output"] = local_vars.get("result", "执行成功")
        except Exception as e:
            result["error"] = str(e)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/backend_frameworks.md (reported line 28)May include surrounding context.

│ └── server.js # 服务器启动 ├── tests/ # 测试 ├── package.json └── .env

text

**命名规范:**

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/backend_frameworks.md (reported line 788)May include surrounding context.

│ └── server.js # 服务器启动 ├── tests/ # 测试 ├── package.json └── .env

text

**命名规范:**

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/security_best_practices.md (reported line 499)May include surrounding context.

│ └── server.js # 服务器启动 ├── tests/ # 测试 ├── package.json └── .env

text

**命名规范:**

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/security_best_practices.md (reported line 500)May include surrounding context.

│ └── server.js # 服务器启动 ├── tests/ # 测试 ├── package.json └── .env

text

**命名规范:**

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
80% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/nodejs_best_practices.md (reported line 173)May include surrounding context.

md
// PUT /api/users/:id
router.put('/:id', userController.updateUser);

// DELETE /api/users/:id
router.delete('/:id', userController.deleteUser);

module.exports = router;

Chaining Abuse

High
Category
Tool Misuse
Confidence
75% confidence
Finding

Tool calls are chained to bypass individual safety checks or escalate capabilities beyond what any single tool call would allow.

Content

Scanner excerpt · references/security_best_practices.md (reported line 478)May include surrounding context.

bash
# 更新系统
sudo apt update && sudo apt upgrade -y

# 安装防火墙
sudo apt install ufw -y

Static analysis

Detected: suspicious.dynamic_code_execution, suspicious.env_credential_access, suspicious.exposed_secret_literal

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
ai-models/_test_code_quality.py:31

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
web-interface/script.js:61

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
references/security_best_practices.md:376