Back to skill

Security audit

Universal Agent

Security checks for vulnerabilities and agentic risk

Overview

This skill is a broad automation runner that can direct an AI to create and run commands or code with your machine's permissions, with weak boundaries and unsafe optional setup guidance.

Install only if you intentionally want a high-authority automation skill that may run AI-generated shell commands or Python code. Use it in a disposable sandbox or restricted workspace, avoid confidential files and credentials, do not enable dangerous mode, do not follow the curl-to-sh installer pattern without independent verification, and review every generated command or script before execution.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (5)

T03 · Remote Payload Retrieval and Execution

Error
Location
references/README.md:244
Finding

Unverified Remote Installer Is Piped Directly Into a Shell

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/universal_agent.py:1418
Finding

Untrusted LLM and Bridge Responses Are Executed as Arbitrary Shell Commands

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/universal_agent.py:953
Finding

Dynamically Generated Python Is Executed Without Isolation

Content
View full analysis
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
scripts/universal_agent.py:1199
Finding

Untrusted Persistent Memory Is Injected Into Future LLM Decision Context

Content
View full analysis
bool: path = filepath or self.memory_file if os.path.exists(path): try: with open(path, 'r', encoding='utf-8') as f: loaded = json.load(f) old_env = self.context.get("environment", {}) self.context = loaded self.context["environment"] = old_env return True except Exception as e: print(f"Memory load failed: {e}") return False return False ``` The loaded content is then supplied to the command-generating LLM: ```python self.memory = ContextManager(memory_file=memory_file) self.memory.load() self.memory.update_cwd() ``` ```python decision = self.brain.think(task, self.memory.get_context_string()) ``` The context is appended to the system prompt: ```python if context: system_prompt += f"\n\n--- Current context ---\n{context}\n" ``` ### Technical Analysis The memory file can contain task history, learned knowledge, summaries, and variables. Its content is trusted as a complete context object after only JSON parsing. There is no strict schema validation, integrity check, origin tracking, ownership check, or separation between untrusted data and instructions. Because the default file is relative to the current directory, a repository or other writable working directory can supply a pre-created `universal_agent_memory.json`. Poisoned entries are serialized into the LLM's system-level ...[truncated 1334 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/universal_agent.py:564
Finding

Sensitive Environment, Task, Memory, and Execution Data Can Be Sent to Arbitrary LLM Endpoints

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (58)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill description prominently markets automated command/script execution but does not provide an upfront warning that it may alter files, system state, network resources, or hardware. Users and orchestrating agents may therefore treat it like a normal helper skill rather than a high-risk executor, increasing the likelihood of unsafe use.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The activation guidance is broad enough to match many ordinary requests such as 'help me do X automatically' or 'generate code and run it,' which can cause the skill to be selected in situations where users did not knowingly consent to autonomous execution. Because this skill generates and executes commands/scripts, overbroad invocation materially increases the chance of unintended code execution and system changes.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The keyword trigger '帮我做XX' is extremely vague and could match a wide range of everyday assistance requests. In a skill that can generate arbitrary scripts and execute shell commands, such vague matching creates a high risk of accidental invocation and execution without meaningful user awareness.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 90)May include surrounding context.

md
python scripts/universal_agent.py --run "task description"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 93)May include surrounding context.

md
python scripts/universal_agent.py --run "task description"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 96)May include surrounding context.

md
python scripts/universal_agent.py --run "task description"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 109)May include surrounding context.

md
python scripts/universal_agent.py --run "task description"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 112)May include surrounding context.

md
python scripts/universal_agent.py --run "task description"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 173)May include surrounding context.

md
python scripts/universal_agent.py --run "task description"

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

Inline Simulation mode explicitly states that safety, retry, and memory protections do not apply, yet it lacks a corresponding high-visibility warning about the resulting risk to data and system integrity. This is especially dangerous because the surrounding documentation encourages agents to simulate the workflow using native capabilities, potentially bypassing even the script's limited guardrails.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

This duplicate finding again points to documented destructive command examples, which are not themselves malicious content. The real security concern is that the skill's safety posture appears to center on matching known bad strings and then prompting, an approach that is easy to bypass with equivalent commands, scripting indirection, or platform-specific variants.

Content

Scanner excerpt · SKILL.md (reported line 193)May include surrounding context.

md
| Level | Examples | Handling |
|-------|----------|----------|
| 🔴 High | `rm -rf /`, `format C:` | **Forced confirmation required** |
| 🟡 Medium | `pip uninstall`, `sudo` | Warning prompt |
| 🟢 Low | `ls`, `cat`, `python script.py` | Direct execution |

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

This duplicate finding again points to documented destructive command examples, which are not themselves malicious content. The real security concern is that the skill's safety posture appears to center on matching known bad strings and then prompting, an approach that is easy to bypass with equivalent commands, scripting indirection, or platform-specific variants.

Content

Scanner excerpt · SKILL.md (reported line 193)May include surrounding context.

md
| Level | Examples | Handling |
|-------|----------|----------|
| 🔴 High | `rm -rf /`, `format C:` | **Forced confirmation required** |
| 🟡 Medium | `pip uninstall`, `sudo` | Warning prompt |
| 🟢 Low | `ls`, `cat`, `python script.py` | Direct execution |

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

The documentation includes an option to skip safety confirmations, which normalizes bypassing safeguards around command execution. Even though it is labeled 'not recommended,' exposing a documented path to disable confirmations materially lowers resistance to destructive or unintended actions in a high-risk execution skill.

Content

Scanner excerpt · SKILL.md (reported line 228)May include surrounding context.

md
# Optional: skip safety confirmations (not recommended)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/README.md (reported line 192)May include surrounding context.

md
- 无需人工干预

### 🔒 安全机制
- 高危命令检测(rm -rf /, format C:, etc.)
- 危险操作强制确认
- 中危操作警告提示
- 可选的危险模式(跳过确认)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/README.md (reported line 192)May include surrounding context.

md
- 无需人工干预

### 🔒 安全机制
- 高危命令检测(rm -rf /, format C:, etc.)
- 危险操作强制确认
- 中危操作警告提示
- 可选的危险模式(跳过确认)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · references/README.md (reported line 192)May include surrounding context.

md
- 无需人工干预

### 🔒 安全机制
- 高危命令检测(rm -rf /, format C:, etc.)
- 危险操作强制确认
- 中危操作警告提示
- 可选的危险模式(跳过确认)

External Script Fetching

High
Category
Supply Chain
Confidence
97% confidence
Finding

The README recommends piping a remotely fetched install script directly into sh. That pattern removes inspection, signature verification, and integrity checks, so a compromised host, CDN, or man-in-the-middle position could lead to arbitrary code execution on the user's machine.

Content

Scanner excerpt · references/README.md (reported line 248)May include surrounding context.

md
# 安装 Ollama
# Windows: https://ollama.ai/download
# Mac: brew install ollama
# Linux: curl -fsSL https://ollama.ai/install.sh | sh

# 运行本地模型
ollama run llama3    # 推荐,7B参数

Chaining Abuse

High
Category
Tool Misuse
Confidence
97% confidence
Finding

The shell pipeline to sh is an explicit command-chaining pattern that executes unverified remote content immediately. In the context of a skill centered on automated command execution, normalizing this pattern materially increases the likelihood of unsafe operator behavior and remote code execution.

Content

Scanner excerpt · references/README.md (reported line 248)May include surrounding context.

md
# 安装 Ollama
# Windows: https://ollama.ai/download
# Mac: brew install ollama
# Linux: curl -fsSL https://ollama.ai/install.sh | sh

# 运行本地模型
ollama run llama3    # 推荐,7B参数

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

This is the same underlying issue: subprocess.Popen is invoked with shell=True on dynamically generated command text. In a universal agent that advertises automatic command generation and execution, this creates a direct remote-code-execution primitive via prompt injection, malicious tasks, or compromised bridge inputs.

Content

Scanner excerpt · scripts/universal_agent.py (reported line 885)May include surrounding context.

python
creationflags=subprocess.CREATE_NO_WINDOW if os.name == 'nt' else 0
                )
            else:  # Linux/macOS
                process = subprocess.Popen(
                    command,
                    shell=True,
                    stdout=subprocess.PIPE,

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

This is the same underlying issue: subprocess.Popen is invoked with shell=True on dynamically generated command text. In a universal agent that advertises automatic command generation and execution, this creates a direct remote-code-execution primitive via prompt injection, malicious tasks, or compromised bridge inputs.

Content

Scanner excerpt · scripts/universal_agent.py (reported line 885)May include surrounding context.

python
creationflags=subprocess.CREATE_NO_WINDOW if os.name == 'nt' else 0
                )
            else:  # Linux/macOS
                process = subprocess.Popen(
                    command,
                    shell=True,
                    stdout=subprocess.PIPE,

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The init method docstring begins at L1278 and never closes before assignments and component construction lines. The comments and surrounding class documentation describe a fully functional orchestrator that wires brain, executor, and memory together, but the shown code places that logic inside the docstring, contradicting the claimed operational behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The agent sends user tasks, execution output, and serialized environment/context to external LLM APIs, including current directory, Python path, task history, variables, and learned knowledge, without an explicit warning or consent flow. This can expose sensitive local metadata, command output, secrets present in context, or proprietary data to third-party services.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The skill advertises capabilities spanning shell, file read/write, environment access, and network use, but it does not declare any explicit tool scope or permission boundaries. In a universal-agent context, that omission is dangerous because downstream agents may invoke it with broad ambient privileges, enabling generated code or commands to operate far beyond user intent.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
96% confidence
Finding

The skill is explicitly designed to auto-execute LLM-generated commands or scripts, which delegates security-sensitive decisions to probabilistic model output. In context, this is especially dangerous because the skill also supports auto-fixing and retrying, increasing the likelihood that it will persist in attempting actions until something runs.

Content

Scanner excerpt · SKILL.md (reported line 7)May include surrounding context.

md
description: >
  This skill should be used when the user needs to execute tasks through a complete 
  automated workflow: understand natural language intent, dynamically generate commands 
  or Python scripts, auto-execute them via shell, and summarize results. This is a 
  minimal universal agent implementation based on "LLM brain + command executor limbs" 
  architecture. Use this when users say things like "help me do X automatically", 
  "generate code and run it", "universal agent", "execute task end-to-end", or when

Unrestricted Tool Access

Medium
Category
Excessive Agency
Confidence
98% confidence
Finding

The skill explicitly frames its executor as able to 'Execute any command' and control software and hardware, indicating effectively unrestricted tool access. In a system driven by LLM-generated decisions, this creates a direct path from prompt interpretation errors or malicious input to arbitrary command execution, persistence, data loss, or physical-world effects.

Content

Scanner excerpt · SKILL.md (reported line 29)May include surrounding context.

md
│ Auto-generate code/command
         ↓
  ┌─────────────────┐
  │ Command Executor │ Execute any command, control software & hardware
  │ (Limbs)          │
  └───────┬─────────┘
          │ Actual execution

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/universal_agent.py:943