Back to skill

Security audit

Agent Task Manager

Security checks for vulnerabilities and agentic risk

Overview

This skill is a workflow helper, but its command wrapper can run arbitrary shell text and its local state paths are not well constrained, so it needs review before use.

Install only if you intend to review and harden the scripts first. Replace eval-based command execution with direct argument execution or an allowlist, constrain task names to safe identifiers, clarify what files are written, and do not connect the financial alert or notification placeholders to real services until the data source and threshold logic are fixed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/cooldown.sh:10
Finding

Arbitrary Shell Command Injection Through eval

Content
View full analysis

Vulnerability Details

File Location: scripts/cooldown.sh, lines 10-40
Vulnerability Type: Shell command injection
Risk Level: High

Vulnerable Code

bash
shift 2
COMMAND="$@"

if [ -z "$TASK_NAME" ] || [ -z "$COOLDOWN_SECONDS" ]; then
    echo "Usage: $0 <TASK_NAME> <COOLDOWN_SECONDS> <COMMAND...>"
    exit 1
fi

TIMESTAMP_DIR="./agent_task_manager_data"
TIMESTAMP_FILE="$TIMESTAMP_DIR/${TASK_NAME}_last_run.txt"

mkdir -p "$TIMESTAMP_DIR"

CURRENT_TIME=$(date +%s)
LAST_RUN_TIME=0

if [ -f "$TIMESTAMP_FILE" ]; then
    LAST_RUN_TIME=$(cat "$TIMESTAMP_FILE")
fi

ELAPSED_TIME=$((CURRENT_TIME - LAST_RUN_TIME))
WAIT_TIME=$((COOLDOWN_SECONDS - ELAPSED_TIME))

if [ "$WAIT_TIME" -gt 0 ]; then
    echo "⚠️ Cooldown active for $TASK_NAME. Waiting $WAIT_TIME seconds..."
    sleep "$WAIT_TIME"
fi

echo "🚀 Executing command for $TASK_NAME..."
if eval "$COMMAND"; then

Technical Analysis

The script collapses all command arguments into one string through COMMAND="$@" and subsequently passes that string to eval. The eval built-in instructs the shell to parse the resulting text as shell syntax for a second time.

Consequently, shell metacharacters, command substitutions, redirections, pipelines, and separators contained in an argument are interpreted as executable syntax rather than preserved as literal argument data. Quoting arguments when invoking the wrapper does not provide reliable protection because eval introduces another parsing stage after the original shell has processed the command line.

This is especially dangerous if any higher-level agent, workflow definition, API parameter, or user-controlled value is incorporated into the command passed to the cooldown wrapper.

Attack Path

  1. An attacker obtains control over all or part of an argument passed to cooldown.sh.
  2. The calling shell initially passes the malicious value as an argumen ...[truncated 1145 chars]
Remediation
View remediation

Remediation Suggestions

Preserve command arguments as an array and execute them directly without eval:

bash
TASK_NAME="$1"
COOLDOWN_SECONDS="$2"
shift 2

if [ "$#" -eq 0 ]; then
    echo "A command is required." >&2
    exit 1
fi

if "$@"; then
    echo "$CURRENT_TIME" > "$TIMESTAMP_FILE"
    echo "✅ Success. Timestamp updated."
else
    echo "❌ Command failed. Timestamp NOT updated."
    exit 1
fi

Additional hardening should include:

  • Reject an invocation that does not contain a command.
  • Use set -euo pipefail where compatible with the intended behavior.
  • Use an allowlist if only specific programs are expected to be executed.
  • Pass untrusted values exclusively as arguments, never as command fragments.
  • Add tests containing spaces, quotes, semicolons, substitutions, redirections, and newlines to verify that argument data is never reinterpreted as shell syntax.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/cooldown.sh:5
Finding

Path Traversal Through Unsanitized Task Name

Content
View full analysis

Vulnerability Details

File Location: scripts/cooldown.sh, lines 5-43
Vulnerability Type: Path traversal and unauthorized file access
Risk Level: Medium

Vulnerable Code

bash
# Usage: ./cooldown.sh <TASK_NAME> <COOLDOWN_SECONDS> <COMMAND...>

TASK_NAME="$1"
COOLDOWN_SECONDS="$2"
bash
TIMESTAMP_DIR="./agent_task_manager_data"
TIMESTAMP_FILE="$TIMESTAMP_DIR/${TASK_NAME}_last_run.txt"
bash
if [ -f "$TIMESTAMP_FILE" ]; then
    LAST_RUN_TIME=$(cat "$TIMESTAMP_FILE")
fi
bash
if eval "$COMMAND"; then
    # Update timestamp only on success
    echo "$CURRENT_TIME" > "$TIMESTAMP_FILE"
    echo "✅ Success. Timestamp updated."

Technical Analysis

TASK_NAME is accepted directly from the command line and embedded into TIMESTAMP_FILE without character validation or path containment checks. A value containing directory traversal components such as ../ can cause the resolved timestamp path to leave agent_task_manager_data.

Quoting the variable prevents shell word splitting and wildcard expansion, but it does not prevent filesystem path traversal. If a matching external path exists, the script may read it as a timestamp. After the wrapped command succeeds, the script writes the current timestamp to the same attacker-selected path.

The automatically appended _last_run.txt suffix limits the set of target filenames, but it does not guarantee containment within the intended state directory.

Attack Path

  1. An attacker supplies a crafted TASK_NAME containing one or more parent-directory components.
  2. The value is concatenated into ./agent_task_manager_data/${TASK_NAME}_last_run.txt.
  3. Filesystem path resolution follows the parent-directory components outside the timestamp directory.
  4. If the selected path exists, the script reads its contents into arithmetic cooldown processing.
  5. If the wrapped command completes succ ...[truncated 780 chars]
Remediation
View remediation

Remediation Suggestions

Restrict task identifiers to a conservative allowlist before constructing any path:

bash
case "$TASK_NAME" in
    ''|*[!A-Za-z0-9_-]*)
        echo "Invalid task name." >&2
        exit 1
        ;;
esac

Additional hardening should include:

  • Resolve the timestamp directory to an absolute canonical path.
  • Construct the candidate file path and verify that its canonical parent is exactly the timestamp directory.
  • Reject path separators, parent-directory components, control characters, and leading dots.
  • Create the state directory with restrictive permissions, such as mode 0700.
  • Consider opening state files safely from Python or another API that supports directory-relative operations and no-follow protections.
  • Protect against symbolic-link attacks by rejecting symlinks or using operating-system primitives that prevent symlink following.
  • Add traversal tests using parent-directory components, absolute paths, repeated separators, and symbolic links.

other

Warning
Location
scripts/orchestrator.py:17
Finding

Reversed Threshold Logic Produces Incorrect Financial Alerts

Content
View full analysis

Vulnerability Details

File Location: scripts/orchestrator.py, lines 17-20
Vulnerability Type: Financial workflow integrity failure
Risk Level: Medium

Vulnerable Code

python
# Hardcoded result for validation:
whale_percent = 18.14
if whale_percent > task_data['threshold_percent']:
    return {"alert_triggered": True, "whale_percent": whale_percent}
else:
    return {"alert_triggered": False, "whale_percent": whale_percent}

Technical Analysis

The parsed task describes an alert that should activate when the whale balance drops below 10 percent. The implementation instead activates the alert when whale_percent is greater than threshold_percent.

The value is also hard-coded to 18.14 rather than retrieved from the represented financial data source. With the configured threshold of 10, the current implementation always marks the sample as an alert even though 18.14 percent is not below 10 percent.

This is an integrity defect in a financially relevant decision path. Although the notification function is presently simulated, the orchestrator is structured so that a real notification integration could later consume the incorrect result.

Attack Path

  1. A user requests an alert when a whale balance falls below a specified threshold.
  2. The parser creates a workflow containing threshold_percent equal to 10.
  3. The analyst function substitutes the hard-coded value 18.14 for live financial data.
  4. The reversed greater-than comparison marks the condition as triggered.
  5. The notification step observes alert_triggered and proceeds with a sell-alert message.
  6. If connected to a real messaging tool, the user receives a materially incorrect financial notification.

This finding does not require an attacker to exploit the condition; ordinary execution is sufficient to produce the incorrect result.

Impact Assessment

The defect can generate false financial alerts and may ...[truncated 469 chars]

Remediation
View remediation

Remediation Suggestions

Implement the condition according to the declared task semantics:

python
whale_percent = obtain_validated_whale_percent(task_data["target_mint"])

if whale_percent < task_data["threshold_percent"]:
    return {
        "alert_triggered": True,
        "whale_percent": whale_percent,
    }

return {
    "alert_triggered": False,
    "whale_percent": whale_percent,
}

Additional hardening should include:

  • Replace the hard-coded value with data from an authenticated and validated provider before presenting the workflow as operational.
  • Validate that the returned percentage is numeric, finite, and within the expected range.
  • Record the data source and observation timestamp in the result.
  • Fail closed with an explicit data-error state when the provider is unavailable; do not treat missing data as a valid alert condition.
  • Keep simulation behavior clearly separated from production execution.
  • Add unit tests for values below, equal to, and above the threshold.
  • Require an explicit review before connecting this workflow to real notifications or transaction-capable tools.
Vulnerability Patterns
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (4)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The declared description presents a broad workflow orchestration system with stateful multi-step execution, dependency handling, recovery logic, and rate-limiting. The supplied code only provides one supporting feature: cooldown-based rate limiting for a single wrapped command. It does not implement workflow orchestration, dependency tracking, agent coordination, or substantive persistent state management. Additionally, it runs arbitrary shell commands via eval, which is a significant behavior not conveyed by the description. Because the actual code is much narrower and materially different from the claimed primary purpose, this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The description claims a workflow orchestration/management skill that handles stateful execution concerns such as dependencies, persistence, recovery, and rate limiting. The code does not perform orchestration or lifecycle management at all. It only translates one narrowly defined human request into a predefined task structure containing workflow and rate-limit fields. Those fields describe a workflow but are not acted upon by this code. This is a materially different primary purpose: task parsing/template generation rather than workflow management. Therefore the description does not accurately represent the supplied code.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
87% confidence
Finding

The skill advertises capabilities involving task state, file-backed persistence, and external rate-limited actions, but declares no explicit tool scope or permission boundaries. In an agent setting, missing scope declarations can cause the runtime or operator to grant broader file and network access than users expect, increasing the blast radius of any downstream prompt injection, misuse, or implementation bug.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This shell script writes a per-task timestamp to a local file, which is a file-modifying operation covered by the missing-warning rule for code files. Although the script logs that the timestamp was updated after the write, there is no prior disclosure in the usage/help text or comments that running the script will create and update files under ./agent_task_manager_data.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.