Back to skill

Security audit

Agent Step Sequencer

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed multi-step agent scheduler, but it can automatically relaunch agent work from heartbeat checks and has under-scoped execution controls that users should review carefully.

Install only if you are comfortable with a heartbeat-triggered scheduler that can keep launching your configured agent after a plan is approved. Use a trusted STEP_AGENT_CMD, avoid STEP_RUNNER unless you fully control the path, and avoid non-idempotent steps such as purchases, public posting, or destructive changes until concurrency locking is added.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/step-sequencer-check.py:108
Finding

Concurrent Heartbeat Invocations Can Execute the Same Agent Step Multiple Times

Content
View full analysis

Vulnerability Details

File Location: scripts/step-sequencer-check.py:108-110; related state transition in scripts/step-sequencer-runner.py:141-150
Vulnerability Type: Race condition and missing execution locking
Risk Level: High
Category: T09: Insecure Skill Coding Practices

Vulnerable code in scripts/step-sequencer-check.py:108-110:

python
# PENDING or IN_PROGRESS: invoke runner
save_state(state_path, state)
invoke_runner(state_path, scripts_dir)

Related code in scripts/step-sequencer-runner.py:141-150:

python
now = datetime.now(timezone.utc).isoformat()
step_runs[step_id] = {
    "status": "IN_PROGRESS",
    "tries": tries + 1,
    "lastRunIso": now,
}
state["stepRuns"] = step_runs
save_state(state_path, state)

agent_cmd = get_agent_cmd() + [prompt]

Technical Analysis

The heartbeat check treats IN_PROGRESS steps the same as PENDING steps and unconditionally invokes another runner. No exclusive file lock, process lock, atomic claim operation, execution lease, process identifier, or unique run identifier prevents two check processes from launching the same step concurrently.

The runner writes IN_PROGRESS before launching the configured agent, but this state does not prevent another heartbeat or manual check invocation from starting a second runner. The JSON state file is also loaded and rewritten without synchronization or atomic replacement. Concurrent processes can therefore execute the same instruction and overwrite each other's state updates.

This behavior is especially dangerous because step instructions may authorize externally visible or non-idempotent actions, such as sending messages, making API transactions, modifying files, or deleting resources. The race does not independently grant privileges beyond those already available to the configured agent, but it can duplicate actions using all of that agent's existing permissions.

...[truncated 1840 chars]

Remediation
View remediation

Remediation Suggestions

  1. Serialize state access and execution claims. Acquire an exclusive lock associated with the state file before reading, changing, or launching a step. Keep the lock through the atomic claim operation, and use a separate execution lease if holding the lock for the entire agent run is undesirable.

  2. Do not immediately relaunch active steps. Treat a valid IN_PROGRESS state as “already running.” Relaunch it only when a recorded lease has expired and no matching live execution remains.

  3. Add execution identity fields. Store a cryptographically random runId, start time, lease expiry, and optionally the runner PID. Require each runner to include the same runId when committing DONE or FAILED, rejecting stale updates from older processes.

  4. Use atomic state writes. Write updated JSON to a temporary file in the same directory, flush and synchronize it, and atomically replace the original state file with os.replace. Combine this with locking because atomic replacement alone does not prevent lost updates.

  5. Use compare-and-set semantics. Under the lock, transition only PENDING to IN_PROGRESS. If the observed state or run identifier changed after it was read, abort rather than launching another agent.

  6. Implement stale-run recovery carefully. A heartbeat should recover an IN_PROGRESS step only after a configurable timeout and verification that its lease is stale. Record the recovery reason and allocate a new runId.

  7. Add concurrency tests. Launch multiple check processes against one state file and assert that the agent command runs exactly once, the JSON remains valid, retry counts are accurate, and stale runners cannot overwrite the active run's result.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (25)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill description presents a higher-level planner that waits for confirmation, but the documented behavior also includes immediate orchestration, heartbeat-triggered execution, and external runner invocation. This mismatch is dangerous because operators may grant trust based on the benign-sounding description while the skill actually enables autonomous subprocess execution and persistent state transitions that can continue after the initial interaction.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 21)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 56)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 89)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 115)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 149)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 189)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 206)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 239)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 275)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 321)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 352)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 383)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill explicitly relies on environment variables, file writes, and shell-invoked Python scripts, yet it declares no tool scope or permission boundary. That creates an avoidable trust gap: users and hosts cannot easily assess that the skill can persist state and trigger subprocess-driven automation, increasing the chance of unsafe deployment or over-privileged execution.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script trusts the STEP_RUNNER environment variable to choose which Python file to execute, then runs it with the current interpreter. In agent or automation environments where environment variables can be influenced by upstream components, this becomes arbitrary code execution under the agent's privileges, which is especially dangerous because this skill is explicitly designed to orchestrate multi-step execution.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/step-sequencer-check.py (reported line 35)May include surrounding context.

python
def invoke_runner(state_path: Path, scripts_dir: Path) -> None:
    runner = get_runner_path(scripts_dir)
    if runner.exists():
        subprocess.run(
            [sys.executable, str(runner), str(state_path)],
            cwd=state_path.parent,
        )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/step-sequencer-runner.py (reported line 156)May include surrounding context.

python
agent_cmd = get_agent_cmd() + [prompt]
    try:
        result = subprocess.run(
            agent_cmd,
            capture_output=True,
            text=True,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/step-sequencer-runner.py (reported line 187)May include surrounding context.

python
if not success:
        check_script = get_check_script_path(scripts_dir)
        if check_script.exists():
            subprocess.run(
                [sys.executable, str(check_script), str(state_path)],
                cwd=state_path.parent,
            )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/step-sequencer-runner.py (reported line 196)May include surrounding context.

python
if not success:
        check_script = get_check_script_path(scripts_dir)
        if check_script.exists():
            subprocess.run(
                [sys.executable, str(check_script), str(state_path)],
                cwd=state_path.parent,
            )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 22)May include surrounding context.

python
def run_check(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(CHECK), str(state_path)],
        cwd=state_path.parent,
        env=env,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 192)May include surrounding context.

python
env = os.environ.copy()
        env["STEP_AGENT_CMD"] = "bash -c"

        r = subprocess.run(
            [sys.executable, str(RUNNER), str(state_path)],
            cwd=tmp,
            env=env,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 324)May include surrounding context.

python
env = os.environ.copy()
        env["STEP_AGENT_CMD"] = "bash -c"

        r = subprocess.run(
            [sys.executable, str(RUNNER), str(state_path)],
            cwd=tmp,
            env=env,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 355)May include surrounding context.

python
env = os.environ.copy()
        env["STEP_AGENT_CMD"] = "bash -c"

        r = subprocess.run(
            [sys.executable, str(RUNNER), str(state_path)],
            cwd=tmp,
            env=env,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · test/test_step_sequencer.py (reported line 207)May include surrounding context.

python
def run_runner(state_path: Path, env: dict | None = None) -> subprocess.CompletedProcess:
    env = env or os.environ.copy()
    return subprocess.run(
        [sys.executable, str(RUNNER), str(state_path)],
        cwd=state_path.parent,
        env=env,

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The docstring says 'Does NOT execute work—invokes runner,' implying a non-executing/check-only role. In practice, the script launches another Python program with subprocess.run, which is code execution and has side effects beyond merely reading state.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.