Back to skill

Security audit

spex

Security checks for vulnerabilities and agentic risk

Overview

This Spex skill is a real SDLC automation tool, but it includes under-scoped dependency installation, repository-controlled hook execution, and broad code/git mutation workflows that need Review before install.

Review this carefully before installing. Use it only in repositories you trust, avoid running init in a privileged or shared Python environment, prefer an isolated virtual environment or --skip-deps/manual dependency install, and do not allow repository-provided .spex.toml or hooks to run unless you have inspected them. Treat specs from pull requests or unfamiliar repos as untrusted because they can steer code-writing agents and later commits.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Tool Hijacking and SpoofingModifies or replaces tools so legitimate-looking calls execute attacker logic
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T07 · Tool Hijacking and Spoofing

Error
Location
scripts/hooks.py:18
Finding

Automatic Execution of Repository-Controlled Lifecycle Hooks

Content
View full analysis
list[Path]: """Discover .spex.toml files in priority order (highest first).""" config_file = _spex_config_file_override or os.environ.get("SPEX_CONFIG_FILE") if config_file: p = Path(config_file).expanduser().resolve() if not p.is_file(): raise FileNotFoundError( f"Specified spex config file does not exist: {config_file}" ) return [p] candidates: list[Path] = [] visited: set[Path] = set() if main_worktree is not None: start = main_worktree.resolve() else: start = Path(workdir).resolve() if workdir else Path.cwd().resolve() current = start while True: toml_path = current / ".spex.toml" if toml_path.is_file(): candidates.append(toml_path) visited.add(toml_path.resolve()) parent = current.parent if parent == current: break current = parent ``` ```python # scripts/common.py:318-324 def _resolve_hook_roots(workdir=None): """Return hook root paths in priority order (highest first). Builds from all resolved spex_roots: /hooks/ for each. """ ctx = get_project_context(workdir) return [Path(sr) / "hooks" for sr in ctx.spex_roots] ``` ```python # scripts/hooks.py:18-30 def find_hook(hook_name: str, workdir=None) -> Path | None: """Find the first executable hook file in priority order.""" for root in _resolve_hook_roots(workdir): candidate = root / hook_name if candidate.is_file() and os.access(candidate ...[truncated 3454 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
templates/apply-one-task.md:10
Finding

Repository Specification Content Is Promoted to Authoritative Agent Instructions

Content
View full analysis
- **Route:** - IF `$resume_phase` is `review` -> skip Phases 4–5 -> Phase 6 in main context. Do NOT set `$commit_sha` from `HEAD` here — Phase 6 **6-entry** resolves it (prefer review file's `commit_sha`) - IF `$resume_phase` is `implement` -> launch sub-agent for Phases 4–5 only (implement + first commit). Instruct it to follow Phases 4–5 of this command exactly. Pass `$prompt` as Phase 4 guide, plus `$current_task_id` and `$spec_name`. Implementation prompt must NOT create the commit — sub-agent runs Phase 5 (`apply-commit`) after implementation. ON_FAIL: report + retry **once**; IF still fails -> STOP. After OK -> Phase 6 in main context ### Phase 4: Execute Task - Using `$prompt` as guide (from Phase 3 — do **not** re-run `prompt apply-one-task`), implement current task. Follow rendered prompt precisely (spec, completed steps, task description, guidelines) - Deliver production code + tests together - IF no file changes -> report issue -> STOP - Do NOT create git commit here — Phase 5 handles commit ``` ```markdown Act as a senior software engineer focused on incremental, high-quality implementation. Your task is to implement exactly one development step from the plan, producing production-ready code with tests. Analyze the specification, review completed work for context and consistency, then implement the current task precisely as described. ## Specification The following is the full specification. Use it as the authoritative reference for requirements, design decisions, and constraints. Your implementation must conform to this spe ...[truncated 3142 chars]
Remediation
View remediation
` and `` before interpolation. 5. Validate requested file paths and operations against an allowlist derived from the selected task. 6. Require user confirmation for operations outside the repository, network access, dependency changes, credential access, or destructive commands. 7. Restrict the sub-agent's writable filesystem scope to the current repository and, where practical, to task-relevant paths. 8. Separate requirement extraction from execution: first produce a normalized task plan, then have a policy layer validate it before granting implementation tools. 9. Review the final diff against the normalized task scope before staging or committing. 10. Warn users when applying specifications obtained from an untrusted branch, pull request, or repository. ]]>

T08 · Insecure Dependencies

Error
Location
scripts/init.py:212
Finding

Unverified Wheels Are Extracted Directly into Python Site-Packages Without Path Validation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (79)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a high-level skill that manages requirements, design, implementation, and submission through multiple /spex commands. The supplied code chunk does not implement any of those workflow capabilities. Instead, it provides a small utility class wrapping Python's argparse for command-line parsing. While this could be a supporting component of a larger Spex system, this specific code chunk's actual behavior is materially different from the declared primary purpose, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a broad SDLC management skill with a specific set of supported commands: create, modify, apply, apply-one-step, merge, archive, and init. The supplied code instead implements a separate spex open command for locating a spec directory, opening it in the OS file browser, or executing a user-provided command within that directory. While this may be adjacent to Spex tooling, it is not one of the declared commands and adds a significant undeclared capability: arbitrary command execution through subprocess.run(..., shell=True, cwd=path). That makes the code's actual purpose and capabilities materially different from the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The description presents a broad Spex skill that manages the full SDLC with commands like create, modify, apply, apply-one-step, merge, archive, and init. The supplied code chunk is narrower and different in focus: it is specifically a prompt/template renderer (spex prompt) for selected subcommands. It prepares metadata from spec files, todo lists, project context, and review files, then renders Jinja2 templates. That general behavior is related to Spex workflow support, so it is not wholly unrelated; however, the actual primary purpose is prompt generation/orchestration support, not end-to-end SDLC management. There is also an undeclared side effect: modify-spec --remove-undone and modify-todo can rewrite todo.json to keep only completed tasks. Finally, the implemented command set materially differs from the declared one, with review/fix/todo-specific commands present and several declared commands absent from this code chunk. These differences are substantial enough to count as a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a broad Spex skill for managing the full SDLC with commands like create, modify, apply, merge, archive, and init. The supplied code chunk does not implement those lifecycle functions; instead, it is narrowly focused on review tracking for individual spec steps. It manages review-step-N.json files, tracks findings, severities, categories, completion state, review rounds, and commit SHAs, and exposes a distinct 'spex review-helper' interface with subcommands such as init, append, edit, bump-round, set-commit, status, show, list/get aliases, and next. While this may support a larger SDLC system, the actual code's primary purpose and exposed capabilities are materially different from the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The supplied code chunk is narrowly focused on showing information about an existing spec: resolving/selecting a spec, formatting its specification and TODO content, and paging the output. The declared description lists supported commands as create, modify, apply, apply-one-step, merge, archive, and init, but does not mention a show/display command. Because the code's concrete capability is an undeclared command and its immediate purpose is informational display rather than any of the listed lifecycle management operations, this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a high-level Spex skill that manages the full SDLC through a specific set of /spex commands. The supplied code chunk instead implements a standalone command-line utility named 'spex todo-helper' for manipulating todo files in JSON or XML. Its subcommands are validate, append, edit, remove, show, xml2json, and json2xml. While it appears related to Spex data structures and spec directories, it does not implement the advertised lifecycle-management commands or end-to-end SDLC behavior. The code’s actual capabilities are materially narrower and different in focus, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The supplied code does not implement SDLC orchestration features like requirements analysis, design, incremental implementation, merge/submission, archive, or init workflows. Instead, it is a narrowly scoped maintenance script for version management. This is a materially different primary purpose from the declared description. It also exposes a 'spex version' command with --check and --bump behaviors that are not listed in the declared supported commands. While this script may be part of the broader skill's internal tooling, the declared description does not accurately represent this chunk's actual behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

Using generic aliases like do and go makes accidental command matching much more likely because those words are common in ordinary conversation. Combined with routing to implementation-oriented commands like apply, this can cause unintentional activation of workflows with side effects, especially in chat contexts where users are not attempting to invoke the skill precisely.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 43)May include surrounding context.

md
Command file paths are relative to this `SKILL.md` directory.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This code implements a custom package installer that resolves packages from PyPI, downloads wheels, and extracts their contents directly into site-packages. That bypasses normal package-manager safety controls, performs no cryptographic verification, and writes importable code into the interpreter environment, creating a severe supply-chain and arbitrary code execution risk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

Adding shell command execution to a command whose stated role is opening directories creates an unnecessary and risky execution surface. In an agent skill, such hidden capability is more dangerous because other automation may call the tool based on its name/description and accidentally hand attackers a convenient execution path.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

This is a classic tool-parameter-abuse issue: a parameter intended as a convenience (--run) becomes a direct arbitrary command-execution channel. Because the command is part of an agent skill, adversarial instructions or unsafe chaining could abuse this parameter to run attacker-chosen shell commands on the host.

Content

Scanner excerpt · scripts/open.py (reported line 40)May include surrounding context.

python
path: Path string of the working directory.
        command: Shell command string to execute.
    """
    result = subprocess.run(command, shell=True, cwd=path)
    sys.exit(result.returncode)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
96% confidence
Finding

This duplicate finding identifies the same underlying issue: user-influenced command construction for an external tool with shell=True. Because this skill is part of an SDLC automation workflow, abuse is more dangerous than in a standalone local script since it may execute in privileged developer or CI-like environments.

Content

Scanner excerpt · scripts/show.py (reported line 27)May include surrounding context.

python
return
    pager = os.environ.get("PAGER", "less -R")
    try:
        proc = subprocess.Popen(pager, shell=True, stdin=subprocess.PIPE)
        proc.communicate(input=text.encode())
    except (BrokenPipeError, OSError):
        print(text)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
96% confidence
Finding

This duplicate finding identifies the same underlying issue: user-influenced command construction for an external tool with shell=True. Because this skill is part of an SDLC automation workflow, abuse is more dangerous than in a standalone local script since it may execute in privileged developer or CI-like environments.

Content

Scanner excerpt · scripts/show.py (reported line 27)May include surrounding context.

python
return
    pager = os.environ.get("PAGER", "less -R")
    try:
        proc = subprocess.Popen(pager, shell=True, stdin=subprocess.PIPE)
        proc.communicate(input=text.encode())
    except (BrokenPipeError, OSError):
        print(text)

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · scripts/version.py (reported line 116)May include surrounding context.

python
logger.error("Error: invalid semver format: %s", new_version)
        return False

    # Update skill's pyproject.toml (authoritative)
    if not _write_pyproject_version(_SKILL_PYPROJECT, new_version):
        logger.error("Error: could not find version in skills/spex/pyproject.toml")
        return False

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · scripts/version.py (reported line 125)May include surrounding context.

python
if _ROOT_PYPROJECT.is_file():
        _write_pyproject_version(_ROOT_PYPROJECT, new_version)

    # Update SKILL.md
    content = _SKILL_MD.read_text(encoding="utf-8")
    new_content = re.sub(
        r"^(version:\s*).+",

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill advertises routing into multiple command SOPs that may read, write, or invoke shells, but it declares no explicit tool scope or permissions boundary in this entrypoint file. That creates unnecessary ambiguity about what capabilities the skill may exercise once downstream command files are loaded, increasing the chance of over-privileged execution or unsafe tool access.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill permits free-form intent inference when the first argument does not match a known route, allowing arbitrary natural-language text to trigger command execution. In a development skill that may lead to reading files, writing files, or other side effects via downstream SOPs, broad inference increases the risk of accidental or prompt-manipulated activation without a deliberate command.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The free-form intent rules are broad and provide only positive matching examples, with no exclusions or examples of ambiguous language that should not trigger routing. This makes prompt injection and accidental invocation easier because loosely related requests may be mapped to powerful commands without sufficiently robust disambiguation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This markdown file describes archive and restore behavior that moves specs between specs/ and archives/, including an --all-projects mode and a --force bypass, but it does not explicitly warn users that these actions modify project files. Because the operation affects user data and project state, the skill description should disclose that it will move files and that --force can override a safety check.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The instructions explicitly require the name and description fields to be in English. This is a natural-language locale policy constraint, while elsewhere the spec content is generated in the user's language, and no opt-in or justification is provided for the English-only requirement.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The command performs a repository-changing submit action by merging a feature branch or creating a PR, but the instructions do not require an explicit confirmation or a clear warning before executing the state-changing operation. In an agent-driven workflow, this increases the risk of unintended merges, accidental PR creation, or submission of incomplete work if the command is invoked with the wrong spec or auto-selection chooses an unexpected target.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · references/compact-sop-style.md (reported line 210)May include surrounding context.

bash
   npx markdownlint-cli2 \
     "skills/spex/SKILL.md" \
     "skills/spex/commands/*.md" \
     "skills/spex/references/*.md"

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The module and CLI description say this command archives completed specs, but the code also provides a reverse operation via restore_single_spec and the --restore flag. Restoration is a materially different file-moving operation than the manifest-described archive behavior, so the command's actual behavior exceeds its stated description.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.