Back to skill

Security audit

adversarial-spec

Security checks for vulnerabilities and agentic risk

Overview

This is a real spec-writing pipeline, but it gives installer and model-driven phases broad code execution and repository mutation authority without enough containment.

Review before installing. Avoid the one-line curl | bash path; install only from pinned, reviewed commits. Treat briefs, repository files, generated spec.md, and provider outputs as untrusted. Run provider CLIs in a sandbox or disposable worktree, require safe permission modes, and do not merge results unless the final diff is limited to expected files such as spec.md.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
Findings (5)

T03 · Remote Payload Retrieval and Execution

Error
Location
SKILL.md:20
Finding

Mutable Remote Installer Is Piped Directly into Bash

Content
View full analysis
Remediation
View remediation
install-v1.1.0.sh' | sha256sum -c - less install-v1.1.0.sh bash install-v1.1.0.sh ``` 7. Ensure release automation prevents existing tags and artifacts from being overwritten. ]]>

T08 · Insecure Dependencies

Error
Location
scripts/install.sh:14
Finding

Unpinned Repositories Are Downloaded and Immediately Imported

Content
View full analysis
/dev/null 2>&1; then echo "detected: running from a checkout ($HERE) — skill already present" else if [ ! -d "$TARGET/$SKILL_DIR" ]; then git clone --depth 1 "$SKILL_URL" "$TARGET/$SKILL_DIR" else echo "skip: $TARGET/$SKILL_DIR already present" fi fi if [ ! -d "$TARGET/$COMMON_REPO" ]; then git clone --depth 1 "$COMMON_URL" "$TARGET/$COMMON_REPO" else echo "skip: $TARGET/$COMMON_REPO already present" fi echo echo "Installed:" echo " $TARGET/$SKILL_DIR (skill)" echo " $TARGET/$COMMON_REPO (dependency, sibling layout required)" echo echo "Sanity check (import adversarial_common + adversarial_spec)..." SKILL_SCRIPTS_DIR="$TARGET/$SKILL_DIR/scripts" if [ -n "$HERE" ] && git -C "$HERE" rev-parse --git-dir >/dev/null 2>&1; then SKILL_SCRIPTS_DIR="$(git -C "$HERE" rev-parse --show-toplevel)/scripts" fi PYTHONPATH="$TARGET/$COMMON_REPO:$SKILL_SCRIPTS_DIR" python3 -c "import adversarial_common, adversarial_spec; print('OK: adversarial_common + adversarial_spec imported')" ``` ### Technical Analysis The installer clones both `adversarial-spec` and `adversarial-common` from their mutable default branches. It does not check out an audited revision, verify signed commits or tags, or compare repository contents against a trusted digest. The installer then places the downloaded paths at the front of `PYTHONPATH` and imports both packages. Python imports can execute arbitrary module-level code, so the ...[truncated 1298 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/adversarial_spec.py:617
Finding

Model-Authored Acceptance Directives Are Executed Without Local Security Validation

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/phases/phase_write.py:47
Finding

Agent Can Modify Arbitrary Repository Files That Are Then Committed and Merged

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
SKILL.md:170
Finding

Recommended Provider Configuration Disables Permission Checks for Untrusted Input

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (73)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The manifest describes a spec-writing skill, but the document also includes installation, bootstrap, repository manipulation, and remote-fetch behavior not reflected in the declared purpose. This mismatch hides materially risky behavior from reviewers and policy engines, which can lead to under-scoped approvals for a skill that performs broader system actions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The manifest describes a spec-writing skill, but the document also includes installation, bootstrap, repository manipulation, and remote-fetch behavior not reflected in the declared purpose. This mismatch hides materially risky behavior from reviewers and policy engines, which can lead to under-scoped approvals for a skill that performs broader system actions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The manifest describes a spec-writing skill, but the document also includes installation, bootstrap, repository manipulation, and remote-fetch behavior not reflected in the declared purpose. This mismatch hides materially risky behavior from reviewers and policy engines, which can lead to under-scoped approvals for a skill that performs broader system actions.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The manifest describes a spec-writing skill, but the document also includes installation, bootstrap, repository manipulation, and remote-fetch behavior not reflected in the declared purpose. This mismatch hides materially risky behavior from reviewers and policy engines, which can lead to under-scoped approvals for a skill that performs broader system actions.

Content

No source excerpt is available for this finding.

Chaining Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

Chaining curl output directly into bash combines remote content retrieval and execution into a single unsafe step, eliminating any opportunity for inspection or integrity validation. In a skill that already normalizes shell-heavy behavior, this pattern significantly raises supply-chain compromise risk and could lead to full code execution on the host.

Content

Scanner excerpt · SKILL.md (reported line 24)May include surrounding context.

md
Requires the `adversarial-common` sibling repo (shared engine). One-line install:

curl -fsSL https://raw.githubusercontent.com/chpomob/adversarial-spec/main/scripts/install.sh | bash

or, from an existing checkout:

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Allowing an externally supplied --dev-cmd means the pipeline may invoke arbitrary local commands as the writer role. In a skill that already has shell and filesystem interactions, untrusted or loosely validated provider command configuration can become a command-execution vector or route sensitive project data to unintended third-party tools.

Content

Scanner excerpt · SKILL.md (reported line 48)May include surrounding context.

bash
python3 scripts/adversarial_spec.py \
  --brief <file>           # brief file (default: stdin)
  --dev-cmd <cmd>          # default: pi --provider zai --model glm-5.2
  --review-cmd <cmd>       # default: pi --provider deepseek --model deepseek-v4-pro
  --workdir <dir>          # default: .
  --max-loops <N>          # default: 2

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

The --review-cmd option similarly enables execution of external provider commands that may process repository content and findings. Without strict validation and trust boundaries, this can leak sensitive data to external services or permit arbitrary command execution under the guise of model selection.

Content

Scanner excerpt · SKILL.md (reported line 49)May include surrounding context.

md
python3 scripts/adversarial_spec.py \
  --brief <file>           # brief file (default: stdin)
  --dev-cmd <cmd>          # default: pi --provider zai --model glm-5.2
  --review-cmd <cmd>       # default: pi --provider deepseek --model deepseek-v4-pro
  --workdir <dir>          # default: .
  --max-loops <N>          # default: 2
  --feature <name>         # default: from brief filename

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly acknowledges that its pipeline can commit arbitrary repository changes beyond spec.md. In a toolchain that stages all changes, a prompt-influenced writer or compromised phase could persist unrelated or malicious modifications to source files, turning a documentation workflow into a code-changing one.

Content

No source excerpt is available for this finding.

Chaining Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

The '| bash' pattern is a classic unsafe execution chain because it eliminates the user's opportunity to inspect what was downloaded before execution. In this skill context, the installer also performs git clones and a Python import check, so a compromised upstream script could chain immediately into broader filesystem and code-execution effects.

Content

Scanner excerpt · scripts/install.sh (reported line 4)May include surrounding context.

sh
#!/usr/bin/env bash
# Install adversarial-spec with its single dependency: adversarial-common (sibling layout).
# Usage:
#   Bootstrap (from anywhere):  curl -fsSL https://raw.githubusercontent.com/chpomob/adversarial-spec/main/scripts/install.sh | bash
#   From a checkout:            bash scripts/install.sh [TARGET_DIR]
# Target layout (siblings required):
#   <TARGET>/adversarial-spec            (this skill)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill documents shell, file, environment, and git-capable behavior but declares no explicit tool scope or permission boundaries. That omission increases the chance an orchestrator grants broader capabilities than necessary, making misuse or prompt-influenced side effects harder to constrain.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

Piping a remotely fetched script directly into bash executes unreviewed network content immediately on the user's machine. If the remote source, network path, or hosting account is compromised, arbitrary code runs with the user's privileges with no integrity verification or review step.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 36)May include surrounding context.

text
PHASE 0 ──→ GIT SETUP (branch, stash, init)
PHASE 1 ──→ WRITE  (spec-writer reads brief, writes spec.md)
PHASE 2 ──→ CHALLENGE (spec-challenger critiques for gaps/contradictions)
PHASE 3 ──→ REVISE (spec-writer amends spec.md per findings)
PHASE 4 ──→ VERIFY (spec-challenger checks findings resolved)

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
79% confidence
Finding

The manifest describes a skill that takes a brief and produces a structured spec.md through an adversarial writing pipeline. However, the documented workflow requires checking CONTRIBUTING rules, existing PR reviewer comments, upstream branch topology, and gap analysis before writing the spec, which adds broader code-review and repository-analysis responsibilities not clearly justified by a specification-writer role.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

These lines state that all pipeline-internal text must be English and instruct users to write the brief in English unless they explicitly instruct otherwise. This imposes a language policy on skill use rather than offering a user choice, which matches the locale/language policy violation criteria.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 115)May include surrounding context.

md
## Language discipline

All pipeline-internal text (spec, commit messages, findings, JSON) is **English**.
User-facing conversation stays in the conversation's language. Write the brief in
English unless the user explicitly instructs otherwise. A spec written in French or
other non-English languages will break later pipeline stages (plan, code loop) that
expect English section headings and identifiers.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The text first states precedence as provider-config > explicit role flags > hardcoded fallback commands, which implies explicit flags outrank other non-registry sources. But later it says that without a registry, each role resolves from the explicit flag, then environment variables, then hardcoded fallbacks, creating an inconsistent description of intent and actual resolution order.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The pitfalls section states provider selection is purely configuration-driven, which contradicts earlier documentation that explicit commands, environment variables, and hardcoded fallback commands may be used. This creates a direct mismatch in stated behavior versus documented operational paths.

Content

No source excerpt is available for this finding.

Context Window Stuffing

Medium
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill attempts to fill the context window with filler content, displacing legitimate instructions and safety constraints. This can degrade agent performance or bypass safety boundaries.

Content

Scanner excerpt · references/adversarial-2026-07-14-pre-publication-review.md (reported line 29)May include surrounding context.

md
| A7 | nit | `_retrospective/ISSUES.md:3` | Stale claim that failures are auto-appended to this file — actually goes to `<out_dir>/ISSUES.md` now |
| A8 | nit | `scripts/adversarial_spec.py:78` | `_ensure_ids` dedup generates awkward compound ids like `S1-3-3` |
| A9 | nit | `scripts/phases/phase_verify.py:55` | Prompt hard-codes `git diff HEAD~1..HEAD` — wrong if writer CLI made multiple commits |
| A10 | nit | `scripts/phases/phase_challenge.py:44` | Full spec embedded in prompt with no size guard — large specs can overflow context window |

## Inspector Findings (Codex GPT-5.6-Sol)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The notes explicitly document that the wrapper defaults to a dangerous permission-skipping mode and only briefly mention using --no-danger for untrusted input. In a security-sensitive, adversarial skill context, documenting an unsafe default without strongly requiring the safer flag increases the chance that operators will run untrusted model prompts with elevated permissions, enabling unintended filesystem or command effects.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest describes a skill that takes a brief and produces a structured spec.md in a git-aware workflow. This file additionally implements an opt-in external research pipeline that gathers authoritative external context and feeds it into the writer, which is a distinct network/external-information capability not stated in the manifest description.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

When --deep-research is enabled, the code sends the full brief text to an external research provider. Briefs may contain proprietary plans, credentials, internal URLs, or other sensitive context, and this transfer occurs without any visible consent prompt, redaction step, or in-file user warning at the point of use.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This shell script clones remote GitHub repositories into the user's skills directory, which both transmits network metadata and writes files locally. Although there are status echoes, there is no confirmation prompt or explicit warning in the script comments/output that running it will fetch code from the network and modify the local filesystem.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This code performs a repository commit after modifying spec.md, which is a safety-relevant file/system state change. While the module docstring mentions that the file is edited and committed, there is no user-facing confirmation prompt, log/print disclosure, or inline warning near the operation itself in this code path.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · spec.md (reported line 18)May include surrounding context.

md
- file: adversarial_common/report.py
    description: "Add quota metadata to the final report (which provider was used, quota state at decision time)."
  - file: adversarial_common/jsonio.py
    description: "No changes needed — already handles JSON read/write used by check-ai-quota.py output."
---

# Quota-Aware Provider Registry

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The spec states that all error messages and report fields must be in English, which imposes a language policy on user-facing output at least for errors. This is a natural-language policy concern because it forces a specific language rather than offering a user choice or clearly limiting the behavior to a justified region-specific context.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.