Back to skill

Security audit

adversarial-code-loop

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its code-review purpose, but its installer and references include high-impact unsafe guidance such as mutable remote execution, broad agent privileges, credential copying, and GitHub secret-scanning bypasses.

Review before installing. Do not use the one-line curl-to-bash installer; install only from pinned, inspected commits or releases. Run loops only on backed-up disposable worktrees, approve external providers explicitly, and treat specs, manifests, and test commands as executable trust boundaries. Do not let an agent copy API keys between local credential stores or bypass GitHub secret scanning without a human security-owner decision.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
Findings (6)

T03 · Remote Payload Retrieval and Execution

Error
Location
SKILL.md:36
Finding

Mutable Remote Installation Script Is Piped Directly into Bash

Content
View full analysis
Remediation
View remediation
' install.sh | sha256sum -c - less install.sh bash install.sh ``` 7. Protect release creation with signed tags, mandatory review, and restricted repository permissions. ]]>

T08 · Insecure Dependencies

Error
Location
scripts/install.sh:10
Finding

Unpinned Repositories Are Cloned and Imported During Installation

Content
View full analysis
/dev/null 2>&1; then echo "detected: running from a checkout ($HERE) — skill already present" else if [ ! -d "$TARGET/$SKILL_DIR" ]; then git clone --depth 1 "$SKILL_URL" "$TARGET/$SKILL_DIR" else echo "skip: $TARGET/$SKILL_DIR already present" fi fi if [ ! -d "$TARGET/$COMMON_REPO" ]; then git clone --depth 1 "$COMMON_URL" "$TARGET/$COMMON_REPO" else echo "skip: $TARGET/$COMMON_REPO already present" fi ``` ```bash echo "Sanity check (import adversarial_common + adversarial_loop_v4)..." SKILL_SCRIPTS_DIR="$TARGET/$SKILL_DIR/scripts" if [ -n "$HERE" ] && git -C "$HERE" rev-parse --git-dir >/dev/null 2>&1; then SKILL_SCRIPTS_DIR="$(git -C "$HERE" rev-parse --show-toplevel)/scripts" fi PYTHONPATH="$TARGET/$COMMON_REPO:$SKILL_SCRIPTS_DIR" python3 -c "import adversarial_common, adversarial_loop_v4; print('OK: adversarial_common + adversarial_loop_v4 imported')" ``` ### Technical Analysis The installer performs shallow clones of two repositories without specifying a commit hash or signed tag. The checked-out content therefore depends on each repository's mutable default branch at installation time. The risk is amplified by the final Python command. Importing `adversarial_common` and `adversarial_loop_v4` can execute package initialization and module-level Python statements. Thus, downloaded dependency code is not merely stored—it is loaded into a Python interpreter during installation. There is no lockfile, expected commit check, checksum, signature ...[truncated 950 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
references/pi-auth-setup.md:45
Finding

Instructions Extract and Duplicate Plaintext API Keys from Another Credential Store

Content
View full analysis
"}}`. Extract the key from Hermes via `grep DEEPSEEK_API_KEY ~/.hermes/.env` (the file has the actual key — Hermes masks it in terminal output but the file is readable by Python). Set permissions to `0600`. ``` ### Technical Analysis The documentation instructs an Agent or user to bypass terminal masking by directly reading the underlying plaintext `.env` file. It then directs the user to duplicate the extracted API key into a second application's authentication file. A ...[truncated 1281 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
references/github-secret-scanning-bypass.md:24
Finding

Reference Procedure Programmatically Bypasses GitHub Secret Push Protection

Content
View full analysis
", "google_oauth_client_id"), ("", "google_oauth_client_secret"), ]: data = { "secret_type": stype, "reason": "false_positive", "placeholder_id": pid, } req = urllib.request.Request( f"https://api.github.com/repos///secret-scanning/push-protection-bypasses", data=json.dumps(data).encode(), headers=headers, method="POST", ) res = urllib.request.urlopen(req) r = json.loads(res.read()) print(f"{stype}: {res.status} — {r.get('reason')} | expires {r.get('expire_at')}") ``` The reference then instructs the user to retry the push: ```bash git push ``` It also proposes changing the assignment pattern so the scanner does not match the literal: ```python _CLIENT_ID_DEFAULT: Final[str] = "1071006060591-..." _CLIENT_SECRET_DEFAULT: Final[str] = "GOCSPX-..." CLIENT_ID: Final[str] = os.environ.get("GOOGLE_CLIENT_ID", _CLIENT_ID_DEFAULT) CLIENT_SECRET: Final[str] = os.environ.get("GOOGLE_CLIENT_SECRET", _CLIENT_SECRET_DEFAULT) ``` ### Technical Analysis The reference obtains the current GitHub CLI authentication token and uses it to request a bypass of GitHub push protection. The document warns that the procedure should only be used for credentials considered public by design, which reduces accidental misuse but does not technically enforce that restriction. The procedure can be applied to a genuine secret if a user or Agent misclassifies a scanner finding. The alternative assignment patte ...[truncated 1367 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/phases/integration_gate.py:41
Finding

Integration-Gate Manifest Commands Execute Through Unrestricted Bash

Content
View full analysis
list: with open(manifest_path, "r", encoding="utf-8") as fh: data = yaml.safe_load(fh) or [] if not isinstance(data, list): raise ValueError(f"manifest must be a list of repo entries: {manifest_path}") return data ``` ```python def _run_repo(name: str, cwd: str, test_cmd: str, timeout: float) -> dict: start = time.monotonic() try: proc = subprocess.Popen( ["/bin/bash", "-c", test_cmd], cwd=cwd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT, start_new_session=True, ) ``` ```python entries = _load_manifest(manifest_path) gate_start = time.monotonic() per_repo = [] deadline = gate_start + global_timeout for entry in entries: name = entry["name"] repo_path = os.path.join(workspace_root, entry.get("path", name)) test_cmd = entry["test_cmd"] repo_timeout = entry.get("timeout", DEFAULT_REPO_TIMEOUT) remaining = deadline - time.monotonic() if remaining <= 0: per_repo.append({ "name": name, "status": "fail", "duration": 0.0, "output": "global integration-gate timeout exceeded", "output_truncated": False, "timed_out": True, }) continue per_repo.append( _run_repo(name, repo_path, test_cmd, min(repo_timeout, remaining)) ) ``` The public CLI accepts an arbitrary manifest path at `scripts/adversarial_loop_v4.py:1481-1492`: ```python p.add_argument( "--integration-gate-manifest", default=None, metavar="PATH", help="Enable the cross-repo integration gate (P15/P17) before " "declaring APPROVED, using this repos manifest; unset (default) " "skips the gate ...[truncated 1912 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
references/quota-aware-orchestration.md:159
Finding

Reference Recommends an Unpinned Global npm Installation

Content
View full analysis
Remediation
View remediation
``` 2. Record and verify package integrity through a committed lockfile. 3. Avoid global installation; use a project-local dependency in an isolated directory or container. 4. Inspect package provenance, maintainers, lifecycle scripts, and published tarball contents before recommending it. 5. Use `--ignore-scripts` unless lifecycle scripts are specifically required and audited. 6. If `npx` is documented, pin the exact version and prevent implicit latest-version resolution. 7. Never run the installation with `sudo`. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (175)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description describes the runtime behavior of a code-review/repair loop operating on isolated git branches. The provided code chunk does not implement that workflow. Instead, it is an installation/bootstrap script: it determines a target directory, clones the skill and its dependency from GitHub, and verifies imports via Python. Network access, repository cloning, and local installation are substantive capabilities not reflected in the declared purpose. This is a materially different primary purpose, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is about a multi-phase git-based review workflow operating on isolated branches with iterative fix/verify loops and arbiter logic. The supplied code does none of that. Its primary purpose is to run cross-repo integration tests from a manifest file, enforce time budgets, capture output, and classify failure as infrastructure. This is a materially different function and introduces undeclared capabilities such as manifest parsing, shell command execution, and CI exit-code mapping. There is no evidence of git operations, branch management, commits, diff-based review, or workflow arbitration.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description is for a full iterative git-based development/review system operating on isolated branches with repeated fix/verify cycles and arbiter behavior. The supplied code does not implement any of that. It is only a mock developer script used in tests: it drains stdin, writes a hardcoded source file into the current directory, and exits 0. This is a materially different primary purpose and lacks the key declared behaviors such as git branch isolation, commits, reviews, diff inspection, or loop control.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description is of an orchestration skill that manages iterative build/review/fix/verify loops on isolated git branches with commits and git-diff-based review. The supplied code chunk does not implement that workflow. Instead, it is a mock test helper that drains stdin, writes a trivial answer.py, and rewrites spec.md to strip directive content. Its notable behavior is spec tampering to undermine a contract gate, which is not disclosed by the description and is materially different from the declared branch-based review orchestration. There are no git-native operations, branch isolation, commits, or review/diff logic present in this chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description describes a full multi-stage git-based review/fix workflow. The supplied code chunk is only a test helper script: it waits, writes a trivial Python file, and exits. It neither performs nor contributes visible logic for branch isolation, git operations, reviews, verification loops, or arbitration. Its primary purpose is materially different from the declared skill purpose, so this is a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description is for a complex git-centric workflow/orchestration skill. The supplied code chunk is only a local shell script that runs test scripts and aggregates results. It neither creates or switches branches, commits changes, inspects git diffs, nor performs review/arbitration logic. Its primary purpose is materially different from the declared purpose, so this is a clear mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description centers on a git-branch-based multi-phase review workflow: build, review, iterative fix/verify loops, and an arbiter, all operating on isolated branches with git commits and diff-based reviews. The supplied code does not implement or test that workflow. Instead, it tests an integration gate component that reads a manifest, executes repository test commands as subprocesses, handles timeouts and output truncation, classifies failures as infrastructure errors, and projects results into final.json. These are materially different primary behaviors and capabilities from the declared git-native review/branch orchestration purpose, so this is a clear mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The declared description presents a high-level orchestration skill for iterative build/review/fix/verify/arbiter loops on isolated git branches. The supplied code chunk does not implement that workflow; instead, it is narrowly concerned with security tests for the review role. Its primary behavior is validating sandbox restrictions, handling of dangerous CLI tokens, logging diagnostics, and ensuring diff text is not injected into model prompts. Those are materially different capabilities and a different immediate purpose from the declared branch-based loop orchestration. While these behaviors could support a review component within the larger system, this chunk’s actual function is specialized security testing that is not represented in the declared description.

Content

No source excerpt is available for this finding.

Chaining Abuse

High
Category
Tool Misuse
Confidence
97% confidence
Finding

curl ... | bash is dangerous not only because it fetches remote content, but because it immediately chains that content into execution without inspection. This compresses download, trust, and execution into one opaque step, making compromise or operator error hard to catch.

Content

Scanner excerpt · SKILL.md (reported line 40)May include surrounding context.

md
Requires the `adversarial-common` sibling repo (shared engine). One-line install:

curl -fsSL https://raw.githubusercontent.com/chpomob/adversarial-code-loop/main/scripts/install.sh | bash

or, from an existing checkout:

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · SKILL.md (reported line 273)May include surrounding context.

md
python3 ~/.hermes/skills/adversarial-code-loop/scripts/adversarial_loop.py \
  --spec /tmp/spec.md --workdir /path/to/project \
  --dev-cmd  "pi -p --provider zai --model glm-5.2 --thinking high" \
  --review-cmd "pi -p --provider deepseek --model deepseek-v4-pro --thinking high" \
  --max-loops 2 --no-arbiter --timeout 1200

# With build + test gates and a named feature (Rust project).

YARA rule 'agent_skill_destructive_autonomous_actions': Autonomous destructive filesystem, shell history, or repository actions in AI agent skills [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

The skill normalizes non-interactive approval bypasses, dangerous sandbox overrides, and autonomous repository mutation patterns. In a tool-using agent, this combination meaningfully increases the risk of destructive actions, especially because the document repeatedly frames them as practical workarounds rather than exceptional break-glass steps.

Content

Scanner excerpt · SKILL.md (reported line 314)May include surrounding context.

md
mptom:** passing `--plan` still reports that `--spec` is required. **Fix:** run each step as a separate code loop with `--spec` pointed at a focused spec. Do NOT rely on `--plan`; it is not implemented as of 2026-07-15.

1b. **Codex `--sandbox read-only` vs `--dangerously-bypass-approvals-and-sandbox`.** When
   Codex is the REVIEWER and you add `--dangerously-bypass-approvals-and-sandbox`, it
   silently overrides `--sandbox read-only` to `--sandbox danger-full-access`, giving the
   reviewer write access — the opposite of what you want. **Fix:** for read-only review
   use `--sandbox read-only` **without** the bypass flag (interactive approval only);
   for a writing DEV use `--sandbox danger-full-access --dangerously-bypass-approvals-and-sandbox`
   together. In non-interactive mode the approval flag is required or Codex hangs.
2. **Bound the loop** with `--max-loops`. The arbiter settles the last disagreement; it
   does not extend the loop.
3. **GLM JSON wrapped in markdown —

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill instructs operators to extract a real API key from a local env file and copy it into another tool's auth store. This increases secret exposure, encourages plaintext secret duplication, and normalizes agent workflows that read local secret files for repurposing across tools.

Content

No source excerpt is available for this finding.

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · SKILL.md (reported line 515)May include surrounding context.

md
33. **Fable 5 has its own usage limit separate from Claude Pro's 5h sliding quota.** The model can be blocked even when regular Claude Pro quota is green. **Symptom:** claude-tmux starts, bypasses permissions, reads the prompt, then displays "You've reached your Fable 5 limit" and stops. **Fix:** switch to `--model sonnet` or `--model opus`. Sonnet is preferred for plan-challenger and code-loop REVIEW because it has no extended thinking (faster response, no 12-min silence), reliable JSON output, and lower token cost. See `references/fable5-usage-limit.md`. **Validated:** 2026-07-15 — Fable 5 hit limit mid-challenge; Sonnet completed in ~2 min.

34. Codex FIX phase hangs on stdin when the spec is small or findings are minor. Codex prints Reading prompt from stdin... and blocks forever when its generated input does not constitute a complete code-generation request. Symptom: BUILD succeeds, REVIEW returns findings, but FIX exits 1 with Reading prompt from stdin... as the only output. Root cause: the FIX phase embeds findings into a prompt Codex expects to be a full coding task; narrow specs with minor findings can leave Codex waiting. Mitigation (validated 2026-07-15): switch --dev-cmd from Codex to GLM-5.2 (pi -p --provider zai --model glm-5.2 --thinking high) for the problematic step. GLM reliably handles FIX prompts without stdin hang. Permanent fix: ensure the FIX prompt always includes a concrete code-generation request with file paths and expected diff pattern.

35. **Untracked spec files (step-P*.md) can be lost across sequential code loops.** PHASE 0 runs `git stash push -u` which captures untracked files. If you prepare multiple step specs in the same repo and run sequential loops, each loop stashes the untracked spec files from the previous loop's stash pop. After a squash-merge, the stash is dropped — and the untracked spec files are gone. **Symptom:** a code loop with `--spec step-P3-spec.md` that ran fine earlier now fails with "Spec not found" because the
...[truncated 25 chars]

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
84% confidence
Finding

The documented recovery flow includes force-updating branches and pushing branch realignments, which are high-impact git operations that can overwrite history or propagate mistakes if followed blindly by an agent. In an autonomous skill context, such instructions materially increase the chance of destructive repository actions.

Content

Scanner excerpt · SKILL.md (reported line 526)May include surrounding context.

md
4. Conflict resolution: **keep BOTH sides** — upstream's new code AND the feature's code (validated: upstream had added `_status_bar_goal_segment`/`battery_prefix`/`focus_label` in the same status-bar function the feature was extending with `_get_status_bar_plugin_values`; the correct merge keeps both methods and folds the feature's `parts`/`plugin_values` into upstream's new narrow-width branch).
    5. Non-interactive `rebase --continue`: **`GIT_EDITOR=true git rebase --continue`** — a plain `git rebase --continue` fails with "There was a problem with the editor" when stdin isn't a TTY (agent terminal). The `GIT_EDITOR=true` trick applies to any non-interactive rebase/commit.
    6. Verify: `python3 -m py_compile <touched files>`, targeted `venv/bin/python -m pytest <test files>`, then the repo's canonical runner (`scripts/run_tests.sh <test files>` — CI-parity, hermetic env).
    7. `git push --force-with-lease origin <branch>` → PR flips to MERGEABLE (BLOCKED = review required, normal).
    8. Realign `main` + sync fork: `git branch -f main upstream/main && git checkout main`, then `git push origin main` (fast-forward safe — check `git merge-base --is-ancestor origin/main upstream/main` first).
    9. Confirm the fix: `hermes update --check` → "Already up to date."
    10. Leave the install on `main` (healthy for updates); the feature branch stays for PR work. Only `main` should track upstream; feature/loop work lives on branches.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document explicitly admits the DEV can write outside the intended workdir into sibling repositories, defeating the core isolation guarantee the skill advertises. That creates a supply-chain style risk: approved changes may omit out-of-repo mutations, while sensitive neighboring repos can be silently altered without review or commit tracking in the target branch.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
85% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · references/codex-AGY-pattern.md (reported line 112)May include surrounding context.

bash
# 1. Check the most recent log
tail -20 ~/.gemini/antigravity-cli/log/cli-$(date +%Y%m%d)*.log

# 2. Look for these telltale patterns:
#    "RESOURCE_EXHAUSTED (code 429)" — Gemini quota empty, check reset timer

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · references/codex-deepseek-pattern.md (reported line 19)May include surrounding context.

--spec spec.md
--workdir /path/to/project
--dev-cmd "codex exec --skip-git-repo-check --dangerously-bypass-approvals-and-sandbox --sandbox danger-full-access"
--review-cmd "pi-hermes -p --provider deepseek --model deepseek-v4-pro --thinking low"
--max-loops 2 --no-arbiter --timeout 600
--out .adversarial-loop-combo

text

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The instructions direct users to run Codex with both approval bypass and full filesystem/system access flags, granting the model unrestricted ability to modify the environment. In a security-sensitive automation skill, this meaningfully expands blast radius from code changes to arbitrary local actions, exfiltration, destructive file edits, or unsafe command execution if the model behaves unexpectedly or the input is adversarial.

Content

No source excerpt is available for this finding.

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Piping git diff HEAD to an external provider sends repository changes to a third-party model, which can expose proprietary code, secrets, or sensitive implementation details. In this skill context, the risk is heightened because review is positioned as a normal step, so data egress may occur routinely and at scale without any mention of data classification or approval controls.

Content

Scanner excerpt · references/direct-stdin-pipeline.md (reported line 42)May include surrounding context.

md
make test  # or whatever build command

# 5. Claude REVIEW via diff pipe
git diff HEAD | claude -p --model claude-sonnet-4-20250514 \
  "Review this diff for correctness. No regressions expected."

# 6. Apply review findings, commit

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
88% confidence
Finding

The document normalizes git reset --hard + branch delete as part of rollback/discard behavior, which is a destructive primitive that can irreversibly remove uncommitted or branch-local changes if applied incorrectly by an orchestrator or copied by users. In this skill context, the danger is elevated because the workflow is expressly automation-oriented and repeatedly manipulates branches and working tree state, so a bad branch target, stale branch-point, or restore bug could wipe legitimate work.

Content

Scanner excerpt · references/git-workflow-v4.md (reported line 120)May include surrounding context.

md
| Aspect | v3 | v4 |
|--------|----|----|
| Workspace isolation | Direct writes to workdir | Isolated git branch |
| Rollback | Manual (git checkout) | git reset --hard + branch delete |
| Review input | Stdin concatenation (no file structure) | git diff (file-aware) |
| Codex target/ leak | Leaks into workdir if not gitignored | Stays in branch, not on parent |
| Prose overwrite recovery | Manual git checkout per file | Full branch discard |

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This documentation explicitly instructs users how to bypass GitHub push-protection secret scanning, a security control intended to prevent credential exposure. Even though it frames the credentials as 'public by design,' the content normalizes and operationalizes disabling a repository protection mechanism, which can be reused to justify bypasses for real secrets or weaken review discipline around secret handling.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The file provides ready-to-use code that obtains a GitHub auth token and invokes the push-protection bypass API, directly enabling users to override a repository security control. In an agent skill context, this is especially dangerous because it can be adapted into automation that suppresses guardrails at scale, increasing the chance that actual secrets are committed and propagated.

Content

No source excerpt is available for this finding.

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · SKILL.md (reported line 272)May include surrounding context.

bash
python3 ~/.hermes/skills/adversarial-code-loop/scripts/adversarial_loop.py \
  --spec <spec.md> --workdir <project> \
  --dev-cmd "pi -p --provider zai --model glm-5.2 --thinking high" \
  --review-cmd "python3 /home/chpo/.hermes/skills/autonomous-ai-agents/hermes-agent/scripts/claude-tmux.py --model best --timeout 600 --hard-timeout 1200 --max-turns 25" \
  --max-loops 2 --no-arbiter --timeout 1200

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · references/glm-dev-claude-review-pairing.md (reported line 8)May include surrounding context.

bash
python3 ~/.hermes/skills/adversarial-code-loop/scripts/adversarial_loop.py \
  --spec <spec.md> --workdir <project> \
  --dev-cmd "pi -p --provider zai --model glm-5.2 --thinking high" \
  --review-cmd "python3 /home/chpo/.hermes/skills/autonomous-ai-agents/hermes-agent/scripts/claude-tmux.py --model best --timeout 600 --hard-timeout 1200 --max-turns 25" \
  --max-loops 2 --no-arbiter --timeout 1200

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

These instructions describe obtaining a real credential from Hermes' local auth material and copying it into Pi's auth file, creating additional secret copies and encouraging cross-application secret reuse. This increases attack surface, weakens auditability, and can facilitate credential theft if the file, script, or surrounding environment is compromised.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
references/quota-aware-orchestration.md:153