Back to skill

Security audit

codex

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent Codex delegation guide, but it recommends high-impact autonomous coding modes without enough approval and isolation safeguards.

Install only if you are comfortable letting Hermes delegate code changes to Codex. Prefer sandboxed workspace-write in trusted repositories, avoid danger-full-access except in a disposable or otherwise isolated environment, and require human review before commits, pushes, PR creation, or any run using no sandbox or no approvals.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
SKILL.md:72
Finding

Unsandboxed Autonomous Code Execution Through Danger-Full-Access Fallback

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 72–87
Vulnerability Type: Unrestricted host access beyond the workspace boundary
Risk Level: High

Vulnerable instructions:

text
In that context, prefer:

codex exec --sandbox danger-full-access "<task>"

Use process boundaries as the safety layer instead: explicit `workdir`, clean git
status before launch, narrow task prompts, `git diff` review, targeted tests, and
human/agent confirmation before committing broad changes.

Technical Analysis

When normal workspace sandboxing fails in a Hermes gateway or service context, the Skill recommends rerunning Codex with --sandbox danger-full-access. This mode removes the Codex filesystem sandbox and allows the autonomous coding agent to execute commands with the permissions of the Hermes host process.

The suggested safeguards do not enforce an equivalent security boundary. An explicit workdir only selects the initial working directory; it does not prevent access to paths outside that directory. Reviewing git diff and requesting confirmation before committing occur after code execution and therefore cannot prevent prior host-file access, command execution, credential reads, or network activity.

The instructions do not impose a mandatory, informed user-confirmation gate immediately before entering unrestricted mode. This creates a reachable authorization expansion when sandbox setup fails.

Attack Path

  1. Hermes delegates a coding or review task to Codex in a repository containing third-party-controlled source, issue, pull-request, or instruction content.
  2. Codex workspace sandboxing fails because of the documented gateway or service restrictions.
  3. Following the Skill instructions, the Agent reruns the task using codex exec --sandbox danger-full-access.
  4. Codex processes the delegated task and repository content while running with the full permissions of the Hermes process a ...[truncated 1080 chars]
Remediation
View remediation

Remediation Suggestions

  • Do not automatically fall back to --sandbox danger-full-access when workspace sandboxing fails.
  • Fail closed and explain the sandbox error, or require explicit, informed user approval immediately before every unrestricted execution.
  • Prefer repairing the sandbox configuration or running Codex in a disposable container or virtual machine.
  • Mount only the intended repository into the isolated environment and make unrelated host paths unavailable.
  • Remove unnecessary credentials and environment variables from the Codex process.
  • Restrict outbound network access unless it is explicitly required for the approved task.
  • Run Codex under a dedicated, low-privilege account with no access to unrelated user files.
  • Preserve clean-status checks, narrow prompts, targeted tests, and diff review as defense-in-depth measures, but do not treat them as substitutes for runtime isolation.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
96% confidence
Finding

The skill recommends codex exec --sandbox danger-full-access as the preferred workaround when sandboxing fails in a gateway/service context. That guidance effectively shifts execution to a fully privileged external agent in exactly the environment where isolation is already unreliable, magnifying the risk of host compromise, sensitive file access, and unintended side effects.

Content

Scanner excerpt · SKILL.md (reported line 92)May include surrounding context.

In that context, prefer:

text
codex exec --sandbox danger-full-access "<task>"

Use process boundaries as the safety layer instead: explicit workdir, clean git

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Describing workspace-write as auto-approving file changes means the agent can modify repository contents without interactive confirmation. While less dangerous than full host access, this still enables autonomous code changes that may introduce malicious or harmful modifications if the prompt or environment is adversarial.

Content

Scanner excerpt · SKILL.md (reported line 75)May include surrounding context.

md
| Flag | Effect |
|------|--------|
| `exec "prompt"` | One-shot execution, exits when done |
| `--sandbox workspace-write` (`-s`) | Sandboxed but auto-approves file changes in the workspace (the recommended auto-build mode) |
| `--dangerously-bypass-approvals-and-sandbox` | No sandbox, no approvals (fastest, most dangerous; `--yolo` still works as a hidden alias) |
| `--sandbox danger-full-access` | No Codex sandbox; useful when the host service context breaks bubblewrap |

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
92% confidence
Finding

The documentation advertises a mode with 'No sandbox, no approvals,' which enables an autonomous coding agent to act without containment or confirmation. In this skill's context, that substantially raises the risk of destructive local changes, unsafe command execution, and unintended access to sensitive files or credentials.

Content

Scanner excerpt · SKILL.md (reported line 76)May include surrounding context.

md
|------|--------|
| `exec "prompt"` | One-shot execution, exits when done |
| `--sandbox workspace-write` (`-s`) | Sandboxed but auto-approves file changes in the workspace (the recommended auto-build mode) |
| `--dangerously-bypass-approvals-and-sandbox` | No sandbox, no approvals (fastest, most dangerous; `--yolo` still works as a hidden alias) |
| `--sandbox danger-full-access` | No Codex sandbox; useful when the host service context breaks bubblewrap |

> **Deprecated:** `--full-auto` still works but the live CLI warns to use `--sandbox workspace-write` instead.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly recommends codex exec --sandbox danger-full-access in a service/gateway context, and the nearby example does not include an inline warning at the point of use that this disables sandbox protections. Because this skill is designed to delegate autonomous coding tasks, presenting a no-sandbox command as the preferred fallback materially increases the chance of unsafe execution, file modification, or command execution on the host.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 110)May include surrounding context.

Parallel Issue Fixing with Worktrees

text
# Create worktrees
terminal(command="git worktree add -b fix/issue-78 ~/.hermes/cache/scratch/issue-78 main", workdir="~/project")
terminal(command="git worktree add -b fix/issue-99 ~/.hermes/cache/scratch/issue-99 main", workdir="~/project")

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The parallel issue-fixing workflow instructs Codex to 'Commit when done' and then later shows git push and PR creation steps, but it does not place a clear warning or approval gate before repository-modifying actions. In an autonomous agent context, this can lead to unintended commits, pushes, or PRs against real repositories without human review.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

The rules section recommends --sandbox workspace-write for building because it auto-approves changes, normalizing unattended modification of repository files. In an agent skill, this reduces friction for broad code changes and can bypass a human's opportunity to catch unsafe or unintended edits early.

Content

Scanner excerpt · SKILL.md (reported line 148)May include surrounding context.

md
1. **Always use `pty=true`** — Codex is an interactive terminal app and hangs without a PTY
2. **Git repo required** — Codex won't run outside a git directory. Use `mktemp -d && git init` for scratch
3. **Use `exec` for one-shots** — `codex exec "prompt"` runs and exits cleanly
4. **`--sandbox workspace-write` for building** — auto-approves changes within the sandbox (`--full-auto` is deprecated for this)
5. **Background for long tasks** — use `background=true` and monitor with `process` tool
6. **Don't interfere** — monitor with `poll`/`log`, be patient with long-running tasks
7. **Parallel is fine** — run multiple Codex processes at once for batch work

Static analysis

No suspicious patterns detected.