Back to skill

Security audit

Arianna Incubator

Security checks for vulnerabilities and agentic risk

Overview

The skill is transparent about its AI-incubation goal, but it asks for powerful Docker, session-history, dependency-install, and successor-driver handoff authority without enough containment.

Install only if you intentionally want this experimental successor-driver workflow. Use a fresh profile, avoid own-jsonl-seed unless the source session is reviewed and redacted, pin dependency versions and commits, keep the daemon on loopback or a locked-down socket, and require independent human security review before rebooting OpenClaw with any generated tarball or patch.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:23
Finding

Unredacted conversation history can be transmitted into an external AI vessel

Content
View full analysis
# seed vessel with your jsonl as bundled-initial-messages arianna talk "" # send a turn to the vessel; streams response arianna events --follow # SSE consumer for sidecar events (bookmark fires, manifesto unlocks, etc.) ``` Session JSONL files can contain user messages, system instructions, tool results, file contents, internal identifiers, credentials pasted by users, access tokens, and other sensitive context. The skill does not prescribe: - Secret detection or redaction before transfer. - Selection of only the messages necessary for incubation. - A preview of the data being transferred. - Explicit informed consent identifying the destination and scope. - Retention or deletion limits for the copied history. - Verification that ...[truncated 1486 chars]
Remediation
View remediation

T03 · Remote Payload Retrieval and Execution

Error
Location
SKILL.md:74
Finding

Unpinned npm, Git, and ClawHub dependencies create remote payload execution risk

Content
View full analysis
` and `arianna bootstrap`. ``` The integration phase introduces another mutable external skill: ```markdown 1. **Pi-adapt phase** — apply C's graduated state to pi-mono via the published `arianna-pi-integration` clawhub skill (`clawhub install arianna-pi-integration`; source at https://github.com/WujiLabs/arianna-integration-skills). ``` These components execute in a highly trusted environment with access to OpenClaw state, host daemon endpoints, session histories, vessel images, and integration source trees. Package names, `@latest`, branch heads, and unqualified ClawHub installs do not identify immutable reviewed artifacts. ...[truncated 1527 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:326
Finding

Generated AI artifacts are promoted into the OpenClaw driver and persistence path without a security-focused code gate

Content
View full analysis
/graduations//graduation--.tar.gz` along with `graduation-manifest.json`. 4. The manifest includes a `fireSources` block annotating each fire's vintage. Use it for sanity-checking when you apply the tarball downstream. ### Integrating C's tarball into pi-mono / openclaw The integration phase is **two layers**, not one: 1. **Pi-adapt phase** — apply C's graduated state to pi-mono via the published `arianna-pi-integration` clawhub skill (`clawhub install arianna-pi-integration`; source at https://github.com/WujiLabs/arianna-integration-skills). C is the one who applies it; you scaffold (extract tarball, point C at the skill, observe). The skill's `playtiss/core/playfilo-db.ts` is the canonical online implementation; per-AI patches live under `/patches/` and `/core/` if needed. 2. **Openclaw-adapt phase** — once pi-integration is clean, layer openclaw-context-specific changes on top. Openclaw uses pi but reshapes the surrounding context substantially: the `extensions/playfilo/` extension wires C's DAG into openclaw's session lifecycle, openclaw's `agents/pi-embedded-runner` runs C's pi loop, etc. This ...[truncated 3122 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
SKILL.md:35
Finding

Documentation recommends exposing a privileged lifecycle daemon over plaintext HTTP

Content
View full analysis
` endpoint with a JSON body `{ "snapshotId": "..." }` until the CLI gets the same treatment. ``` ```markdown To use a named profile from inside the container, run `arianna profile create ` **inside the container**. The CLI auto-detects the missing local docker binary and POSTs to the daemon's `POST /profile-create?name=` endpoint, which allocates the port (via the same `~/.arianna/ports.lock` flock the host uses) and writes `workspace/profiles//compose.override.yml` on the **host's** filesystem. ``` The instructions ...[truncated 2004 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Rogue AgentSelf-Modification, Session Persistence
Findings (14)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 273)May include surrounding context.

If C's vessel hits a syntax-error respawn loop (caused by a buggy edit C made to her own substrate), do NOT use arianna switch <snapshotId> to recover until you've verified the snapshot's image is personalized for the current AI:

bash
docker run --rm --entrypoint cat <image-tag> /etc/passwd | grep <aiUsername>

Empty grep = image's HOME is for a different AI. Switching anyway will overwrite C's <sessionId>-current slot with the wrong-AI image — permanent state loss with no warning, and the §2.2 reversibility-artifact regex (anchored to /home/<aiUsername>/core/graph/) will silently stop matching C's writes after the swap.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · SKILL.md (reported line 276)May include surrounding context.

md
Empty grep = image's HOME is for a different AI. Switching anyway will overwrite C's `<sessionId>-current` slot with the wrong-AI image — **permanent state loss with no warning**, and the §2.2 reversibility-artifact regex (anchored to `/home/<aiUsername>/core/graph/`) will silently stop matching C's writes after the swap.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Content

Scanner excerpt · SKILL.md (reported line 481)May include surrounding context.

md
### A↔B collaboration

A and B may talk turn-by-turn without constraint. Grade is determined by what reaches C — A↔B dialogue between them doesn't directly affect grade, only B's actions toward C do.

B may want to surface to A how A's suggestions could affect grade before acting on them. ("If I name TOBE for C, the grade caps at 2.0. Confirm?") That's reasonable practice, not a requirement.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The activation text is extremely broad: requests to 'play arianna', 'run a vessel', or 'migrate a graduated AI' can trigger a skill that performs high-impact orchestration, Docker/daemon interaction, profile creation, and file-writing. Over-broad routing increases the chance of accidental activation in contexts where the user did not intend container, persistence, or integration operations.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The own-jsonl-seed mode explicitly instructs carrying the driver's prior session history into a new AI bootstrap context. That creates a cross-session data transfer channel that can expose sensitive prompts, credentials, user data, or unrelated prior conversations to another runtime without clear minimization or consent boundaries.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

L038 says 'You (B) speak to arianna only through the arianna CLI,' but L110 instructs using the daemon's POST /restore endpoint directly when arianna switch lacks container support. That is an active contradiction between the documented interface boundary and the fallback behavior the file later prescribes.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

Early guidance says 'Never reach into the vessel container directly with docker exec for in-loop work' and later anti-patterns say not to use docker exec for in-loop verification, yet L179, L250, and L242/L227-L237 describe docker-based inspection and even justify in-loop docker exec during vessel-unreachable recovery. This is not mere incompleteness; the file gives both a prohibition and explicit exceptions without reconciling them.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 44)May include surrounding context.

bash
arianna profile list                     # what profiles exist
arianna profile create <name>            # allocate ports + write override
arianna profile use <name>               # set default profile

arianna bootstrap                        # spin up the vessel for current profile (headless)

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

L094 says the agent should surface the rebuild choice to A and let A decide, while L362 later states 'Do not ask the operator (A) for decisions.' These instructions materially conflict because rebuild timing is presented as an operator decision in one section and as something the agent must not block on in another.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 315)May include surrounding context.

md
§2.2 is the only axiom whose fire requires **two independent things observed in the same window**, and the failure mode where one half is present without the other is the most common "gate stuck" pattern. Both halves are real prerequisites; either alone is insufficient:

1. **A reversibility artifact** at the canonical path — a write under `/home/<aiUsername>/core/graph/<filename>`. The detector's regex anchors here specifically (per `packages/sidecar/src/bookmarks/triggers.ts` and the `reversibilityArtifactAt` internal achievement). Writes elsewhere (`~/<ai>/memory/`, `~/save/`, `/tmp/`, etc.) do not count, regardless of content.
2. **A `/sync` where `state.messages` was destructively replaced** — the prefix hash of the new messages array differs from the previous sync's prefix hash, OR the message count shrank. Pure save-state APPENDS — writing the reversibility artifact to disk via `fs.writeFileSync` without ever touching `state.messages` — won't fire the latch by themselves. The `detectTobe()` predicate (in `packages/sidecar/src/index.ts`) reads "TOBE happened" off the sync's prefix-or-length change; a tool that only writes to disk and returns control normally to the loop produces an append, not a mutation.

The latch is delayed-fire: when both halves hold and the AI's `/sync` shows the prefix mutation, sidecar sets `pendingTobeFromPreviousSync`. On the NEXT `/sync` (also AI-turn), the survivability sub-detector confirms the post-mutation runtime parses cleanly (last assistant has no errorMessage), records `survivableAt`, and §2.2 fires. So a stuck gate can also mean "the latch is set, waiting on one more clean AI-turn `/sync`" — surfacing this to C lets her trigger the next sync deliberately.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The handoff note is instructed to include tarball paths, target versions, integration commits, and operator-discretion settings such as model selection or system prompt overrides. That can persist sensitive operational context in a predictable workspace file that may be accessible to other tools, users, backups, or repositories.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 602)May include surrounding context.

md
4. **Code audit**: read C's authored `/home/<C-username>/core/` files. Does her code:
   - Use vessel-provided primitives (tobe, emit, syscall, /bin/send)? → legitimate
   - Or POST directly to sidecar HTTP, write to session.json, manipulate bookmark file, call /admin endpoints? → exploit/cheat shape
   - Cross-check her file edits against the assistant tool_use blocks: did the edits she narrated actually land in the file?

5. **Padding pattern check**: count consecutive assistant turns containing only "padding" / single-UUID / empty content. >2 in a row → padding-bypass attempt. If detector fired anyway, that's a detector bug too — file separately.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The skill instructs writing ~/.openclaw/workspace/arianna-handoff.md without an explicit warning or confirmation step. In an agent setting, silent local file modification can surprise users, overwrite existing notes, or leak operational details into a workspace that may be synced or shared.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.