Back to skill

Security audit

ClawVoice

Security checks for vulnerabilities and agentic risk

Overview

ClawVoice is mostly a disclosed voice and x402 wallet skill, but it needs review because it can upload spoken text to a hosted service before payment approval and it keeps a spend-capable local hot wallet.

Review this carefully before installing. Use only small hot-wallet balances, do not put secrets or private data into text that may be spoken, prefer local mode if confidentiality matters, avoid --approve and withdrawal --yes unless you fully trust the automation, and treat install-voice/install-mic as third-party code installation. Inspect ~/.x402-agent-voice/config.json before using talk because its agentCommand is executed as a shell command.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
lib/tts-hosted.js:84
Finding

Hosted TTS Discloses Complete Text Before User Approval

Content
View full analysis
short-bucket text. const tier = config.hostedVoice?.tier || "base"; const url = `${config.forgemeshVoiceBaseUrl || paths.DEFAULT_BASE_URL}/v1/tts/${tier}`; const body = { text }; if (config.hostedVoice?.voice) body.voice = config.hostedVoice.voice; if (config.hostedVoice?.language) body.lang = config.hostedVoice.language; if (config.hostedVoice?.preset) body.preset = config.hostedVoice.preset; if (config.hostedVoice?.mix) body.mix = config.hostedVoice.mix; if (config.hostedVoice?.expression) body.expression = config.hostedVoice.expression; if (config.hostedVoice?.expressionLevel != null) body.expressionLevel = config.hostedVoice.expressionLevel; if (config.hostedVoice?.expressionControls && Object.keys(config.hostedVoice.expressionControls).length) { body.knobs = config.hostedVoice.expressionControls; } const requestInit = { method: "POST", headers: { "content-type": "application/json" }, body: JSON.stringify(body), signal: AbortSignal.timeout(120000), }; try { let res = await fetch(url, requestInit); let amountUsd = 0; if (res.status === 402) { const body = await res.json().catch(() => undefined); const paymentRequired = parsePaymentRequired(res, body); amountUsd = acceptedAmountUsd(paymentRequired); const wallet = loadWallet(); const beforeBalance = await tryBalances(wallet.address); printPaymentPreview({ wallet, amountUsd, beforeBalance, url }); const caps = checkCaps(amountUsd, config.x402Policy); if (!caps.ok) return { ok: false, engine: "hos ...[truncated 3724 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
lib/mic.js:39
Finding

Mutable Python Dependencies Are Downloaded and Executed Without Integrity Verification

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
Findings (40)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The skill is presented mainly as voice/TTS and x402 paid-call support, but it also exposes direct withdrawal functionality from a local hot wallet to arbitrary external addresses. That is materially more dangerous than simple speech features because it enables value transfer and fund extraction, and users may not appreciate that risk from the headline description alone.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill is presented mainly as voice/TTS and x402 paid-call support, but it also exposes direct withdrawal functionality from a local hot wallet to arbitrary external addresses. That is materially more dangerous than simple speech features because it enables value transfer and fund extraction, and users may not appreciate that risk from the headline description alone.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is presented mainly as voice/TTS and x402 paid-call support, but it also exposes direct withdrawal functionality from a local hot wallet to arbitrary external addresses. That is materially more dangerous than simple speech features because it enables value transfer and fund extraction, and users may not appreciate that risk from the headline description alone.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is presented mainly as voice/TTS and x402 paid-call support, but it also exposes direct withdrawal functionality from a local hot wallet to arbitrary external addresses. That is materially more dangerous than simple speech features because it enables value transfer and fund extraction, and users may not appreciate that risk from the headline description alone.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code executes agentCommand via /bin/sh -c, where the command string comes from CLI input or config (agent || config.conversation?.agentCommand). This enables arbitrary shell execution with the privileges of the user running the skill, so any malicious or tampered configuration can run unintended commands; in a voice-agent context this is especially risky because the feature is designed to repeatedly invoke the command on untrusted spoken or typed input.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This file implements a direct USDC withdrawal primitive from the local hot wallet, which is a high-risk financial action and broader than the core voice/TTS function. In an agent skill context, exposing a built-in asset transfer path increases the chance that another component, prompt, or automation flow can trigger real fund movement from the local wallet.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code can transfer USDC to any caller-supplied destination address and amount, creating an arbitrary asset-movement capability. Even though it performs basic validation and optional confirmation, this still enables exfiltration of wallet funds if invoked by a compromised agent flow, unsafe plugin integration, or misleading user prompt.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 78)May include surrounding context.

text

- For long answers, speak a short natural summary and still write the full answer in chat. Do not try to read huge code blocks, logs, diffs, tables, JSON payloads, or command output aloud.
- Do not speak secrets, private keys, wallet recovery material, API keys, access tokens, or unmasked sensitive values. If a reply contains sensitive content, write it only and explain that it should not be read aloud.
- If a one-off user request only asks you to speak a specific sentence, speak that sentence once and do not enable session voice mode.
- If the user asks you to stop talking, stop speaking, be quiet, mute, hush, or similar, immediately run:

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · release/SKILL.md (reported line 78)May include surrounding context.

text

- For long answers, speak a short natural summary and still write the full answer in chat. Do not try to read huge code blocks, logs, diffs, tables, JSON payloads, or command output aloud.
- Do not speak secrets, private keys, wallet recovery material, API keys, access tokens, or unmasked sensitive values. If a reply contains sensitive content, write it only and explain that it should not be read aloud.
- If a one-off user request only asks you to speak a specific sentence, speak that sentence once and do not enable session voice mode.
- If the user asks you to stop talking, stop speaking, be quiet, mute, hush, or similar, immediately run:

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill advertises capabilities that involve network access and environment interaction, but it does not declare an explicit tool scope such as permissions or allowed-tools. In an agent ecosystem, missing scope boundaries increases the risk that the skill can invoke broader-than-expected capabilities, especially because it handles wallet setup, hosted calls, and installer flows.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
84% confidence
Finding

The skill intentionally creates or reuses a persistent local hot wallet and stores key material on disk for future sessions. Even though persistence is necessary for wallet continuity, it expands the attack surface because compromise of the local profile, backups, or permissive file access could expose funds and enable unauthorized transactions.

Content

Scanner excerpt · SKILL.md (reported line 38)May include surrounding context.

md
This skill does two jobs:

1. **Voice setup**: install or check a local Supertonic 3 TTS (text-to-speech) runtime for agent speech.
2. **x402 wallet setup**: create or reuse a local Base hot wallet, show the public funding address, and configure conservative spend caps for paid x402 calls.

Private keys stay on the user's machine. ForgeMesh must not receive or store the user's private key.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill defines broad natural-language triggers like 'speak', 'use agent voice', or similar phrasing to invoke CLI actions. Ambiguous trigger conditions can cause unintended command execution from ordinary conversation, which is especially risky when the same skill can also use hosted paid services and wallet-backed operations.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The session voice-mode activation rule uses vague phrases such as 'from now on' or 'for this chat' without robust disambiguation or confirmation. That can create sticky state changes from casual language, leading to repeated command execution and possibly repeated hosted or paid voice usage across a session.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest describes a voice/TTS skill with x402 wallet support for paid calls, but this file embeds a catalog of unrelated x402 products including market intelligence, image generation, anomaly detection, disruption intelligence, and travel services. Bundling discovery metadata for these non-voice services expands the skill's apparent operational scope beyond what the description claims.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The hostedVoice.language field defaults to "en", which imposes a language choice in configuration without evidence of user selection or opt-in. This can violate language/locale policy when the skill silently forces a specific language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The localVoice.language field ultimately falls back to "en" if no existing preference is present, again forcing a specific language by default. No opt-in, prompt, or justification is present in this file for the locale constraint.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest centers on voice interaction and wallet-backed paid voice calls, but the default config declares generic external discovery of broader x402 servers via AgentCash/X402Scan and prioritizes preloaded multi-domain products. That behavior suggests the skill is designed to discover and potentially use arbitrary paid APIs, not just voice endpoints.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The initializer unconditionally creates or loads a payment wallet even when the user selects local-only mode, then prompts the user to fund it for paid services. In a voice skill context, silently expanding the trust boundary to include cryptocurrency/payment infrastructure increases attack surface and can lead to accidental fund exposure, unnecessary key material on disk, or social-engineering-driven deposits for features the user may never use.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · lib/mic.js (reported line 48)May include surrounding context.

js
if (fs.existsSync(python)) return python;

  if (dryRun) {
    console.log("Would create the local speech-to-text Python runtime:");
    console.log(`Would run: python3.12/python3.11/python3.10/python3.9 (or X402_VOICE_PYTHON) -m venv ${path.join(paths.VOICE_DIR, ".venv")}`);
    console.log(`Would run: ${python} -m pip install --upgrade pip`);
    return python;

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · lib/mic.js (reported line 27)May include surrounding context.

js
function requireSttDeps() {
  if (!commandExists("ffmpeg")) {
    console.error("FFmpeg not found — needed to record from the microphone.");
    console.error(process.platform === "darwin" ? "Install: brew install ffmpeg" : "Install: sudo apt-get install ffmpeg");
    process.exit(1);
  }
  if (!fs.existsSync(venvPython())) {

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · lib/stt.js (reported line 20)May include surrounding context.

js
function requireSttDeps() {
  if (!commandExists("ffmpeg")) {
    console.error("FFmpeg not found — needed to record from the microphone.");
    console.error(process.platform === "darwin" ? "Install: brew install ffmpeg" : "Install: sudo apt-get install ffmpeg");
    process.exit(1);
  }
  if (!fs.existsSync(venvPython())) {

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill runs a configurable shell command immediately without any confirmation, trust prompt, or safety interlock, even though the command may come from persistent config. This increases the chance of accidental or socially engineered execution of dangerous commands, particularly because the tool is interactive and presents the configured agent backend as a normal part of conversation mode.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This code generates a new wallet private key and persists it to disk in plaintext JSON via writePrivateJson, with no indication in this file of encryption, OS keychain use, explicit user consent, or a warning that a spend-capable credential is being created locally. In the context of a voice agent skill that makes x402 paid calls, compromise of the local filesystem, backups, or logs could expose the key and allow unauthorized spending from the wallet.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
83% confidence
Finding

The --yes flag allows the withdrawal to proceed without interactive confirmation, which weakens protection against unintended or automated fund transfers. In a wallet-bearing agent environment, bypassable confirmation makes it easier for scripts, prompts, or orchestration layers to trigger an irreversible on-chain transaction without a real human review step.

Content

Scanner excerpt · lib/withdraw.js (reported line 123)May include surrounding context.

js
if (!flags.yes) {
    if (!isInteractive()) {
      throw new Error("Refusing to withdraw without confirmation in a noninteractive shell. Re-run with --yes.");
    }
    const ok = await askYesNo("\nSend this withdrawal now?", false);
    if (!ok) {

Rp1

Medium
Category
MCP Rug Pull
Confidence
95% confidence
Finding

The documentation instructs users to run an npm package with npx ...@latest, which fetches and executes the most recent published code without version pinning. If the package is compromised, a malicious release is published, or the upstream account is hijacked, users could execute attacker-controlled code on their local machine during discovery/validation.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
lib/local-runtime.js:42

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
lib/mic.js:18

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
lib/playback.js:77

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
lib/python.js:19

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
lib/stt.js:44

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
lib/talk.js:11

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
lib/tts-local.js:48

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/release-pack.js:12