Back to skill

Security audit

XunFei Voice Reply

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent voice-reply integration, but it needs review because it persists reply mode, sends reply text to third-party services, and uses unsafe audio-conversion commands.

Install only if you are comfortable with replies being sent to Xunfei for speech synthesis and then to Feishu as audio. Before use, narrow the trigger commands, require confirmation before writing USER.md or enabling voice mode, pin dependencies, and replace shell-based ffmpeg calls with argument-array execution plus private per-run temp files.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
lib/tts-core.js:78
Finding

Shell Command Injection in FFmpeg Invocation

Content
View full analysis

Vulnerability Details

File Location: lib/tts-core.js:78-87; scripts/voice-reply.js:34-36
Vulnerability Type: OS command injection through unvalidated configuration values and file paths
Risk Level: High

Vulnerable Code

lib/tts-core.js:78-87:

javascript
if (outputPath) {
  const pcmPath = outputPath.replace(/\.(mp3|opus)$/, '.pcm');
  fs.writeFileSync(pcmPath, pcmAudio);

  const { execSync } = require('child_process');

  if (outputPath.endsWith('.mp3')) {
    execSync(`ffmpeg -y -f s16le -ar 16000 -ac 1 -i ${pcmPath} -ar ${this.audioConfig.outputSampleRate} -ac 1 ${outputPath} 2>/dev/null`);
  } else if (outputPath.endsWith('.opus')) {
    execSync(`ffmpeg -y -f s16le -ar 16000 -ac 1 -i ${pcmPath} -c:a libopus -b:a ${this.audioConfig.outputBitrate} -ar ${this.audioConfig.outputSampleRate} -ac 1 ${outputPath} 2>/dev/null`);
  }

scripts/voice-reply.js:34-36:

javascript
execSync(`ffmpeg -y -i ${mp3Path} -c:a libopus -b:a ${config.xunfei.audio.outputBitrate} -ar ${config.xunfei.audio.outputSampleRate} -ac 1 ${opusPath} 2>/dev/null`, {
  stdio: 'pipe'
});

Technical Analysis

The code constructs shell command strings using template interpolation and passes them to execSync. The interpolated outputBitrate and outputSampleRate values originate from editable config.json content and are loaded without schema or type validation. The reusable synthesize method also accepts an arbitrary outputPath, which is inserted into a shell command without quoting.

Because execSync executes the supplied string through a shell, shell metacharacters in any attacker-controlled value can terminate or alter the intended FFmpeg command. Quoting values alone would remain error-prone; the appropriate control is to avoid invoking a shell and pass each argument separately.

The default configuration is benign, so exploitation requires the attacker to modify configuration, influence an invocation of the exported library, or compro ...[truncated 1224 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace every string-based execSync call with execFileSync or spawnSync using an argument array and no shell:
javascript
const { execFileSync } = require('child_process');

execFileSync('ffmpeg', [
  '-y',
  '-f', 's16le',
  '-ar', '16000',
  '-ac', '1',
  '-i', pcmPath,
  '-c:a', 'libopus',
  '-b:a', validatedBitrate,
  '-ar', String(validatedSampleRate),
  '-ac', '1',
  outputPath
], {
  stdio: 'pipe'
});
  1. Require outputSampleRate to be an integer selected from an explicit allowlist, such as 16000, 24000, or 48000.
  2. Validate outputBitrate against a strict allowlist or a pattern such as ^[1-9][0-9]*k$, followed by an acceptable numeric range check.
  3. Resolve and normalize output paths, then verify that they remain inside a dedicated private output directory.
  4. Reject unexpected configuration keys and types using a JSON schema or equivalent runtime validator.
  5. Add tests containing spaces and shell metacharacters to confirm that values cannot affect command structure.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/voice-reply.js:19
Finding

Predictable Files in a Shared Temporary Directory

Content
View full analysis

Vulnerability Details

File Location: lib/tts-config.js:55-56; scripts/voice-reply.js:19-25; lib/tts-core.js:78-89
Vulnerability Type: Unsafe temporary-file creation, symbolic-link exposure, and cross-run file collision
Risk Level: Medium

Vulnerable Code

lib/tts-config.js:55-56:

javascript
// 输出目录 - 使用 OpenClaw 标准临时目录
outputDir: '/tmp/openclaw'

scripts/voice-reply.js:19-25:

javascript
const tmpDir = config.outputDir;
if (!fs.existsSync(tmpDir)) {
  fs.mkdirSync(tmpDir, { recursive: true });
}

const mp3Path = path.join(tmpDir, 'voice-reply-temp.mp3');
const opusPath = path.join(tmpDir, 'voice-reply.opus');

lib/tts-core.js:78-89:

javascript
if (outputPath) {
  const pcmPath = outputPath.replace(/\.(mp3|opus)$/, '.pcm');
  fs.writeFileSync(pcmPath, pcmAudio);

  const { execSync } = require('child_process');

  if (outputPath.endsWith('.mp3')) {
    execSync(`ffmpeg -y -f s16le -ar 16000 -ac 1 -i ${pcmPath} -ar ${this.audioConfig.outputSampleRate} -ac 1 ${outputPath} 2>/dev/null`);
  } else if (outputPath.endsWith('.opus')) {
    execSync(`ffmpeg -y -f s16le -ar 16000 -ac 1 -i ${pcmPath} -c:a libopus -b:a ${this.audioConfig.outputBitrate} -ar ${this.audioConfig.outputSampleRate} -ac 1 ${outputPath} 2>/dev/null`);
  }

  fs.unlinkSync(pcmPath);
}

Technical Analysis

All executions use the same predictable paths under /tmp/openclaw, including voice-reply-temp.mp3, voice-reply-temp.pcm, and voice-reply.opus. The directory is created recursively without explicitly setting restrictive permissions or verifying its owner and type. Existing files are overwritten without checking whether they are symbolic links or files belonging to another process.

This creates several related risks:

  • A local attacker may pre-create a predictable path as a symbolic link.
  • Multiple voice-reply executions may overwrite each other.
  • Another local process may read or replace generated audio before it is sent.
  • A pre ...[truncated 1598 chars]
Remediation
View remediation

Remediation Suggestions

  1. Create a unique private directory for every execution:
javascript
const os = require('os');
const runDir = fs.mkdtempSync(path.join(os.tmpdir(), 'openclaw-voice-'));
fs.chmodSync(runDir, 0o700);
  1. Generate unique PCM, MP3, and Opus filenames inside that directory rather than reusing global names.
  2. Open files with restrictive permissions such as 0600, exclusive creation, and no-follow behavior where the platform supports it.
  3. Verify that the temporary directory is a real directory, is owned by the expected account, and is not a symbolic link.
  4. Use try/finally to remove every intermediate file and the per-run directory on both success and failure.
  5. Avoid publishing a generated path until synthesis and conversion have completed successfully.
  6. If the message tool requires files under /tmp/openclaw, create a private per-run subdirectory there with mode 0700 and validate its ownership before use.
  7. Introduce concurrency tests to ensure simultaneous requests cannot overwrite or transmit each other’s output.

T08 · Insecure Dependencies

Note
Location
references/setup.md:28
Finding

Unpinned WebSocket Runtime Dependency

Content
View full analysis

Vulnerability Details

File Location: references/setup.md:28-32
Vulnerability Type: Non-reproducible third-party dependency installation
Risk Level: Low

Vulnerable Code

markdown
### 4. Node.js 依赖

```bash
npm install ws
text

### Technical Analysis

The setup documentation instructs users to install `ws` without specifying a reviewed version. The project structure contains no reviewed package manifest or lockfile establishing an exact dependency resolution. Consequently, installations performed at different times may retrieve different package versions.

The package name is not a demonstrated typosquat, and the audit found no evidence that the currently published `ws` package is malicious. The risk is the unsafe, mutable dependency-resolution process: a future compromised release, registry-account compromise, or incompatible update could change the code loaded by `lib/tts-core.js` after this Skill has been reviewed.

### Attack Path

1. A user follows the setup instruction and runs `npm install ws`.
2. npm resolves the dependency version available under the active registry and current resolution rules.
3. If the resolved package or one of its dependencies has been compromised, the package may execute lifecycle code during installation or malicious code when imported.
4. `lib/tts-core.js` loads the installed package using `require('ws')`.
5. Compromised dependency code executes with the permissions of the installation or Agent process.

This path is conditional on compromise or vulnerability of the remotely resolved dependency; no such compromise was established during this audit.

### Impact Assessment

A malicious dependency could execute code with the privileges of the user running npm or the Agent, access environment variables containing Xunfei credentials, inspect accessible files, alter generated audio, or communicate over the network.

The practical severity is reduced because exploitation depends on an external supply-chain co
...[truncated 60 chars]
Remediation
View remediation

Remediation Suggestions

  1. Add a package.json that declares an exact, reviewed ws version rather than a floating version range.
  2. Generate and commit a lockfile containing integrity hashes.
  3. Change installation guidance to use npm ci from the locked dependency graph.
  4. Review dependency advisories and update the pinned version through a controlled process.
  5. Consider disabling install scripts where compatible with deployment requirements:
bash
npm ci --ignore-scripts
  1. Configure the expected npm registry explicitly in controlled deployment environments and verify package provenance and integrity before release.
  2. Include dependency scanning in continuous integration and regenerate the lockfile only as part of a reviewed update.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (14)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill uses environment-provided secrets and executable/script capabilities but does not declare an explicit tool scope or permissions boundary. That makes the skill harder to audit and can lead an agent runtime to grant broader access than necessary, especially since the documented flow reads env vars and runs a local Node script.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger phrases include broad natural-language terms like '用语音' and '语音模式', which can appear in ordinary conversation and accidentally activate the skill. In this skill, accidental activation is more sensitive because it can switch persistent reply mode and cause future replies to be sent as audio to Feishu.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The mode-switching instructions do not define clear scope, quoting/mention exclusions, or whether the words must be standalone commands. Without such boundaries, the agent may interpret incidental text as a command and persistently modify USER.md, changing future behavior without deliberate user intent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill directs the agent to persist a reply-mode flag in USER.md and to send generated audio externally to Feishu, but the user-facing instructions do not clearly require notice or consent for those actions. This is risky because it combines persistent state changes with outbound delivery, so a user may not realize their mode preference is stored or that content is being transmitted to an external service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The config hard-codes a Chinese voice (x4_xiaoyan) and all available voice labels are Chinese-language voices, which indicates a fixed language/locale choice. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The file uses Chinese-language comments and labels throughout, including the module description and configuration notes, without indicating that the skill is region-specific or offering a language choice. Under the stated policy, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill transmits the provided text to Xunfei's remote TTS service over WebSocket for synthesis. In a voice-reply skill this is functionally expected, but it still creates a real privacy and data-handling risk because potentially sensitive user content leaves the local environment without any evident consent, minimization, or disclosure in this file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

When outputPath is provided, the code writes a temporary PCM file, executes ffmpeg via child_process.execSync, and then deletes the temporary file. These are safety-relevant filesystem and subprocess operations, but the file provides no visible confirmation prompt, user-facing log, or explanatory warning about these side effects.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code builds shell command strings for ffmpeg using the unquoted, unsanitized outputPath value and executes them with execSync. If outputPath is influenced by upstream user input or message content, this can lead to command injection and arbitrary local command execution, which is more severe than merely invoking an undeclared subprocess.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · references/setup.md (reported line 23)May include surrounding context.

bash
# Ubuntu/Debian
sudo apt install ffmpeg

# macOS
brew install ffmpeg

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The setup instructs the agent to send generated reply text to the Xunfei TTS service, which is an external third party, but does not include any explicit privacy notice, consent requirement, or guidance to avoid sending sensitive content. In an agent context, replies may include user-provided personal, confidential, or regulated data, so silently forwarding that content off-platform creates a real data disclosure risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The flow explicitly sends generated reply text to the external 讯飞 TTS service, which means message content leaves the local environment and is disclosed to a third party. Because the design describes voice mode as a general reply path and does not mention any user-facing notice, consent, or content filtering, sensitive user data could be transmitted unexpectedly during normal use.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script's natural-language description and example usage are entirely in Chinese, indicating a fixed language/locale expectation. Under the policy, forcing a specific language without offering user choice or documenting a justified regional constraint is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The module reads TTS service credentials from environment variables, which is sensitive configuration access. Although the file has comments about configuration precedence, it does not include any user-facing warning, prompt, or explanatory note that the skill depends on and will access these credentials.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec, suspicious.exposed_secret_literal

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
lib/tts-core.js:87

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/voice-reply.js:38

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
lib/tts-core.js:13