Back to skill

Security audit

Multimodal Base

Security checks for vulnerabilities and agentic risk

Overview

This multimodal skill does what it claims, but it has under-scoped file handling that can delete unrelated files and can upload sensitive media/API credentials to configurable external endpoints.

Review before installing in sensitive environments. Only use trusted API endpoints and API keys, avoid processing private recordings or images unless third-party transmission is acceptable, keep outputDir dedicated to generated audio only, avoid invoking cleanup on shared directories, and update dependencies/model download integrity controls before production use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (6)

T09 · Insecure Skill Coding Practices

Error
Location
src/speech-synthesizer.js:255
Finding

Unrestricted cleanup can delete unrelated files

Content
View full analysis
maxAge) { await fs.unlink(filePath); } } } catch (error) { console.warn('Cleanup failed:', error); } } ``` ### Technical Analysis The constructor accepts an arbitrary writable directory as `outputDir`. The `cleanup()` method then enumerates that directory and deletes every entry older than the supplied threshold. It does not verify that: - The directory is a dedicated TTS output directory. - The target path is inside an approved application-owned root. - The file was created by this component. - The filename follows the generated TTS naming convention. - The target is a regular file rather than a symbolic link or another unexpected entry. The caller also controls `maxAge`. A value such as `0` makes virtually every existing file eligible for deletion. ### Attack Path 1. An attacker or untrusted configuration source sets `outputDir` to a sensitive writable directory. 2. The application constructs `SpeechSynthesizer` with that configuration. 3. The attacker or another reachable application path invokes `cleanup(0)`. 4. The method enumerates the configured directory. 5. Each eligible entry is passed to `fs.unlink()` without confirming that it is a generated TTS artifact. 6. Unrelated files are deleted using the privileges of the Node. ...[truncated 407 chars]
Remediation
View remediation
.mp3`. - Use `lstat()` and reject symbolic links and non-regular files. - Do not permit untrusted callers to provide arbitrary `maxAge` values. - Return cleanup errors to the caller or record them through structured security logging instead of silently continuing. ]]>

T09 · Insecure Skill Coding Practices

Warning
Location
src/speech-synthesizer.js:113
Finding

Predictable temporary files allow overwrite races and sensitive plaintext retention

Content
View full analysis
{}); return result; } catch (error) { // 清理临时文件 await fs.unlink(tmpPath).catch(() => {}); throw error; } } ``` ```js async synthesizeSSML(ssml, options = {}) { this.emit('synthesizing', { ssml }); try { const voice = options.voice || this.defaultVoice; const outputFile = path.join( this.outputDir, `tts_ssml_${Date.now()}.mp3` ); await fs.mkdir(this.outputDir, { recursive: true }); // 将 SSML 保存为临时文件 const ssmlFile = path.join(this.outputDir, `tmp_${Date.now()}.ssml`); await fs.writeFile(ssmlFile, ssml, 'utf-8'); // 使用 edge-tts const args = [ '--voice', voice, '--file', ssmlFile, '--write-media', outputFile ]; await this.runEdgeTTS(args); // 清理临时文件 await fs.unlink(ssmlFile).catch(() => {}); const result = { ssml: ssml, audioPath: outputFile, voice: voice, duration: await this.getAudioDuration(outputFile) }; this.emit('completed', { result }); return result; } catch (error) { this.emit('error', { error }); throw new Error(`SSML synthesis failed: ${error.message}`); } } ``` ### Technical Analysis Temporary filenames are derived only from `Date. ...[truncated 1633 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
src/multimodal-pipeline.js:91
Finding

Image and audio size validation is implemented but bypassed by processing entry points

Content
View full analysis
20 * 1024 * 1024) { throw new Error('Image too large (max 20MB)'); } return true; } catch (error) { throw new Error(`Image validation failed: ${error.message}`); } } ``` ```js async recognizeAPI(audioPath) { const formData = new FormData(); formData.append('file', await fs.readFile(audioPath), path.basename(audioPath)); formData.append('model', 'whisper-1'); formData.append('language', this.language); formData.append('response_format', 'verbose_json'); const response = await axios.post( 'https://api.openai.com/v1/audio/transcriptions', formData, { headers: { ...formData.getHeaders(), 'Authorization': `Bearer ${this.apiKey}` }, timeout: 60000, maxBodyLength: Infinity, maxContentLength: Infinity } ); ``` ```js async validateAudio(audioPath) { try { const stats = await fs.stat(audioPath); if (!stats.isFile()) { throw new Error('Path is not a file'); } ...[truncated 2142 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
src/speech-recognizer.js:139
Finding

Streaming speech recognition buffers unbounded input in memory

Content
View full analysis
{}); return result; } catch (error) { // 清理临时文件 await fs.unlink(tmpPath).catch(() => {}); throw error; } } ``` ### Technical Analysis `recognizeStream()` retains every incoming chunk in an array and concatenates all chunks into a second allocation after the stream ends. No cumulative byte limit, stream duration limit, inactivity timeout, or cancellation mechanism is applied. The existing `validateAudio()` function cannot mitigate this condition because it operates on a completed file and is not invoked before the stream is buffered. A non-terminating stream also prevents the method from progressing to recognition or cleanup. ### Attack Path 1. An attacker submits a large or continuously producing readable stream. 2. Each incoming chunk is appended to the `chunks` array. 3. The function never enforces the documented audio-size limit. 4. If the stream ends, `Buffer.concat(chunks)` creates another allocation for the complete content. 5. Process memory is exhausted or garbage-collection pressure makes the service unavailable. 6. If the stream does not end, the operation remains occupied indefinitely. ### Impact Assessment An attacker with access to the streaming recognition interface can consume process m ...[truncated 215 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
src/speech-recognizer.js:215
Finding

Whisper model is downloaded from a mutable remote location without integrity verification

Content
View full analysis
{ writer.on('finish', resolve); writer.on('error', reject); }); } ``` ### Technical Analysis The download URL references the mutable `main` branch and the resulting file is accepted without validating a checksum, signature, expected size, or immutable revision. The response is written directly to the final model path rather than to a temporary file followed by atomic verification and rename. The method also does not explicitly remove a partially downloaded file when the response or write fails. Consequently, an upstream change, repository compromise, proxy compromise, or interrupted transfer can leave untrusted or corrupted data at the trusted model location. ### Attack Path 1. Local recognition is enabled and `downloadModel()` is invoked. 2. The application downloads the current artifact referenced by the mutable `main` path. 3. The remote artifact has changed, has been compromised, or the transfer is interrupted. 4. The application performs no cryptographic integrity check. 5. The downloaded or partial file remains at `modelPath`. 6. A later local-recognition invocation passes that file to the Whisper executable. ### Impact Assessment The direct confirmed impact is model corruption, incorrect recogn ...[truncated 307 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
src/image-processor.js:23
Finding

Configurable image API endpoint receives the bearer credential and image data without trust validation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (37)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 30)May include surrounding context.

md
import { MultimodalPipeline } from './src/multimodal-pipeline.js';

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 121)May include surrounding context.

md
import ImageProcessor from './src/image-processor.js';

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 142)May include surrounding context.

md
import SpeechRecognizer from './src/speech-recognizer.js';

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 159)May include surrounding context.

md
import SpeechSynthesizer from './src/speech-synthesizer.js';

Known Vulnerable Dependency: axios==1.13.6 — 16 advisory(ies): CVE-2026-44494 (axios Vulnerable to Full Man-in-the-Middle via Prototype Pollution Gadget in `co); CVE-2026-44495 (axios Vulnerable to Credential Theft and Response Hijacking via Prototype Pollut); CVE-2025-62718 (Axios has a NO_PROXY Hostname Normalization Bypass that Leads to SSRF) +13 more

High
Category
Supply Chain
Confidence
97% confidence
Finding

The lockfile pins axios 1.13.6, and the supplied advisory set includes multiple high-risk issues including SSRF/proxy bypass and prototype-pollution-related request/response manipulation. In this skill context, axios is a primary network client, so dependency flaws affecting outbound request routing, header handling, or response integrity are directly relevant and increase real attack surface rather than being theoretical.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: form-data==4.0.5 — 1 advisory(ies): CVE-2026-12143 (form-data: CRLF injection in form-data via unescaped multipart field names and f)

High
Category
Supply Chain
Confidence
85% confidence
Finding

form-data 4.0.5 is reported as vulnerable to CRLF injection through unescaped multipart field names and filenames. If this skill ever builds multipart requests from user-controlled values, an attacker may be able to smuggle additional headers or alter request bodies, which can enable request tampering against downstream services.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: ws==8.20.0 — 2 advisory(ies): CVE-2026-45736 (ws: Uninitialized memory disclosure); CVE-2026-48779 (ws: Memory exhaustion DoS from tiny fragments and data chunks)

High
Category
Supply Chain
Confidence
91% confidence
Finding

ws 8.20.0 is flagged for memory disclosure and memory-exhaustion denial of service. This matters here because edge-tts depends on ws, so any skill capability that opens WebSocket connections to external services could be exposed to malicious peers or hostile intermediary traffic that triggers data leakage or process instability.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: axios==1.13.6 — 16 advisory(ies): CVE-2026-44494 (axios Vulnerable to Full Man-in-the-Middle via Prototype Pollution Gadget in `co); CVE-2026-44495 (axios Vulnerable to Credential Theft and Response Hijacking via Prototype Pollut); CVE-2025-62718 (Axios has a NO_PROXY Hostname Normalization Bypass that Leads to SSRF) +13 more

High
Category
Supply Chain
Confidence
95% confidence
Finding

The dependency set resolves to a version of axios flagged with multiple advisories, including SSRF and prototype-pollution-related exploitation paths. In a multimodal skill that may fetch remote content or interact with external services, a vulnerable HTTP client meaningfully increases the risk of request forgery, credential leakage, or response manipulation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documentation clearly describes image understanding and speech recognition using OpenAI-backed services, but the notes section does not warn users that image and audio content may be transmitted to external APIs for processing. This creates a real privacy and compliance risk because users may submit sensitive media without informed consent or awareness of third-party processing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Natural-language strings and defaults are hardcoded in Chinese, including the default image-understanding prompt and chart-analysis instructions, and OCR defaults to a Chinese-first language pack. This can violate language/locale policy when users are not given a choice or explicit opt-in for a specific language.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 187)May include surrounding context.

md
super();
    this.apiKey = config.openaiApiKey || process.env.OPENAI_API_KEY;
    this.model = config.model || 'gpt-4o';
    this.baseURL = config.baseURL || 'https://api.openai.com/v1';
  }

  /**

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · src/image-processor.js (reported line 24)May include surrounding context.

js
super();
    this.apiKey = config.openaiApiKey || process.env.OPENAI_API_KEY;
    this.model = config.model || 'gpt-4o';
    this.baseURL = config.baseURL || 'https://api.openai.com/v1';
  }

  /**

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The understand() method base64-encodes the local image and sends the full image contents to an external API endpoint. If callers are not clearly informed or required to opt in, sensitive images such as IDs, screenshots, medical records, or internal charts may be disclosed to a third party, creating a privacy and compliance risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Comments, labels, and the default image prompt are written in Chinese, including the hardcoded prompt 描述这张图片的内容 and exported text labels such as 用户 and 助手. This imposes a specific language/locale in natural-language behavior without any opt-in or alternative locale selection.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The code recognizes audio from a file, stores the transcript and metadata on the message object, and then saves that message into persistent in-memory history. While there are developer comments, there is no user-facing prompt, log, or visible disclosure that speech content will be transcribed and retained.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

When options.speak is enabled, the code synthesizes speech and records ttsResult.audioPath in the output message, indicating creation of an audio file or artifact. The code lacks any user-facing notice, confirmation, or warning about this file-producing behavior beyond internal comments.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The text export path hardcodes 用户 and 助手 as role labels, which forces Chinese-language output for exported conversations. This is a language-policy issue because no user choice or configuration is provided for locale.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code sets the default recognition language to 'zh', and the file-level descriptions are also Chinese-centric. This imposes a specific language/locale by default without explicit user opt-in or a clear statement that the skill is intentionally region-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

In API mode, the module uploads the full audio file to OpenAI's transcription endpoint without any built-in disclosure, consent flow, or guardrail. If used on sensitive recordings, this can cause unintended external disclosure of personal, confidential, or regulated audio data.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
94% confidence
Finding

The code transmits audio content to an external third-party service over the network. In the context of a speech recognizer, this is functionally expected for API mode, but it becomes security-relevant because audio may contain secrets, PII, or sensitive conversations and there is no built-in user warning or restriction.

Content

Scanner excerpt · src/speech-recognizer.js (reported line 122)May include surrounding context.

js
formData.append('response_format', 'verbose_json');

    const response = await axios.post(
      'https://api.openai.com/v1/audio/transcriptions',
      formData,
      {
        headers: {

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The file-level and class/function documentation describe this skill as a speech recognition module that converts audio to text using Whisper. The downloadModel method adds a separate capability to fetch and write a model from Hugging Face over the network, which is not indicated by the stated purpose/documentation and is not necessary to the act of recognizing already-provided audio.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The synthesizer generates MP3 files in an output directory and creates that directory if needed, but the only surrounding comments are implementation-oriented and not user-facing warnings. For a code file, file writes should have some visible disclosure unless clearly covered elsewhere; this file does not indicate such disclosure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The code saves SSML content to a temporary file with fs.writeFile and deletes it with fs.unlink, but there is no visible confirmation prompt, print/log statement, or comment/docstring warning users that their input will be written to disk. This is a safety-relevant file write/delete operation, especially because the SSML may contain user-provided content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The code executes the external edge-tts binary via spawn, and also relies on ffprobe elsewhere, but provides no user-facing notice that local subprocesses will be launched. Subprocess execution is in scope for missing-warning checks when not clearly disclosed.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The generateSSML method always emits xml:lang="zh-CN", which imposes a specific locale in generated content regardless of the user's requested language or voice. This is a natural-language policy issue because the skill forces a locale without offering user choice or documenting a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec, suspicious.env_credential_access, suspicious.exposed_secret_literal

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/speech-recognizer.js:75

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/speech-synthesizer.js:156

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/image-processor.js:22

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/speech-recognizer.js:31

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
src/image-processor.js:22

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
src/speech-recognizer.js:31