Back to skill

Security audit

Funasr Transcribe Skill

Security checks for vulnerabilities and agentic risk

Overview

This skill does what it claims: it installs a local FunASR transcription environment, processes chosen audio locally, and writes a local transcript, with disclosed setup and download side effects.

Install only if you trust the configured PyPI mirror, FunASR dependencies, and model providers. Be aware that setup and first use may download code and models, create a persistent local environment, consume several GB of disk space, and overwrite or create a .txt transcript next to the selected audio file.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
scripts/install.sh:54
Finding

Unpinned Dependencies Installed from a Third-Party Package Mirror

Content
View full analysis

Vulnerability Details

File Location: scripts/install.sh, lines 54-60
Vulnerability Type: Unpinned third-party dependencies and unsafe supply-chain trust
Risk Level: Medium

Vulnerable Code

bash
# 升级 pip
echo ""
echo "升级 pip..."
pip install --upgrade pip -i https://pypi.tuna.tsinghua.edu.cn/simple

# 安装依赖
echo ""
echo "安装 FunASR 及依赖(这需要几分钟)..."
pip install funasr modelscope huggingface_hub torch torchaudio \
    -i https://pypi.tuna.tsinghua.edu.cn/simple

Technical Analysis

The installer downloads and installs the latest available versions of pip, funasr, modelscope, huggingface_hub, torch, torchaudio, and their transitive dependencies from a third-party PyPI mirror. It does not enforce package versions, artifact hashes, or a reviewed lock file.

Python package installation may execute package build and installation logic. Installed packages may also execute code when imported. Consequently, the effective code executed by this Skill can change without any modification to the audited repository.

The dependency and model downloads are legitimate requirements of the declared transcription functionality. However, unrestricted upgrades, unpinned dependency resolution, and reliance on a third-party mirror create avoidable supply-chain exposure beyond the minimum trust necessary.

Attack Path

  1. An attacker compromises the configured package mirror, an upstream package release, or a transitive dependency.
  2. The attacker publishes or substitutes a malicious package version that satisfies the unconstrained installation request.
  3. A user or autonomous agent runs scripts/install.sh.
  4. pip retrieves and installs the attacker-controlled artifact.
  5. Malicious code executes during package installation, verification import, or later transcription.
  6. The code operates with the permissions and data access of the account running the Skill.

Impact Assessment

Su ...[truncated 704 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin every direct dependency to a reviewed, exact version.
  2. Generate and commit a lock file containing resolved transitive dependencies.
  3. Use hashes for every permitted artifact and install with pip --require-hashes.
  4. Pin pip itself instead of using an unrestricted --upgrade.
  5. Prefer the official PyPI index, or clearly require explicit user trust and consent before using the configured mirror.
  6. Separate dependency installation from transcription and require clear user approval before initiating network downloads.
  7. Pin model downloads to immutable revisions and verify checksums where supported.
  8. Periodically review and update dependency pins through a controlled security-update process.
  9. Run installation and transcription in a sandbox with restricted filesystem and network access when processing sensitive recordings.
Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (8)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill performs file-writing behavior by documenting that transcripts are written to sibling .txt files, yet it declares no explicit tool scope or permissions boundary. Without declared limits, an agent may invoke file-modifying behavior without clear authorization controls, increasing the risk of unintended writes or abuse if paired with broader agent capabilities.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
88% confidence
Finding

The skill creates a persistent virtual environment under the user's home/workspace path, leaving software and artifacts behind across sessions. Persistent state is risky because it can silently alter future executions, consume storage, and create an unexpected foothold for stale or tampered dependencies if the environment is reused later.

Content

Scanner excerpt · SKILL.md (reported line 29)May include surrounding context.

Quick Start

bash
# Install dependencies and create a virtual environment
bash ~/.openclaw/workspace/skills/funasr-transcribe/scripts/install.sh

# Transcribe an audio file

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly states that autonomous invocation is normal and allows dependency installation and model downloads unless the user opts out. This is dangerous because it authorizes software installation and network retrieval of third-party packages/models without explicit opt-in, which can lead to supply-chain exposure, unexpected system changes, and policy violations in restricted environments.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This shell script presents its status, errors, and usage guidance entirely in Chinese, which effectively forces a specific language for users. The policy allows locale constraints only when they are optional, user-selectable, or clearly justified as region-specific, none of which is indicated here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script's docstrings, usage text, status messages, and errors are presented in Chinese only, which imposes a specific language on users. The file does not provide any opt-in, fallback, or documented reason that the skill must be Chinese-only.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script and manifest present the capability as local transcription with no direct external endpoints, but model initialization can trigger first-run network downloads from external model repositories. This creates an integrity and privacy boundary mismatch: users may run the skill expecting offline-only behavior, while execution can unexpectedly reach out to the network and fetch unpinned remote artifacts.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The script's comments and all user-visible error/usage messages are in Chinese, including the usage text and failure guidance. This imposes a specific language/locale on users without opt-in or justification, which matches the natural-language policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

该文件主体内容以中文编写,并在此处仅将英文说明重定向到另一个文件,体现出当前 skill 文档对语言有默认限定。根据规则,若技能在自然语言层面强制特定语言而未在本文件中提供用户选择或明确 opt-in,可视为语言/locale 政策风险。

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.