Back to skill

Security audit

火一五 ASR 转写与纪要

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent ASR and meeting-notes helper, with ordinary but important risks around local installs, Hugging Face tokens, and file output.

Install this only if you are comfortable running ffmpeg, Whisper/WhisperX, and Python packages on your machine. Use a virtual environment, avoid command-line Hugging Face tokens, use minimally scoped tokens, verify any external office-doc skill before Word export, and pick fresh output filenames to avoid overwriting files.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:78
Finding

Unpinned Third-Party Dependencies and Unverified External Skill Execution

Content
View full analysis
/skills/huo15-openclaw-office-doc/scripts/create-word-doc.py ` --output "<下载目录>\ASR_纪要_YYYY-MM-DD_HHmmss.docx" ` --title "会议纪要 - YYYY-MM-DD" ` --content @"<临时 markdown 文件路径>" ` --doc-format 会议纪要 ` --company-name "会议纪要" ` --no-title-block ``` ### Technical Analysis The installation commands do not specify reviewed versions, cryptographic hashes, a lockfile, or an explicitly trusted package index. Consequently, the code installed and imported may change between executions. Package installation can execute package build logic, while later imports and command execution run the installed package with the invoking user's privileges. The Word-generation workflow additionally executes `create-word-doc.py` from another Skill that is not included in the audited project. Its implementation and integrity therefore cannot be verified from this package. This does not establish that the external Skill is malicious, but it creates an unverified supply-chain boundary. ### Attack Path 1. An attacker compromises a referenced package, one of its transitive dependencies, or a package index configured in the local pip environment. 2. Alternatively, an attacker modifies or substitutes the separately installed `huo15-openclaw-office-doc` Skill. 3. A user follows the documented installation or Word-generation instructions. 4. Pip installs attacker-controlled package content, or Python executes the substituted external script. 5. Malicious insta ...[truncated 706 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
scripts/transcribe_with_diarization.py:22
Finding

Hugging Face Token Can Be Exposed Through a Command-Line Argument

Content
View full analysis
[--model base] [--language zh] [--hf_token TOKEN] ``` The argument is declared and then preferred over the environment variable: ```python parser.add_argument("--hf_token", default=None, help="HuggingFace Token(也可通过 HF_TOKEN 环境变量设置)") args = parser.parse_args() audio_file = args.audio_file if not os.path.isfile(audio_file): print(f"错误:文件不存在 - {audio_file}") sys.exit(1) # 获取 HF Token hf_token = args.hf_token or os.environ.get("HF_TOKEN", None) ``` ### Technical Analysis Command-line arguments are not a suitable channel for long-lived secrets. Depending on the operating system and execution environment, process arguments may be visible in process listings, shell history, task schedulers, terminal logs, crash reports, monitoring tools, or automation telemetry. The script does not print the token and passes it only to `whisperx.DiarizationPipeline`. No direct token-exfiltration code was found. The weakness is the supported method of supplying the secret rather than deliberate credential collection. ### Attack Path 1. A user invokes the script with `--hf_token TOKEN`. 2. The complete command may be stored in shell history or captured by terminal, job, or monitoring logs. 3. While the process is running, another sufficiently authorized local user or monitoring service may inspect its argument list. 4. The observer retrieves the Hugging Face token. 5. The exposed credential is used against Hugging Face within the permissions and lifetime assigned to that token. ### Impact Assessment An attacker who ob ...[truncated 418 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (13)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code chunk does not implement any audio/video handling, transcription, diarization, meeting summarization, or OpenClaw formatting. Its sole behavior is stripping ANSI terminal escape codes from text input. While this could be a supporting step for processing terminal logs or .tty-derived text, the declared description presents a much broader and materially different primary purpose centered on media transcription and structured analysis. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The code only implements one narrow part of the declared description: transcoding media to MP3. It does not perform transcription, speaker diarization, meeting-note generation, .tty parsing, or any OpenClaw formatting/classification logic. Because the declared purpose presents a broader, composite skill while the actual code chunk only covers MP3 transcoding, the description does not accurately represent what this supplied code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The supplied code chunk has a narrower and materially different behavior than the declared description. It performs ASR transcription plus word-level alignment and speaker diarization on an existing audio file, requiring a HuggingFace token, and outputs plain speaker-labeled transcript text. It does not convert video/audio formats to MP3, does not generate meeting notes or summaries, does not parse .tty files, and does not produce any structured output for OpenClaw to classify or judge. The code is related to part of the declared domain (speech recognition/transcription and speaker diarization), but the overall declared description substantially overstates capabilities not present in this chunk.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill title and body are written as a Chinese-only workflow, and later sections hardcode Chinese document titles and naming such as "会议纪要" and "ASR_转写". While L143 allows language detection for transcription, the skill does not offer a user choice for the language/locale of the summary, delivery formatting, or saved document content, which can violate a language/locale choice policy.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill instructs use of HuggingFace authentication via huggingface-cli login or an HF_TOKEN environment variable, which creates a credential-handling path without guardrails on storage, redaction, or least privilege. In agent environments, guidance to consume env-stored secrets can lead to inadvertent disclosure in logs, downstream tool calls, or generated output if the workflow is not tightly constrained.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Natural-language policy violations include forcing a specific language without user opt-in. This file presents all user-facing instructions in Chinese and does not offer an alternative language or explain that the skill is intentionally region- or locale-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script invokes ffmpeg with the -y flag, which forces overwriting the output file if it already exists. There is no confirmation prompt, warning message, or other user-facing disclosure that an existing file at the destination path may be replaced.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes audio/video conversion, transcription, diarization, and formatting for OpenClaw, but does not mention credential handling or authenticated access. Reading HF_TOKEN from environment variables adds a sensitive-capability surface beyond the stated purpose of a transcription skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

This code performs transcription and speaker diarization on a user-supplied audio file using WhisperX models, including a HuggingFace-authenticated diarization pipeline. While the script prints progress messages, it does not clearly warn users that sensitive audio content may be processed by third-party model components or that a HuggingFace token is involved, which is relevant for privacy-sensitive media.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The stated purpose is ASR/transcription and formatting output for OpenClaw, with optional saving. The Word-generation path expands the capability into cross-skill document production by calling huo15-openclaw-office-doc/scripts/create-word-doc.py, which is not an obvious requirement of speech recognition itself.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

This README documents scripts that write MP3 output files, but it does not include any user-facing warning about file creation or possible overwrite effects. For markdown files, SQP-2 applies when the skill description omits warnings about behaviors that could affect user data or system state.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

Natural-language strings and argument defaults indicate the skill assumes Chinese as the default language (--language zh). This is a locale/language preference baked into the skill behavior rather than a neutral default or explicit user choice, which can violate language-choice policy unless clearly justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.