Back to skill

Security audit

Aliyun Asr

Security checks for vulnerabilities and agentic risk

Overview

This voice transcription skill is purpose-aligned, but it needs Review because it uploads voice audio to Aliyun, recommends broad cloud permissions, and can leave converted audio files on disk despite claiming no local storage.

Review before installing. Use a dedicated low-privilege Aliyun RAM user, rotate and protect the AccessKey, enable the skill only in channels where users understand voice messages will be sent to Aliyun, and treat converted audio files as sensitive until cleanup behavior is fixed.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
aliyun_pure_asr.py:89
Finding

Converted Voice Recordings Persist in Predictable Local Files

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
SKILL.md:27
Finding

Aliyun Setup Instructions Grant Excessive NLS Permissions

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (13)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The supplied code chunk does not implement the declared ASR/transcription functionality at all. While it also does not show any undeclared or harmful capability, the actual behavior visible here is effectively a no-op placeholder, which is materially different from the declared primary purpose of voice message transcription. Because the assessment is limited to the provided chunk, this should be flagged as a mismatch between description and demonstrated behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The setup and usage text does not clearly warn that user audio will be transmitted to Aliyun for third-party processing, which is a material privacy and compliance concern. In voice workflows, silent external transfer of user content can expose sensitive personal or business information and may violate consent, policy, or regulatory requirements.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill metadata declares no explicit tool scope or permissions, yet the documented behavior clearly implies reading a local config file, making outbound network requests to Aliyun, and potentially invoking shell-related setup steps. In an agent ecosystem, missing scope declarations weaken reviewability and sandbox enforcement, increasing the risk that the skill can access resources beyond what users or operators expect.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 46)May include surrounding context.

3. 安全设置

bash
chmod 600 /root/.openclaw/aliyun-asr-config.json

🚀 使用方法

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill is described as automatically activating for voice messages from any supported channel, without clear boundaries, consent gates, or per-channel constraints. Overly broad auto-triggering can cause unintended capture and transmission of user audio, increasing privacy and misuse risks, especially in multi-channel deployments where expectations differ.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · aliyun_pure_asr.py (reported line 93)May include surrounding context.

python
if audio_file.endswith('.ogg'):
                wav_file = audio_file.replace('.ogg', '.wav')
                # 转换为16kHz单声道WAV
                subprocess.run([
                    'ffmpeg', '-i', audio_file, '-ar', '16000', '-ac', '1', 
                    '-f', 'wav', wav_file
                ], check=True, capture_output=True)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill sends raw audio content to a remote Aliyun ASR endpoint, but the CLI usage and code provide no explicit notice that potentially sensitive voice data leaves the local environment. In a transcription skill this behavior is expected functionally, but the absence of disclosure creates a privacy and data-handling risk, especially for confidential or regulated audio.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The module docstring states in Chinese that this is an 'Alibaba Cloud pure ASR skill' and that it only provides speech recognition. The file provides no indication that users can choose another language or locale, which can be a natural-language policy issue when a skill appears to enforce a specific language context without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The only natural-language documentation in this file is written in Chinese and does not indicate that language selection is optional or tied to a documented regional requirement. Under the language/locale policy rule, this can be considered a locale constraint presented without user opt-in or justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The script reads long-lived cloud credentials from /root/.openclaw/aliyun-asr-config.json, including access_key_secret. While the code prints the expected fields, it does not warn that the file contains sensitive credentials or advise secure storage and restricted permissions.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The manifest and module docstring describe a pure voice-to-text skill with no extra functionality, but the implementation invokes an external program via subprocess to perform media conversion. While related to transcription, process execution is a distinct capability that is not explicitly justified by the stated scope and expands operational risk beyond a straightforward API client.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest-like JSON sets the region to "cn-shanghai", which is a locale-specific setting expressed in natural language/config text. Under the policy rule, forcing a specific locale or region without user choice or documented justification can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The module docstring and function docstring describe the skill exclusively in Chinese, including the interface purpose and argument semantics. This creates a language-specific constraint in the natural-language surface of the skill without indicating that users can choose another language or that the limitation is intentional and justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.