Back to skill

Security audit

Audio Meeting Minutes

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly purpose-aligned, but it needs review because it handles sensitive meeting recordings, asks for cloud tokens through chat, and can generate unsafe HTML from transcript content.

Review before installing. Use only with recordings your organization permits sending to Alibaba Cloud NLS, provide credentials through environment variables or a secret manager instead of chat, and avoid opening or sharing generated HTML from untrusted recordings until the renderer escapes all transcript and summary fields.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
scripts/summarize.py:45
Finding

Untrusted Transcript Content Is Embedded Directly into Agent Instructions

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/report.py:812
Finding

Stored HTML Injection Through Unescaped Transcript and Summary Data

Content
View full analysis
str: try: from jinja2 import Template except ImportError: return _render_without_jinja(summary, transcript, audio_filename, output_path) context = { "meeting_title": summary.get("meeting_title", "Meeting Minutes"), "meeting_type": summary.get("meeting_type", "Meeting"), "audio_filename": audio_filename, "meeting_summary": summary.get("meeting_summary", ""), "key_conclusions": summary.get("key_conclusions", []), "decisions": summary.get("decisions", []), "action_items": summary.get("action_items", []), "participants": summary.get("participants", []), "agenda_items": summary.get("agenda_items", []), "next_steps": summary.get("next_steps", []), "risks_and_concerns": summary.get("risks_and_concerns", []), "follow_up_required": summary.get("follow_up_required", []), "segments": transcript.get("segments", []), "full_text": transcript.get("text", ""), } template = Template(HTML_TEMPLATE) html = template.render(**context) ``` Representative output sinks in `HTML_TEMPLATE` include: ```html

{{ meeting_title }}

{{ seg.text }}
{{ full_text }}
``` ### Technical Analysis Jinja2's direct `Template` constructor does not automatically enable HTML autoescaping. Consequently, transcript content, AI-generated summary fields, participant names, action items, source filenames, and other values are interpreted as HTML markup when inserted into the report. The transcript originates from externally supplied audio, while summary fields are ...[truncated 1879 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:52
Finding

Access Tokens Are Requested Through the Chat Conversation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
Findings (22)

Known Vulnerable Dependency: jinja2==3.1.4 — 6 advisory(ies): CVE-2025-27516 (Jinja2 vulnerable to sandbox breakout through attr filter selecting format metho); CVE-2024-56201 (Jinja has a sandbox breakout through malicious filenames); CVE-2024-56326 (Jinja has a sandbox breakout through indirect reference to format method) +3 more

Critical
Category
Supply Chain
Confidence
95% confidence
Finding

The skill pins jinja2 to 3.1.4, which is associated with multiple sandbox breakout advisories. This is especially relevant because the skill explicitly generates HTML documents, making template rendering part of the intended functionality; if any user-controlled content, filenames, or metadata reaches Jinja template evaluation in an unsafe way, an attacker may achieve template injection or escape sandbox protections. The dangerousness is increased by the skill context because templating is likely central to operation.

Content

No source excerpt is available for this finding.

Tainted flow: 'headers' from os.environ.get (line 34, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/process_nas.py (reported line 41)May include surrounding context.

python
"enable_inverse_text_normalization": "true",
        "enable_voice_detection": "false",
    }
    resp = requests.post(NLS_URL, headers=headers, params=params, data=data, timeout=120)
    resp.raise_for_status()
    body = resp.json()
    if body.get("status") == 20000000:

Lp1

High
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The script reads sensitive runtime configuration from environment variables, including NLS credentials and a NAS path, but the declared permissions omit environment access. This creates a privilege-transparency gap: reviewers and users may not realize the skill can consume ambient secrets or infrastructure configuration from the host environment.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 27)May include surrounding context.

md
## 首次使用需要准备

- **NLS AppKey**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 创建项目后获得
- **NLS AccessToken**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 项目详情 → AccessToken 标签(24h 有效)

无需配置 AI Key,AI 总结由 WorkBuddy 内置完成。

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 28)May include surrounding context.

md
## 首次使用需要准备

- **NLS AppKey**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 创建项目后获得
- **NLS AccessToken**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 项目详情 → AccessToken 标签(24h 有效)

无需配置 AI Key,AI 总结由 WorkBuddy 内置完成。

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

md
## 首次使用需要准备

- **NLS AppKey**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 创建项目后获得
- **NLS AccessToken**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 项目详情 → AccessToken 标签(24h 有效)

无需配置 AI Key,AI 总结由 WorkBuddy 内置完成。

Session Persistence

Medium
Category
Rogue Agent
Confidence
75% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 60)May include surrounding context.

md
## 首次使用需要准备

- **NLS AppKey**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 创建项目后获得
- **NLS AccessToken**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 项目详情 → AccessToken 标签(24h 有效)

无需配置 AI Key,AI 总结由 WorkBuddy 内置完成。

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README states users must input the Chinese phrase “熠小听” to trigger the skill, which imposes a specific language requirement. There is no indication that other languages are accepted, no user opt-in for locale, and no documented justification that the skill is region-specific.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger list includes broad, natural phrases such as meeting recording and transcription terms that users may say in ordinary conversation. This can cause unintended activation of a skill that has shell, network, and filesystem permissions, potentially leading to unexpected processing of local files or prompting for sensitive credentials.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly instructs users to paste cloud credentials into chat, which risks exposing tokens to chat logs, transcript retention systems, support tooling, and other downstream storage. Even though the text warns against pasting long-lived secrets, it still normalizes insecure secret handling and creates an avoidable credential exposure path.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The rendering example hard-codes the transcript language as "zh", and the surrounding instructions consistently prescribe Chinese-language interaction and output format without indicating that users may choose another language. This is a locale/language policy concern because it enforces a specific language by default rather than documenting a justified region-specific constraint or obtaining user opt-in.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: requests==2.32.3 — 4 advisory(ies): CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2026-25645 (Requests has Insecure Temp File Reuse in its extract_zipped_paths() utility func); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs) +1 more

Medium
Category
Supply Chain
Confidence
90% confidence
Finding

The skill pins requests to 2.32.3, which is flagged with published advisories including a .netrc credential leak via malicious URLs. Because this skill has network permission and is designed to process files automatically, any code path that fetches remote resources with attacker-influenced URLs could expose local credentials or trigger vulnerable library behavior. The issue is dependency risk rather than clear exploit logic in this file, but the vulnerable version is real.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script uploads raw audio content from a NAS share to Alibaba Cloud speech recognition without any explicit in-code consent flow, warning, or data-handling disclosure. In a meeting-minutes context, recordings may contain confidential business discussions or personal data, so silent third-party transmission materially increases privacy and compliance risk.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/process_nas.py (reported line 83)May include surrounding context.

python
print(f"处理: {Path(AUDIO).name}")
    print(f"{'='*60}")

    dur = float(subprocess.check_output(
        ["ffprobe","-i",AUDIO,"-show_entries","format=duration","-v","quiet","-of","csv=p=0"],
        stderr=subprocess.DEVNULL
    ).decode().strip())

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/process_nas.py (reported line 92)May include surrounding context.

python
print(f"\n[1/2] 音频转换 + 切块(每块≤1.8MB)...")
    t0 = time.time()
    pcm_path = os.path.join(TMPDIR, f"{name}_full.pcm")
    subprocess.run([
        "ffmpeg", "-y", "-i", AUDIO,
        "-ac","1","-ar","16000","-f","s16le", pcm_path
    ], capture_output=True)

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes an end-to-end skill that automatically transcribes recordings, extracts meeting points, and outputs a professional HTML meeting-minutes document. In this file, the implemented behavior stops after cloud transcription, local transcript saving, and printing JSON for another agent to consume; no summarization or HTML generation occurs here.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This code writes rendered HTML containing meeting summaries and full transcript content to a file, which can expose sensitive meeting data on local storage. The file-writing path has no confirmation prompt, visible logging, or inline warning comment/docstring disclosing that potentially sensitive content will be persisted.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The natural-language prompt is entirely in Chinese and requires output in Chinese formats such as '书面语言' and Chinese field value conventions, which effectively constrains the skill to a specific language. The file does not indicate that users can choose another language or that the Chinese-only behavior is a justified region-specific requirement.

Content

No source excerpt is available for this finding.

Lp4

Low
Category
MCP Least Privilege
Confidence
65% confidence
Finding

Declared permissions with no matching code capability may indicate removed functionality or pre-staging for future abuse.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The module docstring and runtime messages are entirely in Chinese, indicating a fixed language experience. Under the policy, locale or language should not be forced unless the user is given an explicit choice or the constraint is clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
75% confidence
Finding

The stated purpose is automatic audio-to-text conversion and meeting-minutes generation, but the manifest does not indicate any need to execute external binaries. While audio conversion is related to the task, invoking subprocesses is a stronger capability than the description suggests and is not explicitly scoped in the manifest text.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The HTML template hard-codes a Chinese locale and Chinese-language document labels, causing all generated reports to default to zh-CN regardless of user preference. This is a natural-language policy concern because the skill imposes a specific language/locale without offering a choice or documenting a justified regional constraint.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.