T01 · Skill Instruction Hijacking
- Location
scripts/summarize.py:45- Finding
Untrusted Transcript Content Is Embedded Directly into Agent Instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is mostly purpose-aligned, but it needs review because it handles sensitive meeting recordings, asks for cloud tokens through chat, and can generate unsafe HTML from transcript content.
Review before installing. Use only with recordings your organization permits sending to Alibaba Cloud NLS, provide credentials through environment variables or a secret manager instead of chat, and avoid opening or sharing generated HTML from untrusted recordings until the renderer escapes all transcript and summary fields.
scripts/summarize.py:45Untrusted Transcript Content Is Embedded Directly into Agent Instructions
scripts/report.py:812Stored HTML Injection Through Unescaped Transcript and Summary Data
SKILL.md:52Access Tokens Are Requested Through the Chat Conversation
The skill pins jinja2 to 3.1.4, which is associated with multiple sandbox breakout advisories. This is especially relevant because the skill explicitly generates HTML documents, making template rendering part of the intended functionality; if any user-controlled content, filenames, or metadata reaches Jinja template evaluation in an unsafe way, an attacker may achieve template injection or escape sandbox protections. The dangerousness is increased by the skill context because templating is likely central to operation.
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
"enable_inverse_text_normalization": "true",
"enable_voice_detection": "false",
}
resp = requests.post(NLS_URL, headers=headers, params=params, data=data, timeout=120)
resp.raise_for_status()
body = resp.json()
if body.get("status") == 20000000:
The script reads sensitive runtime configuration from environment variables, including NLS credentials and a NAS path, but the declared permissions omit environment access. This creates a privilege-transparency gap: reviewers and users may not realize the skill can consume ambient secrets or infrastructure configuration from the host environment.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
## 首次使用需要准备
- **NLS AppKey**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 创建项目后获得
- **NLS AccessToken**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 项目详情 → AccessToken 标签(24h 有效)
无需配置 AI Key,AI 总结由 WorkBuddy 内置完成。
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
## 首次使用需要准备
- **NLS AppKey**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 创建项目后获得
- **NLS AccessToken**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 项目详情 → AccessToken 标签(24h 有效)
无需配置 AI Key,AI 总结由 WorkBuddy 内置完成。
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
## 首次使用需要准备
- **NLS AppKey**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 创建项目后获得
- **NLS AccessToken**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 项目详情 → AccessToken 标签(24h 有效)
无需配置 AI Key,AI 总结由 WorkBuddy 内置完成。
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
## 首次使用需要准备
- **NLS AppKey**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 创建项目后获得
- **NLS AccessToken**:[获取链接](https://nls-portal.console.aliyun.com/applist) — 项目详情 → AccessToken 标签(24h 有效)
无需配置 AI Key,AI 总结由 WorkBuddy 内置完成。
The README states users must input the Chinese phrase “熠小听” to trigger the skill, which imposes a specific language requirement. There is no indication that other languages are accepted, no user opt-in for locale, and no documented justification that the skill is region-specific.
The trigger list includes broad, natural phrases such as meeting recording and transcription terms that users may say in ordinary conversation. This can cause unintended activation of a skill that has shell, network, and filesystem permissions, potentially leading to unexpected processing of local files or prompting for sensitive credentials.
The skill explicitly instructs users to paste cloud credentials into chat, which risks exposing tokens to chat logs, transcript retention systems, support tooling, and other downstream storage. Even though the text warns against pasting long-lived secrets, it still normalizes insecure secret handling and creates an avoidable credential exposure path.
The rendering example hard-codes the transcript language as "zh", and the surrounding instructions consistently prescribe Chinese-language interaction and output format without indicating that users may choose another language. This is a locale/language policy concern because it enforces a specific language by default rather than documenting a justified region-specific constraint or obtaining user opt-in.
The skill pins requests to 2.32.3, which is flagged with published advisories including a .netrc credential leak via malicious URLs. Because this skill has network permission and is designed to process files automatically, any code path that fetches remote resources with attacker-influenced URLs could expose local credentials or trigger vulnerable library behavior. The issue is dependency risk rather than clear exploit logic in this file, but the vulnerable version is real.
The script uploads raw audio content from a NAS share to Alibaba Cloud speech recognition without any explicit in-code consent flow, warning, or data-handling disclosure. In a meeting-minutes context, recordings may contain confidential business discussions or personal data, so silent third-party transmission materially increases privacy and compliance risk.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
print(f"处理: {Path(AUDIO).name}")
print(f"{'='*60}")
dur = float(subprocess.check_output(
["ffprobe","-i",AUDIO,"-show_entries","format=duration","-v","quiet","-of","csv=p=0"],
stderr=subprocess.DEVNULL
).decode().strip())
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
print(f"\n[1/2] 音频转换 + 切块(每块≤1.8MB)...")
t0 = time.time()
pcm_path = os.path.join(TMPDIR, f"{name}_full.pcm")
subprocess.run([
"ffmpeg", "-y", "-i", AUDIO,
"-ac","1","-ar","16000","-f","s16le", pcm_path
], capture_output=True)
The manifest describes an end-to-end skill that automatically transcribes recordings, extracts meeting points, and outputs a professional HTML meeting-minutes document. In this file, the implemented behavior stops after cloud transcription, local transcript saving, and printing JSON for another agent to consume; no summarization or HTML generation occurs here.
This code writes rendered HTML containing meeting summaries and full transcript content to a file, which can expose sensitive meeting data on local storage. The file-writing path has no confirmation prompt, visible logging, or inline warning comment/docstring disclosing that potentially sensitive content will be persisted.
The natural-language prompt is entirely in Chinese and requires output in Chinese formats such as '书面语言' and Chinese field value conventions, which effectively constrains the skill to a specific language. The file does not indicate that users can choose another language or that the Chinese-only behavior is a justified region-specific requirement.
Declared permissions with no matching code capability may indicate removed functionality or pre-staging for future abuse.
The module docstring and runtime messages are entirely in Chinese, indicating a fixed language experience. Under the policy, locale or language should not be forced unless the user is given an explicit choice or the constraint is clearly justified as region-specific.
The stated purpose is automatic audio-to-text conversion and meeting-minutes generation, but the manifest does not indicate any need to execute external binaries. While audio conversion is related to the task, invoking subprocesses is a stronger capability than the description suggests and is not explicitly scoped in the manifest text.
The HTML template hard-codes a Chinese locale and Chinese-language document labels, causing all generated reports to default to zh-CN regardless of user preference. This is a natural-language policy concern because the skill imposes a specific language/locale without offering a choice or documenting a justified regional constraint.
No suspicious patterns detected.