Back to skill

Security audit

Gejun Math Coach

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent Chinese math-coaching skill, but its automatic voice workflow asks the agent to run local commands and write conversation-derived audio to a shared predictable temp file without clear user consent or cleanup.

Review this skill before installing if you use voice integrations. The math tutoring content is coherent, but voice playback should be treated as opt-in, and the audio workflow should use a private unique temporary file with cleanup instead of /tmp/gejun_tts.mp3. Users who do not want broad activation should narrow the trigger keywords or invoke the skill explicitly.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:308
Finding

Predictable Shared Temporary File Used for Generated Audio

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 308–312
Vulnerability Type: Predictable temporary file path and unsafe shared-file handling
Risk Level: Medium

Vulnerable code:

bash
python3 ~/.workbuddy/skills/voice-coach/scripts/edge_tts_engine.py \
  "${追问文本}" --voice yunjian --speed 0.95 -o /tmp/gejun_tts.mp3

afplay /tmp/gejun_tts.mp3

Technical Analysis

The documented voice workflow always writes generated audio to the fixed, globally predictable path /tmp/gejun_tts.mp3. Temporary directories such as /tmp are normally shared among users and processes. Reusing one predictable filename creates race conditions between concurrent Skill sessions and permits another local process to pre-create, replace, or monitor the file.

Depending on how the external edge_tts_engine.py script opens its output file, a local attacker may be able to place a symbolic link at that path before generation. If the script follows symbolic links and the invoking account can write to the link target, generated data could overwrite or truncate another file accessible to that account. The subsequent afplay command also reads the same path without verifying file ownership, type, permissions, or integrity.

The file is not deleted after playback. Because its contents are derived from conversation text, leaving it behind may disclose user-provided information to other local processes or later sessions when system permissions permit access.

Attack Path

  1. A local attacker identifies that the Skill always uses /tmp/gejun_tts.mp3.
  2. Before or during voice generation, the attacker creates, replaces, or repeatedly swaps that path.
  3. Possible exploitation outcomes include:
    • The attacker places controlled audio at the path after generation but before playback, causing attacker-selected content to be played.
    • One concurrent session overwrites another session's audio, causing cross-session disclosur ...[truncated 1066 chars]
Remediation
View remediation

Remediation Suggestions

Replace the fixed path with a private, uniquely generated temporary file and guarantee cleanup:

bash
audio_file="$(mktemp "${TMPDIR:-/tmp}/gejun_tts.XXXXXX.mp3")" || exit 1
chmod 600 "$audio_file"
trap 'rm -f -- "$audio_file"' EXIT HUP INT TERM

python3 ~/.workbuddy/skills/voice-coach/scripts/edge_tts_engine.py \
  "${prompt_text}" --voice yunjian --speed 0.95 -o "$audio_file" &&
  afplay "$audio_file"

Additional hardening measures:

  • Prefer a per-user runtime directory with restrictive permissions over a globally shared temporary directory.
  • Verify before playback that the output is a regular file owned by the current user and is not a symbolic link.
  • Keep the path quoted in every command.
  • Use && so playback occurs only if generation succeeds.
  • Delete the file immediately after playback and retain the cleanup trap for error and interruption paths.
  • Ensure the external TTS engine creates output atomically and refuses to follow symbolic links where the platform supports that behavior.
  • Use a distinct temporary file for every concurrent request or session.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (6)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list contains broad phrases such as '数学', '题目', '复习', '学习方法', and '教教我', which are common in ordinary conversation and can cause the skill to activate outside the user's intended context. Over-broad activation increases the chance that the skill's strong instructional constraints and optional tool-use behaviors are injected into unrelated sessions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The skill directs that S7 guidance, the four-question reflection flow, and quotations should use a specific voice presentation whenever the voice integration is installed, without requiring explicit user consent. This can surprise users, create privacy/usability issues in shared environments, and implicitly encourage tool invocation beyond what the user asked for.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill explicitly instructs the agent to run local shell and Python commands to generate and play audio, including invoking a script from a user-local path and launching a local media player. For a math coaching skill, this expands behavior from text tutoring into host-side code execution, which can create unnecessary execution surfaces, path abuse risks, and unintended side effects if the environment or referenced script is compromised.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The examples and instructional dialogue are entirely written in Chinese, and the file does not indicate that language selection is optional or limited to a justified region-specific context. This can violate a language/locale policy when a skill implicitly forces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The protocol is written to require Chinese-language interaction throughout the reflection flow and does not provide a user language choice or fallback. This can exclude users, cause misunderstanding of instructions, and lead to poor or unsafe learning outcomes when the user cannot fully understand the guidance, but it is not a direct security exploit in the traditional sense.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document is entirely framed in Chinese and line L001 explicitly labels it as output templates, implying the skill's responses are expected in Chinese by default. There is no indication anywhere in this file that users may choose another language or locale, which can violate language-choice policy for natural-language behavior.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.