Back to skill

Security audit

The English Tutor

Security checks for vulnerabilities and agentic risk

Overview

This English tutor skill is mostly coherent, but it needs Review because it handles private voice/chat data and credentials while recommending unsafe installs and exposing secrets in a helper CLI.

Install only if you are comfortable configuring third-party Feishu, MiniMax, and optional ASR providers, and avoid entering secrets through config_manager.py command-line arguments. Verify downloaded Piper and model artifacts yourself, prefer a virtual environment with pinned dependencies, and enable Bitable memory only if you accept external storage of chat and learning records.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T08 · Insecure Dependencies

Error
Location
SKILL.md:193
Finding

Downloaded Executable and Model Artifacts Are Not Cryptographically Verified

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:203
Finding

Python Dependencies Are Installed Without Version or Hash Pinning

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/transcribe.py:50
Finding

Predictable Shared Temporary Audio File Enables Collisions and File Clobbering

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/config_manager.py:138
Finding

Sensitive Configuration Values Are Accepted Through Process Arguments and Echoed in Plaintext

Content
View full analysis
1 else "check" if cmd == "get": print(get(sys.argv[2] if len(sys.argv) > 2 else "")) elif cmd == "set": if len(sys.argv) < 4: print("Usage: config_manager.py set ") sys.exit(1) set_(sys.argv[2], sys.argv[3]) print(f"✅ {sys.argv[2]} = {sys.argv[3]}") ``` The same module identifies several fields as sensitive: ```python SENSITIVE = { "feishu_app_id", "feishu_app_secret", "feishu_bot_token", "minimax_api_key", "minimax_group_id", "bitable_app_token", } ``` ### Technical Analysis The `set` operation receives the value through `sys.argv` and then prints it verbatim. If the value is a Feishu secret, MiniMax API key, or Bitable token, it can be exposed in: - shell history; - process listings or process-monitoring telemetry; - terminal scrollback; - CI, cron, or orchestration logs; - screen recordings and support transcripts. Although `save()` redacts listed sensitive values before writing `config.json`, that protection occurs only after the secret has already been exposed through the command line and output. The `get` operation can also print a sensitive value loaded from an environment variable because it applies no redaction based on the requested key. ### Attack Path 1. A user runs `config_manager.py set minimax_api_key ` or sets another sensitive field. 2. The shell records the full command, and the operating system may expose the argument list while the process runs. 3. The script prints the secret to stdout. 4. A local user, monitoring agent, log collector, CI operator, or anyone with terminal/log access retrieves the credential. 5. The exposed credential ...[truncated 669 chars]
Remediation
View remediation
") sys.exit(1) key = sys.argv[2] value = getpass.getpass(f"Enter value for {key}: ") if key in SENSITIVE else input("Value: ") set_(key, value) print(f"Configured {key} successfully.") elif cmd == "get": key = sys.argv[2] if len(sys.argv) > 2 else "" if key in SENSITIVE: print("[REDACTED]") else: print(get(key)) ``` ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (53)

Tainted flow: 'key' from os.environ.get (line 71, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

This code uploads the full audio file to a third-party cloud service, which is a real privacy and data-handling risk if users expect local transcription. In a tutoring/transcription skill context, audio may contain sensitive voice or personal data, and the script provides no runtime disclosure or consent check before transmission.

Content

Scanner excerpt · scripts/transcribe.py (reported line 76)May include surrounding context.

python
raise RuntimeError('ASR_API_KEY not set (provider=assemblyai)')

    with open(audio_path, 'rb') as f:
        resp = requests.post(
            'https://api.assemblyai.com/v2/upload',
            headers={'Authorization': key},
            file={'file': (audio_path, f, 'audio/ogg')}

Tainted flow: 'key' from os.environ.get (line 71, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/transcribe.py (reported line 85)May include surrounding context.

python
raise RuntimeError(f'AssemblyAI upload failed: {resp.status_code}')

    transcript_id = resp.json()['upload_url']
    poll = requests.post(
        'https://api.assemblyai.com/v2/transcript',
        headers={'Authorization': key},
        json={'audio_url': transcript_id}

Tainted flow: 'key' from os.environ.get (line 71, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/transcribe.py (reported line 94)May include surrounding context.

python
import time
    while result.get('status') not in ('completed', 'error'):
        time.sleep(2)
        poll = requests.get(
            f"https://api.assemblyai.com/v2/transcript/{result['id']}",
            headers={'Authorization': key}
        )

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · agent/config.js (reported line 10)May include surrounding context.

js
*   - 无任何硬编码默认值 (避免误以为"已配置")
 *
 * 生产环境:cron job 的 env 字段注入
 * 本地开发:.env 文件(不提交到仓库)
 */

const path = require('path');

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · agent/config.js (reported line 17)May include surrounding context.

js
const fs = require('fs');

// ===== 加载 .env(仅用于本地开发,生产靠 cron env 注入)=====
const envPath = path.join(__dirname, '.env');
if (fs.existsSync(envPath)) {
  fs.readFileSync(envPath, 'utf8').split('\n').forEach(line => {
    const t = line.trim();

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · scripts/check_env.py (reported line 92)May include surrounding context.

python
piper_model = os.environ.get('PIPER_MODEL', '')
piper_run_ok = False
if os.path.exists(piper_bin):
    env = os.environ.copy()
    lib_dir = os.path.dirname(piper_bin)
    env['LD_LIBRARY_PATH'] = lib_dir + ('' if not env.get('LD_LIBRARY_PATH') else ':' + env['LD_LIBRARY_PATH'])
    r = subprocess.run([piper_bin, '--version'], capture_output=True, timeout=5, env=env)

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · scripts/feishu_voice.py (reported line 208)May include surrounding context.

python
piper_model = os.environ.get('PIPER_MODEL', '')
piper_run_ok = False
if os.path.exists(piper_bin):
    env = os.environ.copy()
    lib_dir = os.path.dirname(piper_bin)
    env['LD_LIBRARY_PATH'] = lib_dir + ('' if not env.get('LD_LIBRARY_PATH') else ':' + env['LD_LIBRARY_PATH'])
    r = subprocess.run([piper_bin, '--version'], capture_output=True, timeout=5, env=env)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The title and overall description define the skill as an English tutor, but the document does not state that the user can choose another language or opt into this locale constraint. Under the policy, forcing a specific language without user choice can be a natural-language policy violation unless clearly justified as region- or domain-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill explicitly supports storing chat history and vocabulary progress in Feishu Bitable, but the documentation only frames this as an optional feature and does not clearly warn users that their conversation content and learning records may be retained externally. This creates a privacy and consent issue: users may share personal text or voice-derived content without understanding that it could be persisted in a third-party system.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The header states the file only handles text generation/business logic and that TTS-related operations are handled by callers or separate scripts. However, ttsVoice() invokes TTS.synthesize(text) directly, which is an actual TTS operation and likely an external service call, not merely returning instructions for a caller.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The system prompt explicitly instructs the model to act as an American English coach and to use natural American short sentences throughout. This imposes a specific language/locale variant on all users without offering a choice or documenting a justified regional constraint.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The comments at L115-L118 say voice synthesis is handled by the caller or Python scripts and that this file returns text/TTS instructions only. In contrast, ttsVoice() immediately attempts TTS.synthesize(text) and returns a real voiceUrl, which contradicts the stated behavior.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code persistently stores full user inputs, AI replies, and per-word chat logs tied to user_id, with no visible minimization, retention limit, consent check, or redaction. In a chat/voice tutoring context, users may disclose personal, sensitive, or regulated data during conversation, so retaining verbatim content increases privacy risk and the blast radius of any memory-store compromise.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The top-level natural-language description identifies the module as 英语助教 and the in-file documentation is entirely in Chinese, which suggests a fixed language/locale assumption. Under the policy, forcing a specific language without user opt-in or a documented regional justification can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The code persists user-linked learning history, including vocabulary, review cadence, mastery, and timestamps, to an external Bitable service. While this may be functional for the skill, the absence of visible notice, consent, or minimization controls makes it a real privacy issue because it builds a behavioral learning profile tied to a user identifier.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The module sends user chat content, identifiers, timestamps, and optional voice URLs to an external Feishu Bitable service without any visible consent, minimization, or access-control safeguards in this file. This creates a privacy and data-governance risk because sensitive conversation data may be stored off-platform and linked to a user identity.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The top-of-file documentation states the module is for "TTS 语音合成 + Token Plan 检查" and links only TTS API documentation. However, the same file later defines MiniMaxAPI.chat(), which performs text chat completions against a separate LLM endpoint. This is an active mismatch between documented intent and implemented capability, not merely an omitted implementation detail.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The synthesize flow packages the input text into a payload and submits it to api.minimaxi.com for remote TTS generation. Although network use is part of the feature, there is no confirmation prompt or clear user-facing disclosure here that the provided text will leave the local environment.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

With no manifest available, the only stated purpose is the module documentation describing TTS synthesis and token quota checks. Implementing a separate chat-completion client materially expands the skill's capability beyond that stated purpose, including sending arbitrary prompts and messages to a remote LLM endpoint. This is not an obvious implementation detail of TTS synthesis.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The scheduled actions specify "Asia/Shanghai" as the timezone for all recurring tasks, and the manifest does not indicate that this is optional, configurable, or limited to a China-specific deployment. Under the policy, forcing a locale setting without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The documentation explicitly describes sending synthesized audio and a user's Feishu Open ID to external Feishu APIs, but it does not present any clear user-facing privacy notice, consent step, data retention note, or disclosure of what leaves the local environment. In a tutoring skill that may process learner speech/content, this omission can lead to users unknowingly sharing personal identifiers and audio-related data with third-party services.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · references/config-schema.md (reported line 148)May include surrounding context.

espeak-ng(系统级,完全离线)

bash
sudo apt-get install espeak-ng
env

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/check_env.py (reported line 22)May include surrounding context.

python
def run(cmd, timeout=10):
    """运行外部命令,cmd 为列表,无 shell=True"""
    try:
        r = subprocess.run(cmd, capture_output=True, timeout=timeout)
        return r.returncode == 0, r.stdout + r.stderr
    except Exception as e:
        return False, str(e)

Dynamic import via __import__()

Medium
Category
Dangerous Code Execution
Confidence
75% confidence
Finding

Dynamic import() can load arbitrary modules at runtime, bypassing static analysis and potentially importing malicious code.

Content

Scanner excerpt · scripts/check_env.py (reported line 29)May include surrounding context.

python
def can_import(mod):
    try:
        __import__(mod)
        return True
    except ImportError:
        return False

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code file emits its status and remediation messages in Chinese for user interaction, for example the main banner and subsequent diagnostic output. The file does not provide any user opt-in, language selection, or justification that the tool is intended only for a Chinese-speaking or region-specific environment, which creates a natural-language locale policy issue.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
agent/minimax.js:172