T08 · Insecure Dependencies
- Location
SKILL.md:193- Finding
Downloaded Executable and Model Artifacts Are Not Cryptographically Verified
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This English tutor skill is mostly coherent, but it needs Review because it handles private voice/chat data and credentials while recommending unsafe installs and exposing secrets in a helper CLI.
Install only if you are comfortable configuring third-party Feishu, MiniMax, and optional ASR providers, and avoid entering secrets through config_manager.py command-line arguments. Verify downloaded Piper and model artifacts yourself, prefer a virtual environment with pinned dependencies, and enable Bitable memory only if you accept external storage of chat and learning records.
SKILL.md:193Downloaded Executable and Model Artifacts Are Not Cryptographically Verified
SKILL.md:203Python Dependencies Are Installed Without Version or Hash Pinning
scripts/transcribe.py:50Predictable Shared Temporary Audio File Enables Collisions and File Clobbering
scripts/config_manager.py:138Sensitive Configuration Values Are Accepted Through Process Arguments and Echoed in Plaintext
This code uploads the full audio file to a third-party cloud service, which is a real privacy and data-handling risk if users expect local transcription. In a tutoring/transcription skill context, audio may contain sensitive voice or personal data, and the script provides no runtime disclosure or consent check before transmission.
raise RuntimeError('ASR_API_KEY not set (provider=assemblyai)')
with open(audio_path, 'rb') as f:
resp = requests.post(
'https://api.assemblyai.com/v2/upload',
headers={'Authorization': key},
file={'file': (audio_path, f, 'audio/ogg')}
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
raise RuntimeError(f'AssemblyAI upload failed: {resp.status_code}')
transcript_id = resp.json()['upload_url']
poll = requests.post(
'https://api.assemblyai.com/v2/transcript',
headers={'Authorization': key},
json={'audio_url': transcript_id}
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
import time
while result.get('status') not in ('completed', 'error'):
time.sleep(2)
poll = requests.get(
f"https://api.assemblyai.com/v2/transcript/{result['id']}",
headers={'Authorization': key}
)
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
* - 无任何硬编码默认值 (避免误以为"已配置")
*
* 生产环境:cron job 的 env 字段注入
* 本地开发:.env 文件(不提交到仓库)
*/
const path = require('path');
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
const fs = require('fs');
// ===== 加载 .env(仅用于本地开发,生产靠 cron env 注入)=====
const envPath = path.join(__dirname, '.env');
if (fs.existsSync(envPath)) {
fs.readFileSync(envPath, 'utf8').split('\n').forEach(line => {
const t = line.trim();
Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.
piper_model = os.environ.get('PIPER_MODEL', '')
piper_run_ok = False
if os.path.exists(piper_bin):
env = os.environ.copy()
lib_dir = os.path.dirname(piper_bin)
env['LD_LIBRARY_PATH'] = lib_dir + ('' if not env.get('LD_LIBRARY_PATH') else ':' + env['LD_LIBRARY_PATH'])
r = subprocess.run([piper_bin, '--version'], capture_output=True, timeout=5, env=env)
Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.
piper_model = os.environ.get('PIPER_MODEL', '')
piper_run_ok = False
if os.path.exists(piper_bin):
env = os.environ.copy()
lib_dir = os.path.dirname(piper_bin)
env['LD_LIBRARY_PATH'] = lib_dir + ('' if not env.get('LD_LIBRARY_PATH') else ':' + env['LD_LIBRARY_PATH'])
r = subprocess.run([piper_bin, '--version'], capture_output=True, timeout=5, env=env)
The title and overall description define the skill as an English tutor, but the document does not state that the user can choose another language or opt into this locale constraint. Under the policy, forcing a specific language without user choice can be a natural-language policy violation unless clearly justified as region- or domain-specific.
The skill explicitly supports storing chat history and vocabulary progress in Feishu Bitable, but the documentation only frames this as an optional feature and does not clearly warn users that their conversation content and learning records may be retained externally. This creates a privacy and consent issue: users may share personal text or voice-derived content without understanding that it could be persisted in a third-party system.
The header states the file only handles text generation/business logic and that TTS-related operations are handled by callers or separate scripts. However, ttsVoice() invokes TTS.synthesize(text) directly, which is an actual TTS operation and likely an external service call, not merely returning instructions for a caller.
The system prompt explicitly instructs the model to act as an American English coach and to use natural American short sentences throughout. This imposes a specific language/locale variant on all users without offering a choice or documenting a justified regional constraint.
The comments at L115-L118 say voice synthesis is handled by the caller or Python scripts and that this file returns text/TTS instructions only. In contrast, ttsVoice() immediately attempts TTS.synthesize(text) and returns a real voiceUrl, which contradicts the stated behavior.
The code persistently stores full user inputs, AI replies, and per-word chat logs tied to user_id, with no visible minimization, retention limit, consent check, or redaction. In a chat/voice tutoring context, users may disclose personal, sensitive, or regulated data during conversation, so retaining verbatim content increases privacy risk and the blast radius of any memory-store compromise.
The top-level natural-language description identifies the module as 英语助教 and the in-file documentation is entirely in Chinese, which suggests a fixed language/locale assumption. Under the policy, forcing a specific language without user opt-in or a documented regional justification can be a natural-language policy violation.
The code persists user-linked learning history, including vocabulary, review cadence, mastery, and timestamps, to an external Bitable service. While this may be functional for the skill, the absence of visible notice, consent, or minimization controls makes it a real privacy issue because it builds a behavioral learning profile tied to a user identifier.
The module sends user chat content, identifiers, timestamps, and optional voice URLs to an external Feishu Bitable service without any visible consent, minimization, or access-control safeguards in this file. This creates a privacy and data-governance risk because sensitive conversation data may be stored off-platform and linked to a user identity.
The top-of-file documentation states the module is for "TTS 语音合成 + Token Plan 检查" and links only TTS API documentation. However, the same file later defines MiniMaxAPI.chat(), which performs text chat completions against a separate LLM endpoint. This is an active mismatch between documented intent and implemented capability, not merely an omitted implementation detail.
The synthesize flow packages the input text into a payload and submits it to api.minimaxi.com for remote TTS generation. Although network use is part of the feature, there is no confirmation prompt or clear user-facing disclosure here that the provided text will leave the local environment.
With no manifest available, the only stated purpose is the module documentation describing TTS synthesis and token quota checks. Implementing a separate chat-completion client materially expands the skill's capability beyond that stated purpose, including sending arbitrary prompts and messages to a remote LLM endpoint. This is not an obvious implementation detail of TTS synthesis.
The scheduled actions specify "Asia/Shanghai" as the timezone for all recurring tasks, and the manifest does not indicate that this is optional, configurable, or limited to a China-specific deployment. Under the policy, forcing a locale setting without user opt-in is a natural-language policy violation.
The documentation explicitly describes sending synthesized audio and a user's Feishu Open ID to external Feishu APIs, but it does not present any clear user-facing privacy notice, consent step, data retention note, or disclosure of what leaves the local environment. In a tutoring skill that may process learner speech/content, this omission can lead to users unknowingly sharing personal identifiers and audio-related data with third-party services.
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
sudo apt-get install espeak-ng
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def run(cmd, timeout=10):
"""运行外部命令,cmd 为列表,无 shell=True"""
try:
r = subprocess.run(cmd, capture_output=True, timeout=timeout)
return r.returncode == 0, r.stdout + r.stderr
except Exception as e:
return False, str(e)
Dynamic import() can load arbitrary modules at runtime, bypassing static analysis and potentially importing malicious code.
def can_import(mod):
try:
__import__(mod)
return True
except ImportError:
return False
This code file emits its status and remediation messages in Chinese for user interaction, for example the main banner and subsequent diagnostic output. The file does not provide any user opt-in, language selection, or justification that the tool is intended only for a Chinese-speaking or region-specific environment, which creates a natural-language locale policy issue.
Detected: suspicious.exposed_secret_literal