Back to skill

Security audit

Audio2srtlocal

Security checks for vulnerabilities and agentic risk

Overview

The skill builds the advertised local transcription app, but its generated backend exposes unauthenticated network APIs that can read local audio paths and consume disk or compute resources.

Install only if you are comfortable reviewing or patching the generated backend first. At minimum, bind it to 127.0.0.1, restrict CORS to the frontend origin, add a session token, remove arbitrary file_path processing, cap uploads, clean temporary files, and use pinned dependencies or lockfiles before running it on a shared network.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (3)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
references/transcribe_server.py:407
Finding

Unauthenticated Network API Permits Processing of Arbitrary Local Audio Files

Content
View full analysis
web.Application: app = web.Application(middlewares=[cors_middleware]) app.router.add_get("/api/health", health_handler) app.router.add_get("/api/models", models_handler) app.router.add_post("/api/upload", upload_handler) app.router.add_post("/api/transcribe", transcribe_handler) app.router.add_get("/api/tasks/{task_id}", task_status_handler) app.router.add_post("/api/tasks/{task_id}/cancel", cancel_task_handler) app.router.add_post("/api/translate", translate_handler) return app ``` ```python web.run_app(create_app(), host="0.0.0.0", port=port, print=None) ``` ### Technical Analysis The backend listens on every available network interface and does not authenticate any API route. The transcription endpoint accepts a client-control ...[truncated 2190 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
references/transcribe_server.py:365
Finding

Unbounded File Uploads and Missing Temporary-File Cleanup Enable Resource Exhaustion

Content
View full analysis
web.Response: reader = await request.multipart() field = await reader.next() if not field or not field.filename: return web.json_response({"error": "No file uploaded"}, status=400) suffix = Path(field.filename).suffix or ".wav" temp_path = os.path.join(UPLOAD_DIR, f"{uuid.uuid4().hex}{suffix}") with open(temp_path, "wb") as f: while True: chunk = await field.read_chunk(8192) if not chunk: break f.write(chunk) return web.json_response({ "file_path": temp_path, "file_name": field.filename, "file_size": os.path.getsize(temp_path), }) ``` ```python def removeTask(id: string) -> { clearPollingTimer(id); set((state) => ({ tasks: state.tasks.filter((t) => t.id !== id), selectedTaskId: state.selectedTaskId === id ? null : state.selectedTaskId, })); } ``` The frontend removal logic only removes client-side task state and does not request deletion of the corresponding uploaded backend file. ### Technical Analysis The upload handler streams the complete request body to disk without enforcing a maximum request size, per-file limit, aggregate quota, content validation, or upload rate limit. The application creates a process-specific temporary directory but does not remove uploaded originals after successful transcription, failure, cancellation, or frontend task removal. Because the API is unauthenticated and externally bound, the absence of storage and concurrency limits is directly exploitable by reachable clients. Randomized filenames prevent dir ...[truncated 1395 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
references/start.sh:39
Finding

Mutable Dependencies and Automatic Runtime Installation Create Supply-Chain Exposure

Content
View full analysis
/dev/null; then echo " ✓ modelscope installed" return 0 fi echo " [install] modelscope..." pip3 install modelscope -q 2>&1 | tail -3 if python3 -c "import modelscope" 2>/dev/null; then echo " ✓ modelscope installation succeeded" return 0 fi echo " ✗ modelscope installation failed" return 1 } ``` ```text aiohttp>=3.9.0 mlx-whisper>=0.4.3 mlx-lm>=0.31.0 soundfile>=0.12.0 numpy>=1.24.0 ``` ```json { "dependencies": { "@emotion/react": "^11.11.4", "@emotion/styled": "^11.11.5", "@mui/icons-material": "^5.15.20", "@mui/material": "^5.15.20", "react": "^18.3.1", "react-dom": "^18.3.1", "react-dropzone": "^14.2.3", "zustand": "^4.5.2" }, "devDependencies": { "@types/react": "^18.3.3", "@types/react-dom": "^18.3.0", "@vitejs/plugin-react": "^4.3.0", "autoprefixer": "^10.4.19", "postcss": "^8.4.38", "tailwindcss": "^3.4.4", "typescript": "^5.4.5", "vite": "^5.3.1" } } ``` ### Technical Analysis Python dependencies use minimum-version constraints with no upper bounds or hashes, while Node.js dependencies use compatible-version ranges. No reviewed lockfile fixes the complete dependency graph. Consequently, identical deployment commands can resolve different code over time. The startup script also installs the latest available `modelscope` package automatically if an import fails. Package installation executes package-controlled build and installation behavior under the privileges of the user starting the application. This converts a normal application l ...[truncated 1421 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (61)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

A second description-behavior mismatch is present: the skill markets itself as an offline local scaffolder while simultaneously directing runtime setup, model retrieval, and application execution. Misrepresenting execution scope is dangerous because users may approve actions they would otherwise scrutinize, especially where code generation, dependency trust, and local server exposure are involved.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

A second description-behavior mismatch is present: the skill markets itself as an offline local scaffolder while simultaneously directing runtime setup, model retrieval, and application execution. Misrepresenting execution scope is dangerous because users may approve actions they would otherwise scrutinize, especially where code generation, dependency trust, and local server exposure are involved.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 92)May include surrounding context.

md
1. 配置文件:`package.json`, `tsconfig.json`, `tsconfig.node.json`, `vite.config.ts`, `tailwind.config.js`, `postcss.config.js`, `index.html`, `.gitignore`, `LICENSE

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill instructs the agent to run shell commands and perform network-dependent actions such as package installation and model downloads, but it does not declare any tool scope or allowed-tools restrictions. This creates an authorization gap where a deployment-oriented skill can invoke powerful capabilities without explicit least-privilege boundaries, increasing the chance of unsafe execution or misuse.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The top-level description emphasizes local generation and 'zero network access for source code' but does not upfront warn that the process will still install dependencies and download multi-gigabyte models from the network. This can lead users to trigger significant external downloads and package execution without clear prior notice, which is a meaningful safety and trust issue.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The troubleshooting section recommends force-killing whatever is bound to ports 3000 and 8765 using kill -9, without warning that this may terminate unrelated user processes or cause data loss. In a generic local environment, those ports may be used by other applications, so the instruction can cause unintended disruption.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The file instructs the agent to create directories, write many files into TARGET_DIR, and change permissions on start.sh, but it does not require any confirmation, path validation, or warning that local files may be overwritten. In a code-generation/deployment skill, this creates a real risk of unintended filesystem modification, especially if TARGET_DIR is mis-specified or already contains user data.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The component defines user-facing labels such as “全部”, “转录中”, and “失败” directly in Chinese, and additional Chinese strings appear elsewhere in the UI. This forces a specific language/locale without any visible opt-in, language selection, or justification that the skill is region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

This component’s user-facing strings are consistently hard-coded in Chinese, including status messages, button labels, and tooltips. The file does not offer any user language selection or opt-in, which can violate a language/locale policy that requires user choice rather than forcing a specific language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This TypeScript/TSX file contains user-facing strings and descriptive comments in Chinese, including the upload and status labels, with no indication that language selection is optional or that the skill is intended only for a Chinese-speaking locale. That can violate the language/locale policy because the interface appears to require a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code presents user-facing labels, tooltips, and documentation entirely in Chinese, such as the component description and form labels, with no indication that the user can choose another language. The policy explicitly calls for flagging language or locale constraints when they are imposed without opt-in or documented justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The store initializes SRT translation defaults to source language 'ja' and target language 'zh', and elsewhere hard-codes translation to 'zh'. This imposes a specific language/locale behavior by default rather than offering a neutral default or explicit user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code automatically uploads the provided File object to a backend service via apiUploadFile, which transmits user data off the local UI. While comments describe the technical step, there is no visible confirmation prompt or user-facing disclosure in this file before the transfer occurs.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The translateTask function always calls apiTranslate with target language 'zh', regardless of user preference. This is a natural-language policy concern because it forces a specific locale outcome instead of allowing user choice or documenting a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The translateSrt flow sends subtitle segment text to apiTranslate in batches, which may expose potentially sensitive transcript content to the backend service. This file contains no visible user disclosure or confirmation that subtitle contents will be transmitted for translation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This TypeScript file performs multiple network operations that send potentially sensitive user data to a backend, including file paths, transcript segments, and uploaded audio content. Although the functions have developer-facing comments, there is no visible user-facing warning, confirmation, or notice in this file about data being transmitted off the client.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Natural-language strings throughout the script force a specific language for user-facing interaction. The policy allows language constraints only when the skill offers opt-in or clearly documents a justified locale limitation, which is not present here.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · references/transcribe_server.py (reported line 42)May include surrounding context.

python
# 用 macOS 内置 afconvert 转换
    tmp_wav = tempfile.mktemp(suffix=".wav")
    try:
        subprocess.run(
            ["afconvert", "-f", "WAVE", "-d", "LEI16", audio_path, tmp_wav],
            check=True, capture_output=True, text=True,
        )

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The system prompt is hard-coded in Chinese, instructing the model behavior in a single language with no indication of user opt-in or locale selection. This can violate language or locale policy when the skill is expected to accommodate user choice rather than imposing one language by default.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This handler persists uploaded content to a temporary directory on the local filesystem, which is a safety-relevant file write operation. Although the code comments describe saving the file, there is no user-facing confirmation, warning, or disclosure to inform users that uploaded audio will be stored locally.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The /api/transcribe endpoint accepts any server-local file_path and only checks os.path.isfile, allowing a caller to make the service open arbitrary local files. In this skill's context, the backend is also exposed on 0.0.0.0 and permits cross-origin requests, which makes this much more dangerous because remote users or browser-based attackers can drive access to unintended server files and trigger parsing/conversion on them.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Accepting arbitrary local file paths from clients without clear warning is not just a disclosure issue here; it is a real security flaw because it enables server-side file access beyond uploaded content. In combination with remote exposure, attackers can coerce the backend into touching sensitive or unintended files on the host.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The server binds to 0.0.0.0 and the CORS middleware allows any origin, so any host that can reach the service can invoke upload, transcription, task, and translation APIs, including from malicious web pages. For a tool described as a local setup utility, this unnecessarily expands attack surface and compounds the arbitrary-file-path issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The operational instructions are presented entirely in Chinese, which can impose a language requirement on users without opt-in. The file does not state that the skill is intentionally region-specific or offer an alternative language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

Most of the instructional content is written in Chinese, which imposes a specific language on users reading or executing the skill guidance. There is no indication that Chinese is optional, user-selected, or required for a region-specific purpose.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.