T09 · Insecure Skill Coding Practices
- Location
server/voice_server_v3.py:126- Finding
Unauthenticated and Unbounded Upload and Synthesis Endpoints Enable Resource Exhaustion
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This appears to be a real local voice-reply server, but it needs Review because it stores voice/text/audio data and exposes unauthenticated, resource-heavy local API endpoints.
Install only if you are comfortable running a local unauthenticated TTS API for trusted local use. Keep it bound to 127.0.0.1, do not expose the port through a proxy or LAN interface, avoid sending sensitive text or private voice samples unless you accept local retention, and plan to manually manage or clear the voice, output, sidecar, and phrase-cache directories. For production or shared machines, require authentication, input limits, retention limits, and pinned dependencies first.
server/voice_server_v3.py:126Unauthenticated and Unbounded Upload and Synthesis Endpoints Enable Resource Exhaustion
server/voice_engine.py:861Synthesized User Text and Audio Are Persisted Without a Retention Boundary
server/VOICE_SERVER_V3.md:15Unpinned Runtime Dependencies and Model Assets Create Supply-Chain Risk
Referenced artifact was not completely inspected
- `server/voice_engine.py` (generation and cache engine)
The skill describes capabilities that involve shell execution, file reads/writes, environment-variable use, and running a local FastAPI server, but it does not declare any explicit tool scope or allowed-tools restrictions. That creates an authorization gap where an agent may invoke broader tooling than necessary, increasing the chance of unintended command execution, filesystem modification, or access to sensitive local resources.
The activation condition includes 'any user request for voice/audio reply,' which is broad enough to match ordinary conversation and may auto-trigger the skill unexpectedly. Because this skill can start or rely on local server processes, write files, and send outbound media, overbroad activation increases the risk of unintended tool use and side effects from benign user messages.
The register_voice flow persists user-supplied audio to the voices directory via write_bytes, but there is no confirmation prompt, user-facing print/log disclosure, or inline comment/docstring warning that uploaded voice samples will be stored on disk. Because this is a file write involving user data, the code should visibly disclose the persistence behavior.
The synthesize method writes an .opus output file and a .json sidecar containing metadata such as trace_id, voice_name, device, and latency metrics, but the code does not include a user-facing warning or explanatory comment/docstring about this persistence. This is a safety-relevant file write because it stores generated content and associated metadata on the local system.
The synthesize_stream method saves a streamed .opus file and a JSON sidecar that includes chunk contents and processing metadata, yet there is no confirmation prompt, user-facing notice, or explanatory documentation in the code. Since this persists potentially sensitive text-derived data and metadata, the storage behavior should be clearly disclosed.
The cleanup_output method unlinks output and sidecar files, which is a destructive file deletion operation, but there is no confirmation step or explicit warning in comments/docstrings that the action is irreversible. For destructive operations, the code should include some visible disclosure unless the destructive nature is clearly documented as core behavior.
The /health endpoint discloses internal operational details including registered voice names, cache keys, output directory paths, implementation method names, and benchmark data. In an exposed service, this information materially helps reconnaissance by revealing system internals, available assets, and performance characteristics that can be used to target follow-on abuse or theft of tenant-specific metadata.
The /output/cleanup endpoint performs deletion-related cleanup based on a user-supplied path with no visible authentication, confirmation, or safety constraints in this file. In the context of a local voice-reply API that writes output files, such an endpoint can be abused to delete generated artifacts or, depending on engine.cleanup_output() behavior, potentially remove unintended files if path validation is weak.
The documentation states that first startup may download Chatterbox model assets via from_pretrained(), but the skill metadata/description does not warn users about this network behavior. Undisclosed outbound network access can violate operator expectations, break offline-only assumptions, and create supply-chain exposure if users enable the skill in restricted environments.
Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.
def _configure_logging() -> None:
level_name = os.getenv("TARVIS_VOICE_LOG_LEVEL", "INFO").upper()
level = getattr(logging, level_name, logging.INFO)
logging.basicConfig(
level=level,
format="%(asctime)s %(levelname)s %(name)s %(message)s",
No suspicious patterns detected.