Back to skill

Security audit

local-voice-reply

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real local voice-reply server, but it needs Review because it stores voice/text/audio data and exposes unauthenticated, resource-heavy local API endpoints.

Install only if you are comfortable running a local unauthenticated TTS API for trusted local use. Keep it bound to 127.0.0.1, do not expose the port through a proxy or LAN interface, avoid sending sensitive text or private voice samples unless you accept local retention, and plan to manually manage or clear the voice, output, sidecar, and phrase-cache directories. For production or shared machines, require authentication, input limits, retention limits, and pinned dependencies first.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
server/voice_server_v3.py:126
Finding

Unauthenticated and Unbounded Upload and Synthesis Endpoints Enable Resource Exhaustion

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
server/voice_engine.py:861
Finding

Synthesized User Text and Audio Are Persisted Without a Retention Boundary

Content
View full analysis
None: wav_cpu = wav.detach().to("cpu", dtype=torch.float32).contiguous() self.phrase_ram_cache.put(key, wav_cpu) try: self._atomic_save_wav(self._phrase_cache_path(key), wav_cpu, sample_rate) except Exception: self.log.exception("phrase_cache_save_failed key=%s", key[:12]) ``` ### Technical Analysis The `/speak_stream` implementation stores the submitted text in the `chunks` field of a sidecar JSON file next to the generat ...[truncated 2111 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
server/VOICE_SERVER_V3.md:15
Finding

Unpinned Runtime Dependencies and Model Assets Create Supply-Chain Risk

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (11)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 25)May include surrounding context.

md
- `server/voice_engine.py` (generation and cache engine)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill describes capabilities that involve shell execution, file reads/writes, environment-variable use, and running a local FastAPI server, but it does not declare any explicit tool scope or allowed-tools restrictions. That creates an authorization gap where an agent may invoke broader tooling than necessary, increasing the chance of unintended command execution, filesystem modification, or access to sensitive local resources.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The activation condition includes 'any user request for voice/audio reply,' which is broad enough to match ordinary conversation and may auto-trigger the skill unexpectedly. Because this skill can start or rely on local server processes, write files, and send outbound media, overbroad activation increases the risk of unintended tool use and side effects from benign user messages.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The register_voice flow persists user-supplied audio to the voices directory via write_bytes, but there is no confirmation prompt, user-facing print/log disclosure, or inline comment/docstring warning that uploaded voice samples will be stored on disk. Because this is a file write involving user data, the code should visibly disclose the persistence behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The synthesize method writes an .opus output file and a .json sidecar containing metadata such as trace_id, voice_name, device, and latency metrics, but the code does not include a user-facing warning or explanatory comment/docstring about this persistence. This is a safety-relevant file write because it stores generated content and associated metadata on the local system.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The synthesize_stream method saves a streamed .opus file and a JSON sidecar that includes chunk contents and processing metadata, yet there is no confirmation prompt, user-facing notice, or explanatory documentation in the code. Since this persists potentially sensitive text-derived data and metadata, the storage behavior should be clearly disclosed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The cleanup_output method unlinks output and sidecar files, which is a destructive file deletion operation, but there is no confirmation step or explicit warning in comments/docstrings that the action is irreversible. For destructive operations, the code should include some visible disclosure unless the destructive nature is clearly documented as core behavior.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The /health endpoint discloses internal operational details including registered voice names, cache keys, output directory paths, implementation method names, and benchmark data. In an exposed service, this information materially helps reconnaissance by revealing system internals, available assets, and performance characteristics that can be used to target follow-on abuse or theft of tenant-specific metadata.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The /output/cleanup endpoint performs deletion-related cleanup based on a user-supplied path with no visible authentication, confirmation, or safety constraints in this file. In the context of a local voice-reply API that writes output files, such an endpoint can be abused to delete generated artifacts or, depending on engine.cleanup_output() behavior, potentially remove unintended files if path validation is weak.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The documentation states that first startup may download Chatterbox model assets via from_pretrained(), but the skill metadata/description does not warn users about this network behavior. Undisclosed outbound network access can violate operator expectations, break offline-only assumptions, and create supply-chain exposure if users enable the skill in restricted environments.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · server/voice_server_v3.py (reported line 27)May include surrounding context.

python
def _configure_logging() -> None:
    level_name = os.getenv("TARVIS_VOICE_LOG_LEVEL", "INFO").upper()
    level = getattr(logging, level_name, logging.INFO)
    logging.basicConfig(
        level=level,
        format="%(asctime)s %(levelname)s %(name)s %(message)s",

Static analysis

No suspicious patterns detected.