Back to skill

Security audit

Audiolla

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent Audiolla client, but its setup docs can expose an unauthenticated audio-processing service and use mutable Docker images.

Review the setup before installing. Bind Audiolla to 127.0.0.1 unless you intentionally need remote access, set a strong AUDIOLLA_AUTH_TOKEN before exposing it, restrict remote URL fetch/upload to an allowlist, and prefer pinned image digests over latest tags.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
references/setup.md:15
Finding

Default Quick-Install Command Exposes an Unauthenticated Service on All Network Interfaces

Content
View full analysis
Remediation
View remediation
``` 2. Require a strong authentication token during initial installation rather than presenting it as an optional follow-up hardening step. 3. Refuse to start without authentication when the server is configured to listen on a non-loopback address, unless an explicit unsafe-development override is supplied. 4. Place any remotely accessible deployment behind a TLS-terminating reverse proxy. 5. Add per-client rate limits, request quotas, upload quotas, and enforced asynchronous-job concurrency limits. 6. Apply authentication and authorization consistently to file listing, download, deletion, job control, and processing endpoints. 7. Use network firewall rules or a private overlay network such as WireGuard or Tailscale to restrict reachability. 8. Monitor disk, CPU, GPU, and job-queue usage, and automatically reject work when configured resource thresholds are reached. ]]>

T08 · Insecure Dependencies

Warning
Location
references/setup.md:15
Finding

Quick-Install Commands Execute Mutable, Unverified Container Images

Content
View full analysis
Remediation
View remediation
@sha256: ``` 2. Publish and verify container signatures using a mechanism such as Sigstore Cosign. 3. Require verification of build provenance and software bill of materials metadata before deployment. 4. Document a controlled upgrade process in which operators explicitly review and replace the pinned digest. 5. Use automated dependency and image vulnerability scanning, but do not automatically deploy a newly moved `latest` tag. 6. Run the container as a non-root user with a read-only root filesystem where supported. 7. Restrict outbound network access and mount only the directories required for Audiolla operation. 8. Avoid exposing unnecessary devices; grant GPU access only to deployments that require CUDA functionality. ]]>
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (16)

Cloud Metadata Access

High
Category
Server-Side Request Forgery
Confidence
90% confidence
Finding

Code accesses a cloud instance metadata endpoint (e.g. 169.254.169.254). A single request can return temporary IAM credentials, making this a high-value SSRF target for credential theft.

Content

Scanner excerpt · SKILL.md (reported line 762)May include surrounding context.

md
Host patterns are exact match (`bucket.s3.amazonaws.com`) or single-wildcard subdomain (`*.s3.amazonaws.com`, matches any `<x>.s3.amazonaws.com` but NOT `s3.amazonaws.com` itself).

Always-on protections regardless of mode:
- DNS-resolved private / loopback / link-local / metadata-service IPs (`169.254.169.254`) rejected unless `AUDIOLLA_FETCH_ALLOW_PRIVATE=true`
- Only schemes in `AUDIOLLA_FETCH_SCHEMES` accepted; `file://`, `gopher://`, etc. always rejected
- Each redirect's `Location` re-validated through the full policy before following
- Body streamed; abort if it exceeds `AUDIOLLA_MAX_UPLOAD_BYTES`

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Content

Scanner excerpt · references/setup.md (reported line 84)May include surrounding context.

md
### Remote URL fetching (file_url / output_url)

Audiolla can fetch input files from a URL and PUT outputs to presigned URLs (S3, R2, etc.). This is **disabled by default** because the fetch path is a classic SSRF surface — without guardrails, an attacker can use it to read your cloud metadata service or probe internal hosts.

Pick a mode that matches your setup:

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/setup.md (reported line 196)May include surrounding context.

md
- ./data:/data
    environment:
      AUDIOLLA_DEVICE: auto
      AUDIOLLA_AUTH_TOKEN: ${AUDIOLLA_AUTH_TOKEN}     # from .env
      AUDIOLLA_ENGINE_TTL: 10m
      AUDIOLLA_MAX_UPLOAD_BYTES: 209715200
    # For CUDA image only:

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 167)May include surrounding context.

bash
# Liveness — no auth required
curl $AUDIOLLA_URL/healthz
# {"ok": true, "device": "cpu", "engines": ["htdemucs", "matchering", ...]}

# Configured engines + capabilities

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 491)May include surrounding context.

bash
# Measure
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
  -H 'Content-Type: application/json' \
  $AUDIOLLA_URL/v1/audio/loudness \
  -d '{"file_path":"uploads/track.wav"}'

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 588)May include surrounding context.

Read the structure of any Standard MIDI File. Analysis-only, returns JSON.

bash
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
  -H 'Content-Type: application/json' \
  $AUDIOLLA_URL/v1/midi/inspect \
  -d '{"file_path":"midi/song.mid"}'

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 661)May include surrounding context.

bash
# Render a staged MIDI to staged audio
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
  -H 'Content-Type: application/json' \
  $AUDIOLLA_URL/v1/midi/render \
  -d '{"file_path":"midi/song.mid","output_format":"wav","output_path":"audio/song.wav"}'

External Transmission

Medium
Category
Data Exfiltration
Confidence
84% confidence
Finding

This example shows server-side fetching from a remote file_url and uploading results to an output_url, which introduces external network transmission and potential data egress. Although the document notes fetch mode controls, enabling this feature can expose user media or derived outputs to third-party endpoints and expands the attack surface to SSRF-style misuse if server policy is weak or misconfigured.

Content

Scanner excerpt · SKILL.md (reported line 772)May include surrounding context.

Example — fetch from S3, master, PUT to a presigned URL:

bash
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
  -H 'Content-Type: application/json' \
  $AUDIOLLA_URL/v1/audio/master \
  -d '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 825)May include surrounding context.

md
Audio over MCP is **path/URL-based for inputs and for most audio-producing outputs** — JSON-RPC can't carry raw bytes efficiently, so the v1.0.0 contract enforces this on nearly every tool. For inputs: pre-stage via REST `PUT /v1/files/{path}` (or the `put_file` MCP tool for small files) and pass `file_path`, or pass `file_url` when fetch is enabled. For outputs: pick `output_path` to keep the result in staging (chain it as the next call's `file_path`), or `output_url` to PUT it to S3-style storage. Exceptions: `beats` (click track), `melody` (as_midi), and `silence` (trim) still return their secondary audio/MIDI artifact as inline base64 rather than staging it — see the table above.

The MCP endpoint is at `$AUDIOLLA_URL/v1/mcp`. It is JSON-RPC over streamable HTTP; do not try to describe it in OpenAPI or hit it with raw curl — use an MCP client.

### AI restoration — de-reverb / de-echo / de-noise (`/v1/audio/restore/{engine}`)

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 919)May include surrounding context.

md
$AUDIOLLA_URL/v1/presets/master-for-spotify | jq '.steps'

# Run a preset against a staged file
curl -X POST -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
  -H 'Content-Type: application/json' \
  $AUDIOLLA_URL/v1/presets/podcast-cleanup \
  -d '{"file_path":"uploads/interview.wav","output_path":"out/cleaned.wav"}'

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 981)May include surrounding context.

md
| jq '{status, duration_sec, result}'

# List with optional filter
curl -H "Authorization: Bearer $AUDIOLLA_TOKEN" \
  "$AUDIOLLA_URL/v1/jobs?status=completed"

# Cancel

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.