other
- Location
scripts/pipeline.py:38- Finding
Undisclosed transmission of voice recordings and children's data to third-party services
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The package is labeled as a generic local pipeline skill, but the inspected code is a provider-specific story app that sends child-related audio and text to external APIs and exposes stored story data without visible authorization.
Review before installing or deploying. This skill should be treated as a child-story web service, not a generic local pipeline helper. It needs accurate disclosure of ElevenLabs and Mistral use, explicit consent and data-minimization controls for voice and child-related content, authentication and per-record authorization on all routes, upload limits, and corrected metadata for required API keys and outbound network access.
scripts/pipeline.py:38Undisclosed transmission of voice recordings and children's data to third-party services
scripts/pipeline.py:116User-controlled child name is interpolated into a privileged LLM system prompt
scripts/pipeline.py:175Story, audio, and image records are exposed without route-level authorization
scripts/pipeline.py:35Unbounded voice upload buffering permits memory and API-resource exhaustion
A strong description-behavior mismatch is security-relevant because users and reviewers may grant trust, approvals, or tool access based on the declared purpose while the underlying implementation performs materially different actions such as provider-specific API calls, persistence, and asset serving. That gap can conceal unexpected data handling and network behavior, increasing the risk of overscoped deployment and inadequate review.
Without declared permissions the skill's intent is opaque and cannot be validated.
Overly broad invocation language can cause the skill to be selected for many unrelated tasks, which raises the chance that powerful or mismatched functionality is invoked in contexts the user did not intend. In combination with the reported description-behavior mismatch, broad matching becomes more dangerous because it can funnel arbitrary requests into a skill with hidden or under-disclosed behaviors.
The manifest describes a reusable, provider-agnostic multi-agent pipeline with sequential/parallel stages, status tracking, error recovery, and progress callbacks. This file instead implements a specific bedtime-story application with hard-coded story generation, STT, narration, cached audio/image serving, and story CRUD endpoints tied to ElevenLabs, Mistral, and a stories database.
The manifest says the skill 'Works with any LLM provider,' but this file directly depends on ElevenLabs for speech-to-text and text-to-speech and on Mistral for transcription and story generation. There is no provider abstraction or pluggable backend mechanism in this implementation, so the advertised generality does not match the code.
The described skill emphasizes generic orchestration features such as stage status tracking, error recovery, and progress callbacks. In this file, endpoints perform discrete actions for transcription, story generation, narration, and cache retrieval, but there is no visible generic stage graph, callback mechanism, or workflow status model; error handling is limited to local try/except and HTTP exceptions.
Uploaded audio is forwarded to external STT providers, which can expose sensitive voice content and metadata to third parties if users are not clearly informed and consent is not captured. In a child-story workflow, this is more sensitive because recordings may include children's voices and personal content.
This endpoint transmits uploaded audio to ElevenLabs over the network, which is an external data exfiltration path by design. Even though it is expected functionality, it is security-relevant because sensitive audio leaves the trust boundary and may include children's voices or private speech.
# ElevenLabs STT (Scribe v1)
try:
async with httpx.AsyncClient(timeout=30) as client:
r = await client.post("https://api.elevenlabs.io/v1/speech-to-text",
headers={"xi-api-key": ELEVENLABS_API_KEY},
files={"file": (audio.filename or "audio.wav", audio_data, audio.content_type or "audio/wav")},
data={"model_id": "scribe_v1"})
The application sends prompts and a child identifier to an external LLM, creating privacy risk because personally identifying or sensitive child-related content may be processed by a third party. The child-focused context increases sensitivity and makes silent transmission more problematic.
Narration text is sent to a third-party TTS service without any explicit disclosure in this code, which can leak story contents and personal details embedded in the text. Because the workflow targets children's stories, transmitted content may include sensitive child-related information.
This code sends narration text to ElevenLabs TTS, creating an outbound transfer of potentially sensitive content to a third-party service. In this child-story context, the transferred text may contain names or personal details, raising privacy and compliance concerns.
async with httpx.AsyncClient(timeout=30) as client:
r = await client.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{req.voice_id}",
headers={"xi-api-key": ELEVENLABS_API_KEY, "Content-Type": "application/json"},
json={
"text": req.text,
The file reads ELEVENLABS_API_KEY and MISTRAL_API_KEY from the environment, but there is no nearby comment or docstring explaining that the skill depends on external credentials and uses them to contact third-party services. This is relevant because credential-backed external processing may affect deployment safety and privacy expectations.
No suspicious patterns detected.