T09 · Insecure Skill Coding Practices
- Location
scripts/server.py:117- Finding
Unauthenticated Public Webhooks Permit API Resource Abuse and Prompt Injection
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This phone-calling skill is mostly coherent, but it exposes unauthenticated public call webhooks and public audio hosting, so it needs Review before installation.
Install only if you are comfortable with a skill that can place real phone calls, use paid Twilio/OpenAI/ElevenLabs resources, expose a local webhook server to the internet, upload one-way call audio to tmpfiles.org, and forward call summaries by iMessage. Before using it, add Twilio signature validation, restrict known call SIDs, require explicit per-call confirmation, pin or replace localtunnel, avoid public audio hosting, and define cleanup for /tmp audio and transcript files.
scripts/server.py:117Unauthenticated Public Webhooks Permit API Resource Abuse and Prompt Injection
scripts/one_way_call.py:30Call Audio Is Uploaded to a Public Third-Party File Host
SKILL.md:39Unpinned Package Is Downloaded and Executed Through npx
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
def generate_audio(text: str, voice_id: str) -> bytes:
r = requests.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
headers={"xi-api-key": ELEVENLABS_API_KEY, "Content-Type": "application/json"},
json={
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
def tts(text: str, voice_id: str = DEFAULT_VOICE) -> str:
"""Generate audio via ElevenLabs, serve it locally, return URL path."""
r = requests.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
headers={"xi-api-key": ELEVENLABS_API_KEY, "Content-Type": "application/json"},
json={
This mismatch includes undeclared iMessage notification behavior, local transcript/summary storage, and incomplete implementation relative to the stated purpose. Hidden notification channels and undocumented persistence expand data exposure and can leak call content or metadata beyond what a user expects from a phone-calling skill.
This mismatch includes undeclared iMessage notification behavior, local transcript/summary storage, and incomplete implementation relative to the stated purpose. Hidden notification channels and undocumented persistence expand data exposure and can leak call content or metadata beyond what a user expects from a phone-calling skill.
The trigger phrases are broad enough to match common conversation such as 'call', 'phone', or 'reserve', increasing the chance that the skill activates in situations where the user did not intend to place an outbound call. In a calling skill, accidental activation is more dangerous than usual because it can initiate real-world actions, contact third parties, and transmit content to Twilio, ElevenLabs, and possibly other services.
Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.
# Search shared library by accent + gender
curl "https://api.elevenlabs.io/v1/shared-voices?accent=british&gender=female&page_size=10" \
-H "xi-api-key: $ELEVENLABS_API_KEY" | python3 -m json.tool
The script reads generated or user-supplied audio from disk and uploads it to tmpfiles.org, an unrelated third-party file-hosting service. This can expose sensitive call content, creates an unnecessary public distribution point, and expands access beyond Twilio/ElevenLabs without user consent or access controls.
with tempfile.NamedTemporaryFile(suffix=".mp3", delete=False) as f:
f.write(audio_bytes)
tmp_path = f.name
resp = requests.post("https://tmpfiles.org/api/v1/upload", files={"file": open(tmp_path, "rb")})
resp.raise_for_status()
url = resp.json()["data"]["url"].replace("tmpfiles.org/", "tmpfiles.org/dl/")
return url
The skill documentation describes use of environment variables, shell commands, network access, and file writes, but the manifest does not declare any tool scope or permission boundaries. In an agent environment, this increases the risk of over-privileged execution and makes it harder for users or platforms to understand that the skill can place calls, upload content, and persist data.
The skill can automatically place outbound calls and send user-provided content to third-party providers, yet the documentation lacks a clear warning and consent model. Given the context, this is especially risky because voice content, phone numbers, and potentially transcripts may be processed by multiple vendors and could create privacy, harassment, or legal-consent issues.
Automatic iMessage summaries send call-derived content to an additional communication channel outside the skill's stated outbound-calling scope. In this context, call summaries may contain sensitive personal or business information, so undisclosed forwarding materially increases privacy and compliance risk.
Using npx localtunnel without a pinned version pulls and executes the latest package at runtime, creating a supply-chain risk if the package is compromised or behavior changes unexpectedly. Because this tunnel exposes a locally running server to the internet, compromise could directly affect call handling and any secrets or endpoints reachable through that server.
The cron example establishes persistent scheduled execution, which can continue placing calls after the original session ends. In this context, persistence is security-relevant because a misconfigured or forgotten schedule could repeatedly contact recipients, generate charges, and continue processing call content without fresh user review.
Use macOS cron for timed calls:
# Add to crontab — this example calls at 8:45 AM
crontab -e
45 8 24 2 * python3 /path/to/scripts/one_way_call.py --to "+1..." --text "Good morning!" >> /tmp/call.log 2>&1
The document centers recommendations for phone calls around specific accents and regional voice defaults, including marking Australian voices as the default for most calls. This can create a language/locale policy issue because it steers output toward a particular locale without stating that users can choose another locale or that the tool is region-specific.
This script directly places a real outbound phone call as soon as it is run, with no explicit confirmation, dry-run mode, recipient validation, or user-facing warning about the external action. In an agent-skill context, that creates a meaningful safety risk because a prompt, automation, or mistaken invocation can trigger unsolicited calls, spam, harassment, or unintended charges in the real world.
User-provided text is sent to ElevenLabs for text-to-speech generation, which is expected functionality but still an external data disclosure. In this skill context that is somewhat expected, yet it remains a privacy issue if users are not clearly informed that their message content leaves the local environment and is processed by ElevenLabs.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
def generate_audio(text: str, voice_id: str) -> bytes:
r = requests.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
headers={"xi-api-key": ELEVENLABS_API_KEY, "Content-Type": "application/json"},
json={
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
def generate_audio(text: str, voice_id: str) -> bytes:
r = requests.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
headers={"xi-api-key": ELEVENLABS_API_KEY, "Content-Type": "application/json"},
json={
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
def generate_audio(text: str, voice_id: str) -> bytes:
r = requests.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
headers={"xi-api-key": ELEVENLABS_API_KEY, "Content-Type": "application/json"},
json={
"text": text,
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
def generate_audio(text: str, voice_id: str) -> bytes:
r = requests.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
headers={"xi-api-key": ELEVENLABS_API_KEY, "Content-Type": "application/json"},
json={
"text": text,
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
def generate_audio(text: str, voice_id: str) -> bytes:
r = requests.post(
f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
headers={"xi-api-key": ELEVENLABS_API_KEY, "Content-Type": "application/json"},
json={
"text": text,
The code sends call audio to tmpfiles.org even though the skill description only mentions Twilio and ElevenLabs. That hidden dependency materially changes the privacy and threat model because a separate service gains access to potentially sensitive voice content and serves it from a publicly retrievable URL.
The script transmits audio to an external file host without clearly warning the user that the content will be temporarily publicly hosted. This undermines informed consent and can cause accidental disclosure of sensitive message content, especially when users expect only Twilio and ElevenLabs to process the data.
Publishing call audio to a public temporary file-sharing service is risky because anyone with the URL may access the content during its lifetime, and the hosting provider can inspect or retain it. In a calling skill, messages may include personal, business, or regulated information, making the context more dangerous than generic file transfer.
This external transmission sends audio content to tmpfiles.org, which is not necessary to the core user expectation of a Twilio/ElevenLabs calling workflow and introduces a separate exposure channel. Because the content is call media, the transmission can disclose private or sensitive information to an unrelated service and to anyone obtaining the resulting URL.
with tempfile.NamedTemporaryFile(suffix=".mp3", delete=False) as f:
f.write(audio_bytes)
tmp_path = f.name
resp = requests.post("https://tmpfiles.org/api/v1/upload", files={"file": open(tmp_path, "rb")})
resp.raise_for_status()
url = resp.json()["data"]["url"].replace("tmpfiles.org/", "tmpfiles.org/dl/")
return url
The server automatically summarizes call transcripts and forwards the summary to a configured local iMessage recipient, which extends beyond core call handling and creates an unannounced data-exfiltration path. In this skill context, phone conversations may contain sensitive personal or business information, making post-call forwarding meaningfully risky.
No suspicious patterns detected.