Back to skill

Security audit

Podcast Summarizer

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its podcast-summary purpose, but it can fetch arbitrary remote or RSS-provided URLs and send transcript text to external LLM providers without strong safeguards.

Review this skill before installing. Use it only with trusted podcast URLs or feeds, preferably in an isolated environment with network egress controls and resource limits. Do not process private or sensitive audio unless you are comfortable sending transcript text to Gemini or OpenAI when API keys are configured.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/summarize_podcast.py:53
Finding

Unrestricted URL Fetching Enables Server-Side Request Forgery

Content
View full analysis
list[dict]: """Parse RSS feed and return episodes.""" resp = requests.get(feed_url, timeout=30) resp.raise_for_status() root = ET.fromstring(resp.content) channel = root.find("channel") episodes = [] for item in channel.findall("item"): title = item.findtext("title", "") description = item.findtext("description", "") # Find audio URL enclosure = item.find("enclosure") audio_url = enclosure.get("url") if enclosure is not None else None # Try media:content as fallback if not audio_url: for child in item: if "content" in child.tag.lower() and child.get("url"): audio_url = child.get("url") break ``` ```python def download_audio(url: str, output_path: str) -> bool: """Download audio file.""" print(f"Downloading audio from {url[:80]}...", file=sys.stderr) try: resp = requests.get(url, timeout=300, stream=True) resp.raise_for_status() with open(output_path, "wb") as f: for chunk in resp.iter_content(chunk_size=8192): f.write(chunk) return True ``` ```python elif url.endswith(".xml") or "feed" in url.lower() or "rss" in url.lower(): print("Detected RSS feed", file=sys.stderr) episodes = parse_rss_feed(url) if episodes: episode_info = episodes[args.episode - 1] print(f"Selected episode: {episode_info['title']}", file=sys.stderr) elif url.endswith(".mp3") or url.endswith(".m4a"): pri ...[truncated 3070 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/summarize_podcast.py:82
Finding

Unbounded Remote Content Processing Enables Resource Exhaustion

Content
View full analysis
list[dict]: """Parse RSS feed and return episodes.""" resp = requests.get(feed_url, timeout=30) resp.raise_for_status() root = ET.fromstring(resp.content) ``` ```python def download_audio(url: str, output_path: str) -> bool: """Download audio file.""" print(f"Downloading audio from {url[:80]}...", file=sys.stderr) try: resp = requests.get(url, timeout=300, stream=True) resp.raise_for_status() with open(output_path, "wb") as f: for chunk in resp.iter_content(chunk_size=8192): f.write(chunk) return True ``` ### Technical Analysis RSS responses are loaded in full through `resp.content` before XML parsing. No upper bound is imposed on the response body, allowing an attacker-controlled endpoint to force substantial memory consumption. Audio is streamed to a temporary file, which avoids buffering the entire file in memory but does not limit the total bytes written. The loop continues until the server terminates the response. The configured timeout does not provide a reliable response-size limit and may not impose a strict total transfer deadline as long as data continues to arrive. The implementation also does not validate `Content-Length`, media type, expected audio format, or actual downloaded size before invoking Whisper. A very large or malformed file can therefore consume temporary storage and subsequently impose excessive CPU, memory, and processing time on the transcription stage. ### Attack Path 1. The attacker supplies an RSS URL or direct audio URL under their control. 2. For RSS exploitation, the server returns a very large XML respons ...[truncated 1285 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (8)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The documented purpose understates the actual behavior by omitting external LLM API usage, undeclared API key dependence, broader network access patterns, and fallback behaviors that do not match 'transcript and summary' expectations. This mismatch is dangerous because users may authorize or run the skill under false assumptions, exposing secrets, sending data to third parties, or receiving misleading outputs.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill documentation describes behavior that requires network access, shell execution, environment variable access, and likely file output, but it declares no explicit tool scope or permissions. This weakens security boundaries because a user or runtime may not realize the skill can fetch arbitrary remote content, invoke local tools like whisper, and write results to disk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The documentation does not clearly warn that the skill downloads remote audio from user-supplied URLs and may save transcripts or summaries to local files. This can lead to unsafe use, including processing untrusted remote content, unexpected bandwidth or storage consumption, and accidental persistence of sensitive content on disk.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/summarize_podcast.py (reported line 108)May include surrounding context.

python
cmd.extend(["--language", language])
    
    try:
        result = subprocess.run(cmd, capture_output=True, text=True, timeout=1800)
        
        # Whisper outputs to {input_name}.txt
        txt_path = audio_path.rsplit(".", 1)[0] + ".txt"

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Full transcript text is sent to external LLM APIs without an explicit warning or consent step. Even though summarization inherently may require model processing, transcripts can contain personal, confidential, or copyrighted material, so silent exfiltration to external providers creates a real privacy and compliance risk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill reads API keys from environment variables and transmits podcast transcripts to third-party LLM providers. That data flow is part of summarization, but the code does so implicitly and without clear consent controls, which can expose sensitive transcript content or internal deployment credentials to external services in ways users may not expect.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest description says the skill summarizes episodes from Spotify, Apple Podcasts, or RSS feeds, and the CLI help repeats that Apple URLs are supported. However, the Apple Podcasts branch only prints detection and contains a TODO without implementing retrieval or transcription, so the documentation actively overstates actual behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The file switches into Traditional Chinese for the SoundOn section and presents operational guidance in that language only. Under the language/locale policy, forcing a specific language without user opt-in can be a policy violation unless the locale restriction is clearly documented and justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.