T09 · Insecure Skill Coding Practices
- Location
scripts/read-url.js:29- Finding
Article Text and API Credentials Transmitted over Plaintext HTTP
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
CastReader’s core goal is understandable, but it handles arbitrary web pages and extracted text in ways that create review-level security and privacy risks.
Install only if you are comfortable with page text and generated summaries being sent to a remote TTS service and audio being sent through Telegram. Avoid private, internal, authenticated, paywalled, business, chat, document, and ebook URLs unless the skill is fixed to use HTTPS, validate URLs, keep Chromium sandboxing enabled, use safe temp files, and avoid shell interpolation for summaries.
scripts/read-url.js:29Article Text and API Credentials Transmitted over Plaintext HTTP
scripts/extract.js:21Unrestricted URL Navigation Permits Server-Side Request Forgery
scripts/extract.js:29Chromium Sandbox Disabled While Rendering Untrusted Pages
SKILL.md:68Generated Summary Text Is Interpolated into a Shell Command
scripts/read-url.js:35Predictable Shared Temporary Paths Store Extracted Content without Explicit Protection
package-lock.json:875Session Setup Executes an Unsupported Dependency's Installation Script
The declared description promises an end-to-end 'read any web page aloud' skill that extracts article text from any URL and converts it to audio (MP3). However, the supplied code is focused on extracting readable text from the current document: scoring content candidates, filtering navigation/comments/ads, handling shadow DOM, splitting paragraphs, refining highlightable elements, and applying site-specific extraction rules. There is also a WeChat extractor and embedded Mozilla Readability logic. These behaviors support article extraction, but the key promised capabilities—voice synthesis, audio playback, MP3 creation, and likely URL retrieval—are absent from this code chunk. Because the actual code's primary visible purpose is extraction rather than reading aloud or audio generation, the description materially overstates what this code does.
The declared description promises an end-user capability centered on turning any web page URL into spoken audio/MP3. The code shown instead focuses on extracting readable text/content from the current document, parsing metadata, caching extraction results, and handling special cases such as WeRead canvas-based layouts and highlight overlays. While article extraction is one supporting part of a read-aloud tool, the core promised behavior—AI voices, audio playback/synthesis, and MP3 generation—is absent from this code chunk. Additionally, the chunk includes substantial functionality unrelated to the declared description, namely DOM/canvas text mapping and highlight overlay logic for a specific reader environment. Therefore the description does not accurately represent what this supplied code chunk actually does.
The declared description centers on a text-to-speech skill that takes any URL/web page and converts article text into spoken audio/MP3. This code chunk instead implements extraction and synchronization infrastructure: parsing page content from multiple domains, reconstructing paragraph positions, creating highlight overlays, and handling Kindle-specific OCR and glyph decoding. Those are materially different capabilities from 'read aloud' and 'convert to audio.' While text extraction could support a TTS product, the supplied code does not itself perform audio generation, voice selection, or MP3 conversion, and it includes substantial undeclared behavior (highlighting, OCR, multi-site document/chat extraction, extension messaging). Therefore the description does not accurately represent what this code chunk actually does.
The declared purpose centers on turning arbitrary web pages into spoken audio/MP3. However, this code chunk is an extraction pipeline only: it parses DOM content, detects titles/language, handles site-specific selectors, deduplicates paragraphs, supports highlightable element mapping, and waits for SPA content to stabilize. It is heavily tailored to extracting text from chat AI websites and some document/reading sites. There is no evidence of speech synthesis APIs, audio encoding, MP3 creation, network calls to a TTS service, or playback controls. Therefore the supplied code does not accurately represent the declared 'read aloud / convert URL to audio' purpose; it represents a text-extraction subsystem with materially different primary behavior.
The code aligns with the text-extraction portion of the description but does not implement the core advertised audio functionality. It fetches a URL and extracts clean article text using Puppeteer and an extractor bundle, then outputs JSON. There is no code for speech synthesis, AI voice selection, audio encoding, MP3 creation, or reading content aloud. Therefore, the declared description materially overstates the skill’s actual behavior.
The declared description promises an end-to-end webpage-to-audio capability: take any URL, extract article text, and read the page aloud. This code chunk does only the TTS portion for one paragraph from an already-produced extraction file. While generating natural-voice MP3 audio is aligned with part of the description, the primary user-facing promise of accepting a URL and extracting webpage/article content is not implemented here. The code’s actual scope is narrower and materially different: paragraph-level audio generation from structured local input, with reliance on another tool (extract.js) for extraction.
The declared description centers on taking a webpage URL, extracting article text, and reading that page aloud. The actual code chunk only accepts a local text file path as input, reads its contents, calls a text-to-speech API, and writes an MP3 output file. While the TTS/audio generation portion is consistent with 'convert text to audio,' the key advertised capability—handling arbitrary web pages/URLs and extracting article text—is absent from this code. That is a material purpose mismatch, not just an implementation detail.
The description claims the skill can extract article text from any URL and convert it to audio (MP3). The code instead automates Chrome, loads a specific extension, navigates to a page, and sends a TOGGLE_READING message to trigger in-browser read-aloud playback. This supports 'listen to a webpage' at a high level, but materially differs from the declared implementation and capabilities: there is no text extraction pipeline, no MP3 creation/export, and the code relies on a local browser extension and browser automation not disclosed in the description.
The bundle includes targeted extraction for private/editable cloud document platforms such as Google Docs, Notion, Feishu, DingTalk, and Yuque, including logic for editable and dynamically loaded content. That enables access to potentially confidential business or personal documents outside the advertised article/webpage scope, raising serious data exposure risk.
The bundle implements extraction logic for many contexts far beyond ordinary public webpages, including AI chats, cloud docs, ebook readers, and novel platforms. That materially expands the data-access surface of the skill and enables collection of sensitive or copyrighted content inconsistent with the declared purpose, increasing the risk of over-collection and misuse.
The Kindle-specific code captures blob-backed page images, performs OCR, inspects proprietary render/token data, and builds highlight overlays. This goes beyond normal webpage text extraction and can recover book content from protected reader contexts, creating elevated privacy, copyright, and policy risk if invoked on user reading sessions.
Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.
extract-zip is pulled in by @puppeteer/browsers and is used to unpack browser archives during install/runtime setup. If a malicious or tampered archive were ever processed, the cited symlink/path traversal flaws could permit writes outside the intended extraction directory, potentially overwriting files in the build or execution environment.
Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.
Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.
ws is a direct transitive dependency of puppeteer-core and is used for browser protocol communication. A vulnerable WebSocket implementation can expose the process to memory disclosure or memory-exhaustion denial of service if an attacker can interfere with or influence WebSocket traffic to the browser automation channel, which is more plausible in a networked skill that fetches arbitrary web content.
The README explicitly states that extracted page text is sent to a remote TTS API, but it does not clearly warn users that webpage contents may leave the local device. This is risky because users may paste private, internal, paywalled, or sensitive URLs assuming processing is local, causing unintended disclosure of extracted content to a third-party service.
The skill invokes local Node scripts that, by description and setup, can perform network access and environment-dependent operations, yet the manifest declares no explicit tool scope or permission boundaries. This makes the skill harder to audit and allows capability creep beyond what users and the platform metadata clearly disclose.
Overly broad invocation phrases can cause the skill to trigger in situations where the user did not intend URL extraction, script execution, or Telegram delivery. In this skill, accidental activation is more dangerous because activation can lead to network access, local command execution, and outbound messaging rather than a harmless local transformation.
The skill instructs the agent to extract a Telegram chat ID from message metadata and use it in every outbound send operation. That introduces a separate messaging/exfiltration channel not disclosed in the high-level purpose, increasing the risk of sending content to unintended recipients or using the skill as a data-delivery mechanism.
Mandating use of an external message tool to transmit files extends the skill from local content processing into outbound communication. In context, that is more dangerous because webpage-derived or summarized content can leave the current interaction boundary and be delivered over Telegram without a clearly disclosed warning in the skill metadata.
The skill description does not warn users that processed webpage content will be sent as audio over Telegram. This undermines informed consent and increases privacy risk, especially if the extracted page contains personal, proprietary, or otherwise sensitive information.
The summary-only path still performs file creation and Telegram delivery, so even reduced-content flows retain the same outbound data-transfer risk. Because summaries may include sensitive extracted material, the context does not make this safer; it still creates an undeclared exfiltration path.
The WeRead extractor always returns "zh" from extractLanguage(), regardless of actual page content or user preference. This is a natural-language locale policy issue because it enforces a specific language outcome without offering opt-in or documenting a justified regional constraint.
The code contains dedicated extractors for multiple AI assistant sites that selectively harvest assistant responses and related conversation content. In a skill presented as reading webpages/articles aloud, this is scope expansion into potentially sensitive conversational data, which may include private prompts, outputs, and embedded secrets.
Detected: suspicious.env_credential_access