T09 · Insecure Skill Coding Practices
- Location
tools/crawl.py:16- Finding
API Credential and User Data Transmitted Over Plaintext HTTP
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This remote crawling skill is mostly coherent, but it sends API keys and browsing/search data to a plaintext HTTP endpoint and trusts server-provided download URLs, so it needs review before use.
Install only after reviewing the trust boundary. Do not use the public HTTP endpoint for API keys or sensitive browsing targets; prefer a self-hosted HTTPS OpenCrawl server, avoid authenticated/internal/private pages, and understand that crawl content may be processed by remote workers and stored in R2.
tools/crawl.py:16API Credential and User Data Transmitted Over Plaintext HTTP
tools/crawl.py:60Unrestricted Server-Controlled Download URL Enables Client-Side SSRF
requirements.txt:1Open-Ended Dependency Version Prevents Reproducible Installation
The crawl request sends the Bearer API key to a server endpoint whose base URL is taken directly from the OPENCRAWL_API_URL environment variable. If that variable is modified by an attacker or misconfigured, the tool will transmit credentials and user-supplied crawl targets to an arbitrary host, enabling credential theft and covert data exfiltration.
if selector:
body["selector"] = selector
res = requests.post(
f"{API_URL}/api/crawl",
headers={
"Authorization": f"Bearer {API_KEY}",
The search function posts user queries and the Bearer token to an API host controlled by the OPENCRAWL_API_URL environment variable. A malicious or poisoned environment can redirect these requests to an attacker-controlled server, exposing the API key and potentially sensitive search terms.
error_exit("OPENCRAWL_API_KEY environment variable not set")
try:
res = requests.post(
f"{API_URL}/api/search",
headers={
"Authorization": f"Bearer {API_KEY}",
The balance check sends the Authorization bearer token to the host specified by OPENCRAWL_API_URL without trust validation. That makes the token vulnerable to exfiltration if the environment variable is attacker-influenced or if the default endpoint is replaced in deployment.
error_exit("OPENCRAWL_API_KEY environment variable not set")
try:
res = requests.get(
f"{API_URL}/api/balance",
headers={"Authorization": f"Bearer {API_KEY}"},
timeout=10,
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
def check_status():
try:
res = requests.get(f"{API_URL}/api/status", timeout=10)
data = res.json()
except Exception as e:
error_exit(str(e))
The declared purpose is webpage crawling, but the documented behavior also includes multi-engine search and account/platform queries. This mismatch undermines informed consent and policy enforcement because users or orchestrators may approve the skill for one purpose while it performs broader network actions, increasing the chance of unintended data disclosure or abuse.
The README encourages users to send crawl jobs to a public server but does not clearly warn that requested URLs, rendered content, and extracted data are transmitted to and processed by an external third-party service, with results stored remotely. In a browsing/crawling skill, this omission can lead users to unknowingly expose proprietary, personal, or authenticated content to infrastructure outside their control.
The README states that community-contributed workers run real Chrome browsers and mentions cookie isolation, but it does not warn users not to crawl sensitive, internal, or authenticated pages. Because this skill routes browsing activity through external worker infrastructure, users may expose session-bearing content, confidential documents, or internal application data to untrusted systems if they assume the service is safe for private targets.
The skill declares required environment variables and clearly relies on outbound network access, but it does not define any explicit tool scope or permission boundaries. In an agent setting, this weakens least-privilege controls and can cause the skill to be invoked with broader capabilities than users expect, especially since it sends data to a remote third-party service.
The manifest description presents the skill as a crawler, while the body documents additional operational features such as search, balance checks, and status queries. In a security-sensitive agent ecosystem, incomplete disclosure of capabilities is itself risky because it can bypass user expectations, review workflows, or automated classification based on the manifest.
The skill states that rendered content is sent to an OpenCrawl server and then uploaded to Cloudflare R2, but it does not present this as a prominent user warning before use. This is dangerous because users may submit sensitive URLs or page content under the assumption of local processing, resulting in unintended transmission, storage, and possible retention of confidential data by third parties.
This call transmits user-provided crawl targets and an API credential to an external service. External transmission is expected for this tool's functionality, but in the current form it still presents a real privacy and secret-handling risk because the destination is configurable and the transfer is not clearly disclosed or constrained.
if selector:
body["selector"] = selector
res = requests.post(
f"{API_URL}/api/crawl",
headers={
"Authorization": f"Bearer {API_KEY}",
User-supplied crawl URLs and selectors are transmitted to a remote third-party service without any user-facing disclosure at execution time. In a skill context, this is dangerous because users may assume the tool operates locally while actually sending potentially sensitive internal URLs or private targets off-box to an external service.
The manifest says the skill crawls JavaScript-rendered webpages through remote Chrome browsers, which aligns with the crawl path. However, this file also exposes separate search functionality plus API balance and platform status checks, which go beyond the stated crawling-only purpose and represent additional product capabilities.
This request sends user search queries and the API token to an external service. That behavior matches the feature set, but it remains security-relevant because search terms may contain sensitive information and the endpoint is configurable, increasing the chance of accidental or malicious exfiltration.
error_exit("OPENCRAWL_API_KEY environment variable not set")
try:
res = requests.post(
f"{API_URL}/api/search",
headers={
"Authorization": f"Bearer {API_KEY}",
The dependency is specified as requests>=2.28.0, which allows installation of many future versions and does not guarantee reproducible builds. Unpinned dependencies increase supply-chain risk because different environments may resolve to different releases, including versions with breaking changes or newly introduced vulnerabilities.
requests>=2.28.0
Because requests is not pinned, it is impossible to determine from this manifest whether the installed version is one affected by known advisories. In a skill that performs web crawling and likely makes outbound HTTP requests, using an unresolved dependency version can expose the runtime to request-handling flaws, credential leakage issues, or other known defects if a vulnerable release is installed.
No suspicious patterns detected.