T09 · Insecure Skill Coding Practices
- Location
scripts/extract_emails.py:118- Finding
Unrestricted Website Crawling Enables SSRF and Disables TLS Authentication
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a disclosed lead-generation scraper, but it has broad, weakly contained network crawling and persistent contact-data export risks that users should review before installing.
Install only if you are comfortable running a web-scraping lead-generation tool. Use an isolated environment with limited network reach, avoid sensitive internal networks, review proxy settings, treat exported CSV files as untrusted, and define your own retention/deletion and legal-compliance process for collected contact data.
scripts/extract_emails.py:118Unrestricted Website Crawling Enables SSRF and Disables TLS Authentication
scripts/find_customers.py:161CSV Formula Injection Through Untrusted Search and Website Data
scripts/requirements.txt:1Unpinned and Unverified Third-Party Dependencies
The script forwards proxy settings from environment variables or the CLI directly into outbound HTTP requests without validation. An attacker who can influence the runtime environment can force all requests through a malicious proxy, enabling interception, traffic logging, and response tampering; this matters more here because the skill scrapes external sites and may process many targets automatically.
emails = set()
try:
response = requests.get(
url,
headers=get_random_headers(),
timeout=10,
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
return False
try:
response = requests.get(
engine['test_url'],
headers=get_random_headers(),
timeout=5,
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
headers = get_random_headers()
headers['Accept-Language'] = 'en-US,en;q=0.9'
response = requests.get(url, headers=headers, timeout=15, proxies=PROXIES)
soup = BeautifulSoup(response.text, 'html.parser')
for div in soup.select('div.g'):
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
headers = get_random_headers()
headers['Accept-Language'] = 'en-US,en;q=0.9'
response = requests.get(url, headers=headers, timeout=15, proxies=PROXIES)
soup = BeautifulSoup(response.text, 'html.parser')
for div in soup.select('div.g'):
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
headers = get_random_headers()
headers['Accept-Language'] = 'en-US,en;q=0.9'
response = requests.get(url, headers=headers, timeout=15, proxies=PROXIES)
soup = BeautifulSoup(response.text, 'html.parser')
for div in soup.select('div.g'):
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
headers = get_random_headers()
headers['Accept-Language'] = 'en-US,en;q=0.9'
response = requests.get(url, headers=headers, timeout=15, proxies=PROXIES)
soup = BeautifulSoup(response.text, 'html.parser')
for div in soup.select('div.g'):
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
headers = get_random_headers()
headers['Accept-Language'] = 'en-US,en;q=0.9'
response = requests.get(url, headers=headers, timeout=15, proxies=PROXIES)
soup = BeautifulSoup(response.text, 'html.parser')
for div in soup.select('div.g'):
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
headers = get_random_headers()
headers['Accept-Language'] = 'en-US,en;q=0.9'
response = requests.get(url, headers=headers, timeout=15, proxies=PROXIES)
soup = BeautifulSoup(response.text, 'html.parser')
for div in soup.select('div.g'):
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
url = f"{instance}/search?q={quote(query)}&format=json"
headers = get_random_headers()
response = requests.get(url, headers=headers, timeout=15, proxies=PROXIES)
data = response.json()
for result in data.get('results', []):
声明描述的是一个完整的“找客户/获客”技能:先搜索潜在客户公司,再提取联系方式,并维护已收集数据。实际代码只是一个邮箱提取脚本,其核心输入是单个domain,功能集中在访问官网常见页面、正则/HTML中提取邮箱、可选WHOIS查询、简单过滤和优先级排序,最后输出JSON。声明中的第2、3、4项(官网提取、WHOIS查询、邮箱有效性基础过滤)与代码基本一致,但声明中的主要前置能力(关键词+地区搜索潜在客户公司)和后续数据能力(缓存、去重、简单查询)均不存在。因此,代码实际行为只覆盖了描述中的一部分子功能,且未实现该技能的主要用途,应判定为描述与行为不匹配。
声明描述的是一个以“找客户/获客”为主的技能,核心能力包括外部搜索潜在客户、访问官网抓取邮箱、WHOIS 查询和邮箱验证。而提供的代码块 primary purpose 明显是“查询已收集的客户数据”:它从 CacheManager 读取已有 companies 和 emails,按名称/邮箱过滤,展示列表、统计并导出结果。代码没有任何网络访问、爬取、WHOIS、邮箱校验或去重逻辑。虽然声明中也提到了“简单查询已收集客户数据”,这与代码的一部分相符,但代码块整体只覆盖该次要子功能,没有体现声明中的主要获客能力。因此该描述与此代码块的实际行为存在实质性不匹配。
声明描述的是一个较完整的“获客+联系方式提取”工具,核心能力包括邮箱抓取、WHOIS邮箱查询、邮箱校验、缓存和历史数据查询。但提供的代码只是一个多搜索引擎公司搜索脚本:它检测各搜索引擎可用性,发起搜索请求,解析搜索结果页面,过滤常见非目标域名,并对结果按域名去重后输出结构化JSON。代码没有访问搜索结果对应的公司官网页面,也没有任何邮箱提取正则、WHOIS请求、邮箱校验逻辑、缓存存储、数据库/文件索引查询能力。虽然“根据关键词和地区搜索潜在客户公司”这一部分与声明部分一致,且有基本去重,但声明中的多项关键能力并未实现,因此描述明显高于实际行为,构成实质性不匹配。
声明描述的是一个“获客/找客户”技能,核心功能应是搜索公司、抓取联系方式并管理结果数据。实际代码则是一个独立的邮箱验证脚本,主要做格式校验、DNS/MX 检查和一次性邮箱识别。虽然“验证邮箱格式有效性”属于声明中的一个子功能,但这段代码没有体现该技能的大部分主要能力,且其实际主用途明显缩窄为邮箱校验,而非潜在客户搜索与联系方式提取。因此描述与代码行为存在明显不匹配。
The README explicitly promotes collecting, caching, querying, and exporting company contact data, including email addresses and WHOIS-derived information, but provides only a vague 'comply with local laws' note later in the document. That omission increases the risk that users will treat personal or quasi-personal contact data as unrestricted, store it indefinitely, and reuse it without consent, creating privacy, retention, and misuse concerns.
The usage example shows a workflow of finding companies and then sending each one a sales email, which normalizes bulk outreach without any explicit anti-spam, consent, or rate-limit warning. In the context of a lead-generation skill that extracts emails at scale, this materially increases the likelihood of misuse for unsolicited marketing or spam campaigns.
The skill declares broad capabilities to read environment variables, use the network, and write local files, but it does not publish any explicit tool scope or permission boundary. In a skill that scrapes websites, queries WHOIS, and caches customer data locally, that omission increases the risk of unintended data access, persistence, and exfiltration beyond what users expect.
The skill is designed to collect contact data from company websites and WHOIS records and then validate, cache, and export that data, yet the top-level description lacks a clear user-facing privacy and data-handling warning. In this context, the omission is more dangerous because the skill targets personal or business contact information at scale, which can create compliance, consent, and misuse risks.
The documentation specifies persistent local caches for companies, searches, and emails, including permanent storage for some records, but does not pair that with a prominent user-facing notice or retention controls. Because this skill builds a reusable customer-data repository, silent persistence materially increases privacy exposure, accidental disclosure, and secondary use risk.
The module persistently stores scraped company and email data to local JSON files without any consent flow, visibility controls, or protective measures. In this skill’s context, the data being cached consists of lead information and contact emails, so silent retention increases privacy, compliance, and unintended disclosure risk if the host environment is shared or logs/files are later exposed.
The export command dumps all cached companies and emails directly to stdout, which can expose accumulated contact data to terminals, shell history, calling processes, orchestration logs, or other users on the system. Because this skill is specifically designed to harvest and cache contact information, bulk export materially increases the chance of accidental data leakage.
The code disables TLS certificate verification for HTTPS requests, which allows man-in-the-middle attackers to intercept or modify fetched website content without detection. In this skill, that can poison extracted email results, expose browsing targets, and combine dangerously with proxy support to make interception trivial.
url,
headers=get_random_headers(),
timeout=10,
verify=False,
allow_redirects=True,
proxies=PROXIES
)
This file contains natural-language strings that establish the skill identity and user-facing behavior entirely in Chinese, including the main description and command help text, with no indication that another language is supported or that Chinese is required for a region-specific purpose. Per policy, forcing a specific language without opt-in is a locale/language policy violation.
The module docstring says the script only queries collected customer data, but the implementation performs file-writing exports. This mismatch can cause operators, reviewers, or calling agents to underestimate the script's side effects and unintentionally permit bulk extraction of stored emails and customer records.
The script is presented as a customer-query utility, but it also supports bulk export of the entire collected company and email dataset to JSON or CSV files. That creates a data-exfiltration pathway for potentially sensitive contact information with no access control, confirmation prompt, scope restriction, or audit mechanism, which is risky in a lead-generation skill that aggregates emails.
This code defines an export command that writes collected customer records and associated email addresses to JSON or CSV files. Although it prints a success message after writing, there is no user-facing warning in the command description, help text, or nearby comments that the operation exports potentially sensitive collected data to disk.
No suspicious patterns detected.