T09 · Insecure Skill Coding Practices
- Location
scripts/data-scraper.py:134- Finding
Bearer Credentials Can Be Sent to Arbitrary or Insecure Endpoints
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a coherent data-scraping skill, but users should review it because API tokens can be sent to any user-provided endpoint without built-in safeguards.
Install only if you are comfortable with a basic, partially implemented scraper. Do not pass valuable API tokens unless you fully trust the exact HTTPS endpoint, and avoid putting secrets directly on the command line. Treat local output files as potentially sensitive and check the target site's terms, robots.txt, and privacy rules before scraping.
scripts/data-scraper.py:134Bearer Credentials Can Be Sent to Arbitrary or Insecure Endpoints
整体方向与声明相近:代码的主要目的确实是网页/API 数据抓取,并支持批量处理与文件输出,没有发现与声明无关的高风险隐藏能力。但存在明显的描述-行为不一致:1)scrape_url 虽加载 BeautifulSoup 和预定义选择器模板,却没有执行任何实际字段抽取,result['data'] 始终为空,仅返回“完整功能需要安装依赖后实现”的说明,因此“从网页提取结构化数据”这一核心能力并未真正实现;2)输出格式参数宣称支持 json/csv/excel,但 save_output 仅真正实现 JSON,CSV 只是写入占位注释加 JSON,Excel 完全未实现而退回 JSON。相比之下,API 获取和批量处理能力基本存在。因此应判定为部分但实质性的能力夸大,属于 mismatch。
The markdown includes an example that passes an authorization token on the command line and writes fetched data to a local file, but it does not warn users that shell arguments may be exposed in shell history or process listings, nor that scraped/API data will be stored locally. The existing compliance section discusses legal scraping behavior, but not privacy or system-safety implications of credential use and saved outputs.
The skill advertises and instructs use of capabilities that imply network access plus local file read/write, but it declares no explicit tool scope or permissions. In an agent environment, missing scope boundaries increases the risk of over-privileged execution, unintended file access, or unrestricted outbound requests beyond what users expect.
The manifest description and top-level documentation are written in Chinese, and the skill does not indicate that other languages are supported or that the user can choose a preferred locale. This creates a natural-language policy concern because the skill appears to impose a specific language without opt-in or justification.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
uv run scripts/data-scraper.py scrape --urls-file urls.txt --output data.json
uv run scripts/data-scraper.py api --endpoint "https://api.example.com/data" --auth "Bearer TOKEN"
### 高级选项
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
uv run scripts/data-scraper.py scrape --urls-file urls.txt --output data.json
uv run scripts/data-scraper.py api --endpoint "https://api.example.com/data" --auth "Bearer TOKEN"
### 高级选项
The manifest explicitly advertises web/API extraction, batch processing, and data collection capabilities but provides no privacy, authorization, or data-handling notice. In a scraping tool, that omission increases the risk that users deploy it against personal, restricted, or rate-limited sources without clear safeguards or expectations, which can lead to privacy breaches or policy violations.
This code file contains natural-language strings that define the tool and its interface entirely in Chinese, including the module docstring and later CLI help text. That effectively forces a single language for users without offering a choice or documenting a justified locale restriction.
The code comment at L140 says it is processing authentication, and the branch at L145-L148 parses user:pass credentials for basic auth, but the resulting credentials are never attached to the requests.get call at L150. This contradicts the apparent intent of the code/documentation because Bearer auth is applied while basic auth is only recorded in the result metadata.
The skill documentation appears to force a specific language/locale for all user-facing instructions and examples, with no opt-in or alternative language guidance. Under the stated policy, a language restriction should be optional or explicitly justified as region-specific.
The display name and description are presented in Chinese, but the manifest does not indicate that language is selectable or that the skill is intended only for a Chinese-speaking locale. This can violate language/locale policy when users are not given an opt-in or documented locale constraint.
This requirements file includes human-readable comments entirely in Chinese (for example, dependency descriptions on L01, L03, L06, L09, L12, and L15). Under the policy, forcing a specific language without user opt-in or documented justification is a natural-language policy concern.
Using an unpinned dependency range for requests allows future installs to resolve to different versions over time, including versions with newly introduced vulnerabilities or breaking security behavior. In a data-scraping skill that makes outbound HTTP requests, dependency drift increases supply-chain risk and can expose network-facing functionality to known flaws.
# Data Scraper 依赖
# HTTP 请求
requests>=2.28.0
# HTML 解析
beautifulsoup4>=4.11.0
Requests has multiple known advisories, and because the manifest does not pin a version, there is no way to verify whether deployments avoid affected releases. Given this skill performs web/API access, the uncertainty is more dangerous than in an offline-only tool because the dependency sits directly on a network-facing path.
An unpinned beautifulsoup4 dependency creates non-reproducible builds and supply-chain uncertainty, even if the package is not highly security-sensitive by itself. Different resolved versions may change parser behavior or include security defects not accounted for during testing.
requests>=2.28.0
# HTML 解析
beautifulsoup4>=4.11.0
# Excel 输出(可选)
openpyxl>=3.0.0
Leaving openpyxl unpinned can allow installation of versions affected by XML-related issues, including historical XXE-style problems. In a scraping/export tool that may process untrusted spreadsheet content or generate Excel outputs in shared environments, uncontrolled version selection raises the chance of deploying a vulnerable release.
beautifulsoup4>=4.11.0
# Excel 输出(可选)
openpyxl>=3.0.0
# 数据处理(可选)
pandas>=1.5.0
Openpyxl has known historical advisories, and the current manifest does not provide enough version precision to determine whether an affected release could be installed. Because spreadsheet parsers may process attacker-controlled files or XML content, version ambiguity can translate into real risk if unsafe releases are resolved.
An unpinned pandas dependency introduces supply-chain and reproducibility risk by allowing installs to vary across environments. Although the direct exploitability depends on how pandas is used, unresolved version drift can expose the skill to known issues in data parsing or deserialization-related code paths.
openpyxl>=3.0.0
# 数据处理(可选)
pandas>=1.5.0
# JavaScript 渲染支持(可选,高级功能)
# playwright>=1.30.0
Pandas has at least one cited advisory, and without a pinned version the installed release cannot be verified as safe. The practical danger depends on whether the skill uses risky deserialization or parsing features, but the manifest still leaves avoidable uncertainty in a data-processing component.
No suspicious patterns detected.