T09 · Insecure Skill Coding Practices
- Location
src/data_harvester/adapters/base.py:269- Finding
Server-Side Request Forgery Through Unrestricted Harvesting URLs
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This data-harvesting skill appears purpose-aligned, but it needs review because it can access broad data sources and has unsafe URL and credential handling.
Review this carefully before installing. Use it only with data sources you are authorized to access, avoid credentials unless endpoints are HTTPS and trusted, restrict outbound network access if possible, and check export paths because files may be created or overwritten. Do not run scheduled jobs unattended until you confirm their targets and storage locations.
src/data_harvester/adapters/base.py:269Server-Side Request Forgery Through Unrestricted Harvesting URLs
src/data_harvester/adapters/base.py:336API Source Credentials Can Be Sent Over Cleartext HTTP
src/data_harvester/openclaw_integration/api_client.py:53OpenClaw API Key Can Be Transmitted to an Unvalidated Cleartext Endpoint
Package name closely resembles a popular package, suggesting possible typosquatting. Attackers publish malicious packages with similar names to trick developers into installing them.
Using shell=True for build/test commands creates a tool-parameter abuse risk because the shell will parse metacharacters, expansions, and chained commands. Even though the current call sites use string literals, this pattern is unsafe in an agent skill/build context where commands, environment, PATH, or working tree may later become partially attacker-controlled, making arbitrary command execution more plausible.
def run_command(cmd, cwd=None):
"""运行命令并返回结果"""
print(f"执行: {cmd}")
result = subprocess.run(cmd, shell=True, cwd=cwd, capture_output=True, text=True)
if result.stdout:
print(f"输出: {result.stdout[:500]}...")
Using shell=True in a generic command wrapper is a real tool-parameter abuse risk because it gives the shell control over parsing, expansion, and metacharacter handling. In an agent or automation context this becomes more dangerous, since later code changes or externalized parameters could allow attacker-controlled input to execute arbitrary commands on the host.
def run_command(cmd, cwd=None):
"""Run command and return result"""
print(f"Execute: {cmd}")
result = subprocess.run(cmd, shell=True, cwd=cwd, capture_output=True, text=True, encoding='utf-8', errors='replace')
if result.stdout:
# Safe print for stdout
The README promotes web scraping, API access, database querying, scheduled execution, and file export, but gives no warning about privacy, authorization, rate limiting, data retention, or local system impact. This omission can cause unsafe use of a powerful automation skill, including collection of sensitive data or unattended recurring activity without informed consent. Because the skill is explicitly a data harvester with scheduling and export features, the missing safety guidance materially increases risk.
The README instructs users to run npx clawhub install data-harvester without pinning a specific package version. That can lead to installation of a newer or compromised release than the author intended, creating a supply-chain risk for users who follow the setup instructions. In the context of an install command for a skill ecosystem, this is more dangerous because users are likely to execute it directly.
The skill is explicitly designed for web scraping, API access, database queries, file reads, scheduling, and exporting data, yet the description provides no warning about network access, data collection scope, credential handling, or where harvested data may be written. This omission increases the risk that users enable broad data access or persistent exports without understanding the privacy and security consequences.
The installation command invokes npx clawhub without pinning a specific version, so users may execute whatever package version is current at install time. Because npx runs fetched code, a compromised or malicious upstream release could lead to arbitrary code execution on the user's system during installation.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
采集网页 https://example.com 保存为 data.json 定时采集 https://api.example.com/data 每天 09:00 导出数据为 Excel 报表
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
采集网页 https://example.com 保存为 data.json 定时采集 https://api.example.com/data 每天 09:00 导出数据为 Excel 报表
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
# 运行脚本
start_time = time.time()
result = subprocess.run(
[sys.executable, str(script_path)],
cwd=Path(__file__).parent,
capture_output=True,
This code file contains user-facing natural language that forces a specific language locale, beginning with the module description and continuing throughout status messages. Under the policy, hard-coding one language without user opt-in or a documented justification is a natural-language policy violation.
The script's operational prompts, errors, and success messages are presented in Chinese only, with no option for the user to select another language. This is a language/locale policy issue because it enforces a locale preference across the primary user interaction surface.
The helper executes shell commands via subprocess.run with shell=True, which is dangerous because any future variable interpolation into cmd would be interpreted by the shell and could enable command injection. In a build script that runs from a project workspace, this also increases exposure to PATH hijacking and unintended shell behavior if the environment or command strings are influenced by untrusted project contents or CI inputs.
def run_command(cmd, cwd=None):
"""运行命令并返回结果"""
print(f"执行: {cmd}")
result = subprocess.run(cmd, shell=True, cwd=cwd, capture_output=True, text=True)
if result.stdout:
print(f"输出: {result.stdout[:500]}...")
The module docstring says only 'initialize Git repository and prepare first commit', which suggests basic repository setup. In practice, the code also sets hard-coded Git identity values, stages all files, and creates a highly specific commit message containing project branding, release status, publishing plans, and revenue forecast, which goes beyond the stated intent of simple initialization.
The helper executes shell commands with subprocess.run(..., shell=True), which is dangerous because any future use of untrusted input in cmd can become command injection. In this file most commands are hardcoded, but the design pattern itself is unsafe and one command interpolates a multi-line commit message into a shell string, increasing fragility and injection risk if that content ever becomes variable.
def run_command(cmd, cwd=None):
"""Run command and return result"""
print(f"Execute: {cmd}")
result = subprocess.run(cmd, shell=True, cwd=cwd, capture_output=True, text=True, encoding='utf-8', errors='replace')
if result.stdout:
# Safe print for stdout
The top-level package description is written in Chinese, and the file provides no indication that the skill is region-specific or that users can opt into that locale. Under the policy rule, a natural-language locale constraint should either be optional for users or clearly justified as region-specific.
The package description field is user-facing metadata and is written only in Chinese. Because the file does not state that the package is intended exclusively for a Chinese-language context, this appears to impose a language choice without opt-in.
The manifest explicitly advertises automated data collection, processing, and export across multiple sources, but provides no warnings or constraints around privacy, consent, rate limiting, legal use, or potential system impact. In a skill centered on web scraping and data harvesting, this omission increases the risk of misuse for unauthorized collection or bulk exfiltration of sensitive data.
The configuration schema allows authenticated external sources via arbitrary URLs and a generic auth object, but does not disclose how credentials are stored, protected, or transmitted, nor does it warn about outbound network access. This creates meaningful risk of credential misuse, accidental secret exposure, or use of the skill to access internal or sensitive network resources.
This code file contains natural-language module documentation and multiple user-visible error/log messages exclusively in Chinese. The file does not indicate that the skill is China-region-specific or offer any language/locale opt-in, which can violate language/locale policy requirements.
User-visible exception messages such as adapter-disabled, fetch failure, and validation errors are emitted only in Chinese throughout the file. Without an explicit locale scope or opt-in, forcing a single language is a natural-language policy issue.
This code presents the skill name and user-facing help/output strings in Chinese, and there is no indication elsewhere in the file that users can choose another language or that the tool is intentionally limited to a Chinese-speaking locale. That creates a natural-language policy concern under the language/locale rule because the skill implicitly forces one language without opt-in.
The user-facing natural-language content in docstrings and log/error strings is consistently Chinese throughout the file, with no indication that another language can be selected. Under the policy, forcing a specific language without user choice can be a natural-language policy violation, especially because these strings are likely surfaced to operators and users.
The top-level docstring and class docstring describe the component as a '数据采集器'/'智能数据采集器主类', which implies collection functionality. However, the implementation also initializes exporters, exposes an export() method, and can create scheduled jobs, expanding behavior beyond what the documentation states.
This code opens the configured CSV output path in write mode, which will create or overwrite a file on disk. Although there is internal logging, there is no user-facing confirmation prompt or explicit warning in comments/docstrings that exporting will modify files, so the safety-critical file write lacks disclosure.
No suspicious patterns detected.