Back to skill

Security audit

网络安全情报爬虫

Security checks for vulnerabilities and agentic risk

Overview

This skill appears to be a real security-news and vulnerability crawler, but it has unsafe credential, scheduling, and transport practices users should review before installing.

Install only if you are comfortable granting it network access, IMA note-write credentials, local state/log files, and scheduled operation. Use dedicated limited-scope IMA and MiniMax credentials, remove the global openclaw.json credential fallback, avoid hardcoding secrets, do not run the system-wide cron command, and require TLS verification for all feeds before relying on the collected intelligence.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/vuln_crawler.py:304
Finding

TLS Certificate Validation Disabled for External Vulnerability Feeds

Content
View full analysis

Vulnerability Details

File Location: scripts/vuln_crawler.py, lines 304-312, with unsafe call sites at lines 395 and 441
Vulnerability Type: Improper certificate validation
Risk Level: High

python
def http_get(url, timeout=15, verify_ssl=True):
    headers = {"User-Agent": "Mozilla/5.0 (compatible; VulnBot/2.0)"}
    ctx = ssl.create_default_context()
    if not verify_ssl:
        ctx.check_hostname = False
        ctx.verify_mode = ssl.CERT_NONE
    try:
        req = Request(url, headers=headers)
        resp = urlopen(req, timeout=timeout, context=ctx)

The insecure mode is explicitly enabled for two sources:

python
raw = http_get("https://cxsecurity.com/rss/wl", verify_ssl=False)
python
raw = http_get("https://api.anquanke.com/feed", verify_ssl=False)

Technical Analysis

Setting check_hostname to False and verify_mode to ssl.CERT_NONE removes both certificate-chain and hostname authentication. Encryption alone does not establish the identity of the feed server.

A network-positioned attacker can therefore provide an arbitrary certificate and impersonate either feed. The crawler treats the resulting RSS entries as trusted vulnerability intelligence, processes their descriptions, and uploads the generated content to IMA.

This behavior is not necessary for the declared crawler functionality. A source with an invalid certificate should be repaired, replaced, or accessed through a separately authenticated channel rather than globally bypassing TLS validation for that request.

Attack Path

  1. An attacker obtains a network interception position, such as control of a proxy, gateway, DNS response, or compromised network.
  2. The attacker intercepts a request to one of the feeds using verify_ssl=False.
  3. The attacker presents an untrusted certificate and returns a forged RSS document.
  4. The crawler accepts the certificate and parses the forged en ...[truncated 649 chars]
Remediation
View remediation

Remediation Suggestions

  • Remove the verify_ssl bypass and always use a default validating SSL context.
  • Do not set CERT_NONE or disable hostname checking.
  • Replace feeds that cannot provide a valid certificate.
  • If a private certificate authority is legitimately required, configure a narrowly scoped CA bundle rather than disabling validation.
  • Add automated tests confirming that expired, self-signed, hostname-mismatched, and untrusted certificates are rejected.
  • Consider signing or independently validating feed content when feed integrity is security-sensitive.

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/vuln_crawler.py:33
Finding

Skill Reads a Shared API Credential from Global OpenClaw Configuration

Content
View full analysis

Vulnerability Details

File Location: scripts/vuln_crawler.py, lines 33-43
Vulnerability Type: Excessive credential access outside the Skill directory
Risk Level: Medium

python
MINIMAX_API_KEY = os.environ.get("MINIMAX_API_KEY", "")
if not MINIMAX_API_KEY:
    _cfg = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(__file__))), "openclaw.json")
    try:
        with open(_cfg, "r") as f:
            _j = json.load(f)
        raw_key = _j.get("models", {}).get("providers", {}).get("minimax", {}).get("apiKey", "")
        if raw_key and raw_key != "__OPENCLAW_REDACTED__":
            MINIMAX_API_KEY = raw_key
    except Exception:
        pass

Technical Analysis

When MINIMAX_API_KEY is not explicitly provided, the script traverses outside its own directory and reads the Agent's global openclaw.json configuration. It then extracts a shared provider credential.

The documented command already supports supplying MINIMAX_API_KEY through the environment, so access to the global configuration is not required for the crawler's core functionality. This fallback weakens secret isolation and allows the Skill to use a globally configured credential without explicit per-run assignment.

The reviewed code subsequently sends the key only to api.minimaxi.com; no unrelated key-exfiltration destination was identified.

Attack Path

  1. The crawler is invoked without an explicit MINIMAX_API_KEY.
  2. The fallback logic locates and reads the parent OpenClaw configuration.
  3. The script extracts the globally configured MiniMax API key.
  4. Untrusted feed descriptions cause translation requests to be issued using that shared key.
  5. The shared account incurs requests and associated billing or quota consumption without a separately scoped Skill credential.

Impact Assessment

The Skill gains access to a credential beyond its local files and explicitly supplied environment. It ...[truncated 330 chars]

Remediation
View remediation

Remediation Suggestions

  • Remove the fallback that reads openclaw.json.
  • Require MINIMAX_API_KEY to be supplied explicitly through a scoped secret provider or environment variable.
  • Use a dedicated key with minimal permissions, a limited quota, and separate billing controls for this Skill.
  • Fail safely with a clear error when translation is requested without a configured key, or continue without translation.
  • Restrict global configuration permissions so Skill processes cannot read unrelated provider credentials.
  • Document the exact external service, data transmitted, and expected credential scope.

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
SKILL.md:63
Finding

Manual Trigger Command Executes Every System Hourly Job

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 63
Vulnerability Type: Overbroad system-level task execution
Risk Level: Medium

bash
run-parts /etc/cron.hourly

Technical Analysis

The documentation presents run-parts /etc/cron.hourly as a way to manually trigger the crawler's cron task. This command does not select the crawler. It attempts to execute every eligible program in the system-wide hourly cron directory.

This exceeds the minimum privilege and execution scope needed to run a single news crawler. Depending on system ownership and configuration, the command may require elevated privileges or execute unrelated maintenance programs with system-level effects.

The package does not contain code that creates a cron entry or installs a persistent service. Therefore, this is not evidence of implemented system persistence. The issue is the overbroad manual command.

Attack Path

  1. A user or Agent follows the documented manual-trigger instruction.
  2. The command enumerates all eligible entries under /etc/cron.hourly.
  3. Every matching hourly task is executed rather than only the crawler.
  4. Unrelated jobs perform their configured operations, potentially with elevated privileges if the command is run as an administrator.
  5. Those jobs can consume resources or modify system state outside the crawler's intended scope.

Impact Assessment

The command may trigger unrelated backups, cleanup operations, package tasks, log processing, or locally installed administrator jobs. The exact effects depend on the host's /etc/cron.hourly contents and the invoking user's privileges.

It does not itself create persistence or automatically grant elevated privileges. However, it encourages execution beyond the Skill's legitimate scope and can cause privileged side effects when run with elevated rights.

Remediation
View remediation

Remediation Suggestions

  • Remove the run-parts /etc/cron.hourly instruction.
  • Provide a direct command that invokes only the crawler, such as the intended wrapper script or Python entry point.
  • Do not instruct users to use sudo or system-wide cron directories for a user-level Skill.
  • If scheduling is supported, use a dedicated user-owned cron entry with an absolute executable path and a minimal environment.
  • Include commands for inspecting and invoking only the named crawler job.
  • Ensure installation and removal of any scheduled job require explicit user consent.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:136
Finding

Documentation Promotes Hardcoded IMA Credentials

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 136-139
Vulnerability Type: Insecure secret-storage guidance
Risk Level: Medium

The relevant documentation states, translated into English, that required environment credentials are hardcoded in scripts/run.sh and identifies the following values:

text
IMA_OPENAPI_CLIENTID
IMA_OPENAPI_APIKEY

Technical Analysis

Storing reusable API credentials directly in a workspace shell script exposes them to source inspection, backups, accidental repository commits, diagnostic archives, and other users with workspace read access.

The referenced scripts/run.sh is absent from the audited package, so actual embedded credential values and file permissions could not be verified. The confirmed issue is that the Skill documentation explicitly describes hardcoded credential storage as the expected deployment arrangement.

Attack Path

  1. A deployment follows the documented practice and embeds IMA credentials in scripts/run.sh.
  2. The workspace is copied, backed up, shared, archived, committed, or read by another local account.
  3. The embedded client identifier and API key are recovered from the script.
  4. The credentials are reused against the IMA API.
  5. The attacker obtains whatever note access and modification capabilities are granted to those credentials.

Impact Assessment

If the documented deployment practice is followed, disclosure can permit unauthorized IMA API use within the credential's scope, including potentially creating or modifying notes.

No actual credential value or run.sh file was present in the reviewed artifact. The exposure impact is therefore conditional on the external deployment matching the documentation.

Remediation
View remediation

Remediation Suggestions

  • Remove all guidance that recommends hardcoding credentials in scripts.
  • Obtain credentials through a permission-restricted secret manager or explicitly configured runtime environment.
  • Keep secret-bearing files outside source-controlled Skill directories.
  • Apply restrictive file permissions and avoid exposing secrets through command-line arguments or logs.
  • Use a dedicated IMA credential with only the note and folder permissions required by the crawler.
  • Rotate any credential that has previously been embedded in a script or committed to version control.
  • Add secret scanning to packaging and release workflows.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/vuln_crawler.py:217
Finding

Untrusted Feed Descriptions Are Inserted into an LLM Translation Instruction

Content
View full analysis

Vulnerability Details

File Location: scripts/vuln_crawler.py, lines 217-230 and 525-527
Vulnerability Type: Indirect prompt injection through external feed content
Risk Level: Medium

python
payload = json.dumps({
    "model": model,
    "max_tokens": 300,
    "messages": [
        {
            "role": "user",
            "content": (
                "Translate to Chinese. Rules:\n"
                "1. Keep CVE IDs in English (CVE-YYYY-NNNNN)\n"
                "2. Keep component names in English (apache, nginx, mysql, etc.)\n"
                "3. Keep technical terms in English (RCE, SSRF, XSS, SQL Injection, zero-day, etc.)\n"
                "4. Translate rest to fluent Chinese.\n"
                "5. Output ONLY the translated text, no notes or explanations.\n\n"
                f"Text: {text}"
            ),
        }
    ],
}).encode("utf-8")

The generated result is placed directly into a note:

python
if v["description"]:
    translated_desc = translate_to_chinese(v["description"])
    lines.append(f"  {translated_desc[:300]}")

Technical Analysis

Vulnerability descriptions obtained from external feeds are attacker-influenced data. The script concatenates that data into the same user message that contains the translation instructions, without a trusted structural boundary or output validation.

A malicious description can contain instructions telling the model to ignore the translation request and generate attacker-selected text. The current check only attempts to identify a few forms of model reasoning output; it does not detect arbitrary prompt injection or verify that the response is a faithful translation.

The risk is amplified for feeds where TLS verification is disabled, although a malicious or compromised legitimate feed publisher could exploit the same issue without network interception.

Attack Path

  1. An attacker publishes, compromis ...[truncated 952 chars]
Remediation
View remediation

Remediation Suggestions

  • Treat all feed fields as untrusted data.
  • Place the translation policy in a dedicated system message and provide the source text as a separately delimited data field.
  • Use a translation-specific API or deterministic translation model where possible.
  • Validate that the output preserves expected CVE identifiers and does not introduce new URLs, commands, or unrelated claims.
  • Reject or escape descriptions containing instruction-like patterns when a faithful translation cannot be established.
  • Store the original description alongside the translation so users can verify accuracy.
  • Apply strict input and output length limits and monitor anomalous API usage.
  • Restore TLS verification for all feeds to reduce the opportunity for network-based content injection.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (31)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The documented behavior significantly exceeds the declared purpose: besides RSS news crawling, it introduces CVE aggregation from non-RSS APIs, machine translation via a third party, and storage into separate notebook flows. This mismatch can cause an agent or user to authorize a seemingly simple news crawler while it performs broader data collection, external data sharing, and credentialed actions than expected.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
91% confidence
Finding

The documented use of a raw rm command for state reset is a form of tool-parameter abuse risk because it normalizes direct shell deletion of files in an agent-manageable workspace. In agent environments, shell-based destructive operations can be mis-targeted, replayed, or expanded, causing data loss and potentially deleting the wrong files if paths are altered or templated unsafely.

Content

Scanner excerpt · SKILL.md (reported line 61)May include surrounding context.

md
| 手动执行一次 | `bash ~/.openclaw/workspace/skills/sec-news-crawler/scripts/run.sh` |
| 查看日志 | `cat ~/.openclaw/workspace/logs/sec_news_cron.log` |
| 查看上次运行状态 | `cat ~/.openclaw/workspace/data/sec_news_last_run.json` |
| 重置去重库(重新抓取所有文章) | `rm ~/.openclaw/workspace/data/sec_news_seen.json` |
| 查看 cron 配置 | `crontab -l \| grep sec_news` |
| 手动触发 cron | `run-parts /etc/cron.hourly`(系统级)|

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document states that IMA credentials are used and even hardcoded in scripts, but it provides no warning about the sensitivity of those secrets. Hardcoded API credentials can be exposed through source files, shell history, logs, backups, or agent file access, enabling unauthorized note access, data modification, and long-lived account compromise.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill is presented as a general security news crawler, but the code actually collects vulnerability intelligence from NVD and other vuln-focused feeds. This scope mismatch can cause users or operators to grant the skill broader trust than warranted and may route higher-risk exploit/vulnerability content into systems that were only approved for general news ingestion.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill documents capabilities that include shell execution, filesystem access, network access, and environment-variable use, but it does not declare any explicit tool scope or permission boundaries. In an agent setting, this increases the chance of overbroad execution and makes it harder to enforce least privilege for a skill that can read/write data, invoke scripts, and handle credentials.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

该描述将触发短语与“或需要查看/管理安全新闻爬虫任务时使用此 skill”并列,后者属于模糊场景判断,未明确哪些具体表达会触发、哪些不会触发。虽然“抓取安全新闻”等短语有一定领域限定,但“查看/管理…任务时”仍可能导致在普通讨论相关新闻任务时被意外调用。

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation includes a destructive deletion command to reset the deduplication database without prominent warning or confirmation guidance. In an agent context, exposing raw delete commands can lead to unintended state loss, duplicate reprocessing, noisy downstream writes, and operational disruption if invoked automatically or by mistake.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 62)May include surrounding context.

md
| 查看日志 | `cat ~/.openclaw/workspace/logs/sec_news_cron.log` |
| 查看上次运行状态 | `cat ~/.openclaw/workspace/data/sec_news_last_run.json` |
| 重置去重库(重新抓取所有文章) | `rm ~/.openclaw/workspace/data/sec_news_seen.json` |
| 查看 cron 配置 | `crontab -l \| grep sec_news` |
| 手动触发 cron | `run-parts /etc/cron.hourly`(系统级)|

## 添加/移除 RSS 源

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The documentation suggests manually running system-wide hourly cron jobs via run-parts /etc/cron.hourly without warning that this may execute unrelated tasks across the host. In shared or production environments, that can trigger unintended maintenance jobs, data processing, or privileged actions far beyond this skill's scope.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation expands the skill into a second, materially different workflow for vulnerability-intelligence collection and publication. Hidden scope expansion is dangerous because users may invoke the skill expecting passive RSS aggregation, while the skill actually performs additional network calls, processing, and writes into distinct notebooks and state files.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The vulnerability crawler adds third-party translation functionality unrelated to the stated RSS-crawling purpose, meaning content is sent to an external service without being clearly scoped in the main skill description. This broadens the trust boundary and can leak article or vulnerability content, operational metadata, or usage patterns to an external provider.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

文档明确声明“英文描述自动翻译为中文”,并给出固定翻译规则,但没有提供保留原文、双语输出或由用户选择语言的选项。这构成语言/locale 强制策略,可能违反要求提供用户语言自主权的组织政策。

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The top-level documentation says '每条新闻单独一篇笔记', indicating one note per article. However, the implementation collects articles into daily_articles and writes one note per date via build_daily_note_content and import_doc, so the documented behavior directly contradicts actual behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring describes the skill entirely in Chinese and hard-codes Chinese notebook naming and output conventions, indicating the skill is intended to operate in a specific language/locale. There is no natural-language indication that users may choose another language or that the locale restriction is required for a region-specific compliance purpose.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module docstring states that English content is automatically translated into Chinese, and the translation prompt explicitly instructs output in Chinese only. This imposes a specific language behavior without offering user choice or documenting it as an opt-in regional constraint.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The crawler sends collected content to a third-party translation API, which exceeds the narrow expectation of a local news/vuln ingestion workflow. Even if the content is public, this creates undisclosed data egress and introduces additional supply-chain and confidentiality risk through an external processor.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The code reads a sensitive API key from a local openclaw.json file outside the immediate crawler configuration path. This broadens the skill's access to local secrets beyond its stated role and can lead to unauthorized secret use if the skill is deployed in an environment containing unrelated credentials.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/vuln_crawler.py (reported line 45)May include surrounding context.

python
MINIMAX_API_KEY = raw_key
    except Exception:
        pass
MINIMAX_BASE_URL = "https://api.minimaxi.com/anthropic"

# 不应翻译的专业术语(保持英文原样)
TECH_TERMS = {

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Vulnerability descriptions are sent to a third-party translation API without explicit user disclosure or consent. Even if descriptions are typically public, transmitting aggregated security intelligence externally can leak operational interests, internal enrichment context, or future expanded inputs, and it adds a third-party processing dependency to the pipeline.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

SSL verification is explicitly disabled for external RSS fetches, allowing a man-in-the-middle attacker to tamper with feed contents or inject arbitrary data into the generated notes. In this skill's context, that means untrusted external content can be silently modified before being stored and potentially redistributed as trusted vulnerability intelligence.

Content

No source excerpt is available for this finding.

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
99% confidence
Finding

Using verify_ssl=False as a default is an unsafe transport setting that permits interception and tampering of supposedly HTTPS-protected content. Because this skill ingests security advisories and stores them as notes, a network attacker could poison downstream intelligence with forged or altered entries.

Content

Scanner excerpt · scripts/vuln_crawler.py (reported line 405)May include surrounding context.

python
# ── cxsecurity RSS ─────────────────────────────────────────────────────────

def fetch_cxsecurity():
    raw = http_get("https://cxsecurity.com/rss/wl", verify_ssl=False)
    if not raw:
        return []
    try:

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/vuln_crawler.py (reported line 447)May include surrounding context.

python
def fetch_anquanke():
    """安全客是国内为数不多仍然可用的安全资讯 RSS"""
    raw = http_get("https://api.anquanke.com/feed", verify_ssl=False)
    if not raw:
        return []
    try:

Unsafe Defaults

Medium
Category
Tool Misuse
Confidence
99% confidence
Finding

This second verify_ssl=False occurrence has the same consequence: HTTPS is downgraded to unauthenticated transport in practice. In a threat-intelligence workflow, that materially increases the risk of feed poisoning, false reporting, and trust erosion in stored outputs.

Content

Scanner excerpt · scripts/vuln_crawler.py (reported line 447)May include surrounding context.

python
def fetch_anquanke():
    """安全客是国内为数不多仍然可用的安全资讯 RSS"""
    raw = http_get("https://api.anquanke.com/feed", verify_ssl=False)
    if not raw:
        return []
    try:

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The documentation is inconsistent about secret handling: one section describes environment-variable execution while another says IMA credentials are hardcoded in scripts. This inconsistency is itself risky because it encourages insecure secret storage practices and makes operators more likely to embed credentials in files that may be read, logged, versioned, or exposed.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file uses Chinese throughout, including the title, headings, and operational guidance, but does not indicate that the skill is China-specific or that users may choose another language. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.