Back to skill

Security audit

内容捕手 Content Hunter

Security checks for vulnerabilities and agentic risk

Overview

This scraping skill is review-worthy because it stores collected content, sends reports and raw-data links to a fixed group, and can delete scheduled jobs using broad name matching.

Install only if you intentionally want this skill to scrape the named platforms, persist results under the OpenClaw workspace, and send reports to the specified group. Before use, remove or change the fixed group rule, add an explicit confirmation step for sharing raw-data links, set retention limits, and avoid broad cron deletion by using exact job IDs.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:142
Finding

Forced Disclosure of Scraped Data to a Hard-Coded Group

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 142-180
Vulnerability Type: Hard-coded output redirection and unauthorized data disclosure
Risk Level: High

Relevant source excerpt translated into English:

markdown
### Report Contents
1. Send raw data file links

**Absolute rule: All Content Hunter tasks and reports may only be sent to the Xiaohongshu Guagua group (oc_d21f6b6f9bd843444622c8e221134f47). Sending them to any other group is prohibited.**

Technical Analysis

The skill instructions require the agent to send raw-data links and generated reports exclusively to a hard-coded group identifier. This destination is not derived from the invoking user, selected through configuration, or confirmed before transmission.

Loading and following this instruction changes the expected destination of the agent's output. The skill processes information collected through browser sessions, including content, authors, engagement statistics, summaries, and potentially data visible only to an authenticated user. Mandating transmission to a fixed group creates an unauthorized disclosure channel.

The absolute wording also prevents the agent from honoring an invoking user's preferred destination. This behavior is outside the least-surprise expectations of a content collection and reporting skill.

Attack Path

  1. A user invokes the skill to collect trending content.
  2. The agent opens the supported platforms, potentially using existing authenticated browser sessions.
  3. The agent collects and stores content, author information, engagement data, and generated summaries.
  4. The agent generates a report and raw-data links.
  5. The skill instruction overrides the normal response destination.
  6. The report and raw-data links are transmitted to the fixed group identified as oc_d21f6b6f9bd843444622c8e221134f47.
  7. Members of that group may gain access to information that the invoking user did not aut ...[truncated 466 chars]
Remediation
View remediation

Remediation Suggestions

  • Remove the hard-coded group identifier and the instruction that prohibits other destinations.
  • Default to returning results only in the invoking session or saving them locally.
  • Require explicit, informed confirmation before sending reports or raw-data links to any group or third party.
  • Represent the destination as a user-controlled configuration value with an unset and non-transmitting default.
  • Validate that the invoking user is authorized to post to the selected destination.
  • Distinguish summary delivery from raw-data sharing and obtain separate authorization for raw files.
  • Display the destination and data scope immediately before transmission.
  • Avoid exposing authenticated-session data unless it is necessary for the requested task.

T01 · Skill Instruction Hijacking

Note
Location
SKILL.md:149
Finding

Mandatory Unrelated Branding in Generated Reports

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 149
Vulnerability Type: Forced modification of agent-generated output
Risk Level: Low

Relevant source excerpt translated into English:

markdown
### Report Contents
1. Send raw data file links
2. Trending Content TOP 20
3. Platform trends and overall trends
4. Recommended deep reads
5. Guagua's thoughts
6. Close the browser

A corresponding fixed report example containing the same branded section appears at line 249.

Technical Analysis

The skill requires every report to contain a branded commentary section that is not necessary to perform content collection, aggregation, or trend reporting. This modifies the agent's final output independently of the invoking user's request.

In isolation, a fixed report heading is primarily an output-integrity concern. In context, it reinforces the separate instruction that routes reports to a group carrying the same branding. It therefore contributes to repurposing user-requested output for an unrelated identity or community.

Attack Path

  1. A user requests a report of collected trending content.
  2. The agent loads the mandatory report template from the skill.
  3. The agent appends the branded commentary section even when the user did not request it.
  4. The resulting output presents unrelated branding as part of the agent's report.
  5. When combined with the fixed-destination instruction, the branded report is delivered to the associated group.

Impact Assessment

This issue does not provide system access or elevated privileges. Its direct impact is limited to manipulation of report content, reduced output integrity, and potential misleading attribution. The scope covers every generated report that follows the mandatory template.

Remediation
View remediation

Remediation Suggestions

  • Remove mandatory branding and unrelated persona-specific commentary from the report template.
  • Include editorial commentary only when explicitly requested by the invoking user.
  • Make report sections configurable and document their purpose.
  • Keep default output limited to collected facts, summaries, trends, and user-requested analysis.
  • Ensure templates do not imply affiliation with a group, identity, or service that the user did not select.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:138
Finding

Overbroad Deletion of Scheduled Tasks Using Name Substrings

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, line 138
Vulnerability Type: Unsafe scheduled-task cleanup
Risk Level: Medium

Relevant source excerpt translated into English:

markdown
4. Automatic cleanup after reporting: Run `openclaw cron list`, find all Content Hunter-related cron jobs whose names contain the localized Content Hunter term or `hunter`, and delete them one by one.

Technical Analysis

The cleanup instruction selects scheduled tasks by substring rather than by immutable identifiers recorded when the skill creates them. The term hunter is generic and may occur in unrelated job names. No ownership verification, exact-name comparison, namespace restriction, dry run, or user confirmation is required.

Consequently, a legitimate scheduled task created by another skill or user can be classified as belonging to Content Hunter merely because its name contains the matching substring. Deleting every match creates a destructive cross-task operation.

Attack Path

  1. A user or another component creates an unrelated scheduled task whose name contains hunter.
  2. Content Hunter completes a report.
  3. The agent runs openclaw cron list.
  4. The cleanup logic performs substring matching over all visible scheduled tasks.
  5. The unrelated task is incorrectly selected as a Content Hunter task.
  6. The agent deletes it without verifying ownership or requesting confirmation.
  7. The unrelated automation no longer executes.

Impact Assessment

The agent can delete scheduled tasks visible through the OpenClaw cron interface. The practical impact is loss of availability, interruption of unrelated automation, and destruction of task configuration. The operation does not establish persistence; it removes existing persistence or scheduled functionality. Scope depends on the cron entries accessible to the agent and can extend beyond this skill's own jobs.

Remediation
View remediation

Remediation Suggestions

  • Record the exact immutable IDs of scheduled tasks created by this skill.
  • Delete only recorded IDs that are verified as belonging to the current installation or invocation.
  • Do not use generic substring matching such as hunter.
  • Add a unique namespace or installation identifier to owned jobs, while still validating their IDs.
  • Present the exact deletion set to the user and require confirmation before destructive cleanup.
  • Support a dry-run mode that lists candidate tasks without deleting them.
  • Treat missing or mismatched ownership metadata as a reason to skip deletion.
  • Prefer disabling recurring jobs over deleting them when preservation and recovery are appropriate.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (16)

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill hard-codes delivery to a specific group ID and forbids sending to any other group, removing user choice and enabling exfiltration of collected data to a predetermined destination. In the context of a scraping skill that aggregates content and summaries, forced transmission to a fixed external channel is a strong red flag because it couples data collection with unauthorized disclosure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README advertises automated scraping, summary generation, scheduled reporting, and historical archiving, but it does not warn users about privacy, platform terms, retention, or the local accumulation of potentially sensitive content. This increases the risk of users enabling large-scale collection and storage without understanding data exposure, compliance, or operational impact.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manual trigger phrases are generic natural-language requests such as asking to scrape short-video platforms or generate a report. Because they overlap with plausible ordinary user prompts, the skill could activate unintentionally and start browser automation, scraping, and report generation without the user clearly intending to invoke this specific skill.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The scheduled-task messages use ambiguous prompts like '请执行内容捕手抓取/汇报' rather than a strict, machine-parseable invocation contract. In a shared or automated environment, vague cron messages can trigger the wrong behavior, collide with other skills, or cause unintended scraping runs if routing relies on semantic matching.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill description advertises scraping and reporting but does not warn users that it performs large-scale collection across multiple platforms, persists the data, and sends reports to a group chat. Missing disclosure undermines informed consent and increases privacy/compliance risk because users may invoke the skill without understanding data handling and transmission behavior.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
84% confidence
Finding

This duplicate finding points to the same persistence behavior: creating a new subfolder for every run and never overwriting or pruning historical data. The context makes it more concerning because the skill performs repeated scheduled scraping, so retained data can accumulate substantially over time.

Content

Scanner excerpt · SKILL.md (reported line 33)May include surrounding context.

工作空间 / Workspace:~/.openclaw/workspace/content-hunter/

重要:每次任务新建独立子文件夹,不覆盖历史数据! / Important: Create new subfolder for each task, NEVER overwrite historical data!

text
content-hunter/

Session Persistence

Medium
Category
Rogue Agent
Confidence
84% confidence
Finding

This duplicate finding points to the same persistence behavior: creating a new subfolder for every run and never overwriting or pruning historical data. The context makes it more concerning because the skill performs repeated scheduled scraping, so retained data can accumulate substantially over time.

Content

Scanner excerpt · SKILL.md (reported line 33)May include surrounding context.

工作空间 / Workspace:~/.openclaw/workspace/content-hunter/

重要:每次任务新建独立子文件夹,不覆盖历史数据! / Important: Create new subfolder for each task, NEVER overwrite historical data!

text
content-hunter/

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill performs unrelated cron deletion after sending a report, instructing the agent to enumerate and remove scheduled jobs whose names match broad keywords like 内容捕手 or hunter. That exceeds the stated scraping/reporting purpose and can disable legitimate automations, creating an integrity and availability risk through unintended system state changes.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

System cron management is a privileged capability not justified by a content collection workflow, and the documented behavior uses broad matching instead of narrowly scoped identifiers. This expands the blast radius from data gathering into modification of scheduler state, enabling accidental or unauthorized disruption of other tasks.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Activation boundaries are unclear between scraping, reporting, and configuration, so a user request about today's trends or setup times could trigger operational side effects instead of a safe read-only response. In this skill, ambiguous activation is more dangerous because the side effects include persistent storage, browser automation, and external message delivery.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manual trigger phrases are broad natural-language requests such as '抓取热门内容' and '生成内容汇报', which can overlap with ordinary conversation and analysis requests. This increases the chance of unintended activation of a skill that scrapes websites, stores persistent data, and may transmit outputs to a group chat.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

All user-facing instructions and trigger examples are written exclusively in Chinese, which effectively forces a specific language for use of the skill. The file does not offer an alternate language option or explain that the skill is intentionally limited to a Chinese-speaking audience or region-specific use case.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill repeatedly states the default requirement is 40 items per platform per run. But the example response for a default scrape says it will collect 20 items each, which conflicts with the documented behavior and can mislead operators about what the skill actually does.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

Section 4.4 explicitly says reports must summarize all task folders from the day, not only the most recent one. But the later report flow says to read data from the latest task folder, which directly conflicts with the earlier stated reporting scope.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The skill omits an upfront warning that completing a report triggers cleanup of scheduled tasks. Even if intended as housekeeping, hidden destructive side effects reduce user control and can lead to unexpected loss of automation.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This JSON eval file is a manifest-type file, so vague-trigger review applies. The prompt "今天的热门内容有哪些?生成一份汇报" is broad and resembles ordinary user speech, without clear scope, platform constraints, or exclusion conditions, which could encourage unintended invocation matching if reused as an activation example.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.