Back to skill

Security audit

news_scraper

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a user-directed Chinese news scraping and reporting tool, but its install and model dependency provenance are under-scoped and its advertised summarization capabilities are overstated.

Review the install path before use. Prefer running the reviewed local scripts in an isolated virtual environment, pin and verify Python dependencies, avoid sensitive search topics because requests go to third-party services, choose output paths carefully, and pin or locally vendor the optional summarization model before using abstractive mode.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:166
Finding

Unpinned and Unverifiable Python Dependency Installation

Content
View full analysis

Vulnerability Details

File Locations:

  • SKILL.md:166-176
  • README.md:40-57

Vulnerability Type: Unpinned third-party dependencies and unverifiable package installation
Risk Level: Medium

Vulnerable Code

SKILL.md:166-176:

bash
pip install requests beautifulsoup4
bash
pip install jieba
bash
pip install transformers torch

README.md:40-57:

bash
# Basic installation
pip install news-hot-scraper

# Install abstractive summarization dependencies
pip install news-hot-scraper[abstractive]
bash
# Clone the repository
git clone https://clawhub.com/workbuddy/news-hot-scraper.git
cd news-hot-scraper

# Install basic dependencies
pip install -r requirements.txt

# Install optional abstractive summarization dependencies
pip install transformers torch

Technical Analysis

The installation instructions retrieve packages from external registries without pinning exact versions or verifying artifact hashes. Consequently, the code installed by users can change after the Skill has been reviewed.

The instructions also tell users to install a package named news-hot-scraper, but the audited directory does not contain packaging metadata that establishes a verifiable relationship between that registry package and the reviewed source. The README additionally references a requirements.txt file that is absent from the audited project.

Python packages can execute code during installation, import, or normal use. A compromised package release, dependency-confusion event, package-name takeover, or malicious transitive dependency could therefore introduce arbitrary code into the user's environment.

Attack Path

  1. An attacker compromises one of the named packages, its publishing account, or a transitive dependency.
  2. Alternatively, an attacker publishes or takes control of the unverified news-hot-scraper package name.
  3. The user fo ...[truncated 934 chars]
Remediation
View remediation

Remediation Suggestions

  1. Add a reviewed dependency manifest to the repository and pin every direct dependency to an exact version.
  2. Generate and commit a lock file that also constrains transitive dependencies.
  3. Record cryptographic hashes and install with hash verification, such as:
    bash
    pip install --require-hashes -r requirements.lock
    
  4. Verify ownership and provenance of the news-hot-scraper registry package before recommending it.
  5. Add packaging metadata to the audited repository if registry installation is intended.
  6. Remove the reference to the nonexistent requirements.txt, or add the reviewed file to the project.
  7. Run dependency vulnerability and provenance checks in CI.
  8. Recommend installation in an isolated virtual environment without administrative privileges.

T08 · Insecure Dependencies

Warning
Location
scripts/news_summarizer.py:111
Finding

Runtime Retrieval of an Unpinned Remote Machine-Learning Model

Content
View full analysis

Vulnerability Details

File Locations:

  • scripts/news_summarizer.py:111-116
  • references/summarization_methods.md:198-212
  • references/summarization_methods.md:245-247

Vulnerability Type: Mutable remote model dependency without revision or integrity pinning
Risk Level: Medium

Vulnerable Code

scripts/news_summarizer.py:111-116:

python
if self.abstractive_pipeline is None:
    print("正在加载摘要模型(首次加载需要时间)...")
    # 使用中文摘要模型
    self.abstractive_pipeline = pipeline(
        "summarization",
        model="google/mt5-small-chinese",
        device=-1  # 使用 CPU
    )

references/summarization_methods.md:198-212:

python
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

def load_summarizer(model_name):
    """
    加载摘要模型和分词器

    Args:
        model_name: 模型名称或路径

    Returns:
        tokenizer, model
    """
    tokenizer = AutoTokenizer.from_pretrained(model_name)
    model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
    return tokenizer, model

references/summarization_methods.md:245-247:

python
tokenizer, model = load_summarizer("google/mt5-small-chinese")
text = "这是一段长新闻文本..."
summary = generate_summary(tokenizer, model, text)

Technical Analysis

Abstractive summarization initializes a Transformers pipeline using only the mutable repository identifier google/mt5-small-chinese. No immutable commit revision or expected artifact hash is supplied.

On first use, the Transformers ecosystem may download model configuration, tokenizer data, and model artifacts from a remote model repository. Because the repository revision is not pinned, the effective component used at runtime can change after this Skill has been audited.

This creates a supply-chain trust boundary. A compromised model repository, changed default branch, or tampered distribution path could cause users to retrieve unexpected artifacts. Even ...[truncated 1820 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin the model to a reviewed immutable commit:
    python
    self.abstractive_pipeline = pipeline(
        "summarization",
        model="google/mt5-small-chinese",
        revision="REVIEWED_COMMIT_HASH",
        device=-1,
        trust_remote_code=False
    )
    
  2. Download and review the model artifacts in advance, then load them from a controlled local directory.
  3. Record and verify cryptographic hashes for all model, tokenizer, and configuration artifacts.
  4. Prefer safe serialization formats such as safetensors where supported.
  5. Explicitly keep remote custom code disabled.
  6. Pin and regularly patch the Transformers, tokenization, and tensor-runtime dependencies.
  7. Perform model loading in a restricted environment with minimal filesystem permissions and unnecessary network access disabled.
  8. Update the examples in references/summarization_methods.md to use the same immutable revision and integrity controls.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The code largely matches the scraping portion of the description: it retrieves trending/news items from several Chinese platforms using both an API-based method and direct scraping, and stores source URLs. However, key declared capabilities are absent. There is no summarization logic at all—the 'summary' field is always left empty, and there are no extractive or abstractive summarization methods or libraries. Likewise, 'cite sources' is only minimally reflected by storing URLs, not by generating citations. Also, direct scraping is not implemented for all declared platforms: 抖音 and 今日头条 are available only through the aggregator API, not via direct scraping. Therefore the description materially overstates the skill's implemented functionality.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared purpose emphasizes end-to-end hot-topic scraping from multiple Chinese platforms, supporting both API-based and direct scraping methods, then summarizing and citing sources. This code chunk implements only the summarization/reporting portion over preexisting JSON input. There are no HTTP requests, platform clients, browser automation, scraping logic, or API integrations. Additionally, although helper methods for extractive and abstractive summarization exist, the actual batch workflow (generate_summary_for_news) summarizes only the title field, with extractive mode simply taking the first 100 characters of the title rather than using the extractive sentence-scoring logic over full text. The Markdown report includes source URLs if already present, but source collection/citation generation is not implemented. Therefore the description materially overstates the code’s actual capabilities and primary purpose.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language instructions, examples, and operational guidance are presented almost entirely in Chinese, which effectively forces a specific language for users. The file does not offer an opt-in language choice or explain that the skill is intentionally limited to Chinese-speaking or China-specific users.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill describes network access, reading reference files, and writing JSON/Markdown outputs, but it does not declare any explicit tool scope or permission boundaries. In an agent environment this can lead to overbroad runtime authority, making it easier for the skill to access files or perform network operations beyond what users reasonably expect.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The description explicitly limits the skill to Chinese platforms and Chinese summarization tooling, and the document is primarily written in Chinese, but it does not state that this locale restriction is intentional or offer alternatives for users needing other languages. Under the policy, forcing a specific language/locale without opt-in can be a violation unless clearly justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill encourages sending user-supplied keywords and requests to third-party APIs and target sites without clearly warning that those queries and associated request metadata will leave the local environment. This can expose sensitive research topics, internal terms, or operator IP/user-agent metadata to external services, which is especially relevant for an agent skill performing networked collection.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

This markdown file includes examples that load pretrained models and process user-provided news text, and later recommends caching raw text inputs for summarization. The document does not warn that using external pretrained model sources or caching submitted text may affect user privacy or data handling, which is relevant under the markdown-specific missing-warning criterion.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The file's user-facing docstrings, help text, and status/error messages are all written in Chinese, which imposes a specific language on users without any opt-in or alternative locale. Under the stated policy, forcing a language without user choice is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest says the skill scrapes hot news topics, generates summaries, and cites sources, including extractive and abstractive summarization. In this file, every scraped item leaves 'summary' as an empty string and the code only saves raw title/URL metadata to disk, so the implemented behavior falls materially short of the declared functionality.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest explicitly includes 抖音 and 今日头条 among supported Chinese platforms. However, the direct scraping dispatcher only handles weibo, zhihu, bilibili, tencent, and thepaper, and reports unsupported for anything else, so part of the claimed platform coverage is missing in actual code.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This is a natural-language policy concern because the skill forces a specific language for its description, usage instructions, CLI help text, and runtime messages. There is no opt-in, fallback, or documentation that this tool is intentionally limited to a Chinese-speaking or region-specific context.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The summarize_extractive docstring describes extracting key sentences from the source text, but generate_summary_for_news does not call this method for extractive mode. Instead, extractive mode at L166-L168 simply uses title[:100], so the documented summarization behavior is not what the tool actually performs for news items.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest says the skill scrapes hot news topics, generates summaries, and cites sources. In this file, the main summarization path ignores article text/content and in extractive mode just assigns title[:100]; even abstractive mode summarizes the title only. That is materially narrower than genuine news summarization and may mislead users about what the skill actually does.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The method documentation says it generates summaries for news data, and the inline comment at L170 says abstractive summaries are generated 'based on the title'. Across both modes, the code never summarizes article content/body text and instead operates solely on title. This contradicts the module-level description of a news summarization script supporting extractive and abstractive summarization in the usual sense of summarizing text content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

This markdown file documents use of a third-party aggregation API and direct web scraping, and later examples show writing results to JSON and Markdown files. While the README includes legal/compliance notes, it does not clearly warn users that running the skill will transmit requests to external services and create local output files that may contain collected data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill states that results are written to JSON/Markdown files but does not warn users that scraped content and generated summaries may persist on local storage. This can create unintended data retention, particularly if collected topics, article text, or summaries contain sensitive or regulated information.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

At L212, the example sets Accept-Language to zh-CN,zh;q=0.9,en;q=0.8, which biases requests toward a specific language/locale. The file does not explain that this is optional, configurable, or justified by a region-specific use case, so it can be read as forcing a locale without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

All instructional content in the file is presented in Chinese, and the document does not indicate that Chinese is optional or that the content is intentionally limited to a Chinese-only audience. Under the language/locale policy rule, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

This code performs a file write operation using a path supplied by --output, but the only disclosure is a success message after the write completes. There is no prior warning, confirmation, or comment/docstring cautioning that an existing file may be overwritten.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.