Back to skill

Security audit

Ai Research Scraper

Security checks for vulnerabilities and agentic risk

Overview

This AI news scraper is not clearly malicious, but it needs review because it runs an unverified external search script and includes under-disclosed translation code.

Review before installing. Confirm the external tavily-search skill is trusted, pinned, and protected from modification, and run the scraper with limited filesystem and environment access. Remove or disable the unused translation scripts unless you specifically want text sent to those translation providers, and avoid passing sensitive prompts, private notes, or proprietary content through this skill.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T08 · Insecure Dependencies

Error
Location
scripts/scraper.py:40
Finding

Execution of an Unpinned External Skill Dependency

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Error
Location
scripts/final_scraper.py:12
Finding

Alternate Scraper Executes an Unverified External Skill

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Error
Location
scripts/simple_scraper.py:10
Finding

Legacy Scraper Executes an Unpinned External Search Component

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (51)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The finding indicates third-party translation API use with signed requests and output of translation results, which is unrelated to the claimed AI research aggregation role. This is dangerous because authenticated external requests may involve secrets, user text disclosure, and a materially different trust model than a passive scraper would imply.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes a skill focused on collecting recent AI research information from known AI websites and providing concise summaries with links. This file instead implements a general-purpose Baidu translation client that sends arbitrary text to an external translation API, which is not an obvious or declared requirement of a research scraping skill.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes a skill for fetching and summarizing recent AI research information from well-known AI websites with concise links. This code instead sends user-provided text to Google's translation API and returns translated text, which is a different end-user capability and not an implementation detail of research scraping.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes a skill for scraping and summarizing recent AI research information from well-known AI websites, but this file implements a Youdao translation API test that translates a fixed sentence to Chinese. Translation functionality is materially different from research scraping, summarization, and link aggregation, so this behavior is outside the described skill purpose.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill advertises and demonstrates shell and network-capable execution paths but does not declare any explicit tool scope or permissions. In an agent ecosystem this weakens least-privilege boundaries, making it easier for the skill to access network and command execution capabilities without transparent review or policy enforcement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description and overview are written entirely in Chinese, and the file does not indicate that the skill is region-specific or that users may choose another language. This creates a natural-language locale constraint without opt-in, which matches the policy-violation category.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest describes a skill focused on collecting recent AI research information from well-known AI websites and providing concise summaries with links. This reference file documents a translation API and testing flow, which is a distinct capability not implied by the stated purpose of news/research scraping and summarization.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The documentation claims use of Google Cloud Translation API but provides sample code for the unofficial googletrans library, which routes data differently and may rely on undocumented upstream behavior. This mismatch can cause operators to make incorrect trust, compliance, and data-flow assumptions, potentially sending content through an unintended third-party path.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The translation section describes sending text to an external translation service without warning that user or scraped content may be transmitted off-system. In a research-scraping skill, collected text could include proprietary prompts, internal notes, or licensed content, so silent external transmission creates privacy, compliance, and data-governance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The Tavily API documentation instructs external search requests and states that failures trigger fallback web scraping, but it does not warn users that queries and possibly subsequent browsing activity will be sent to third parties. This is risky because the skill's context is live research collection, where search terms, targets, and fallback behavior can leak user intent and trigger unreviewed outbound network access.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

The documented request to https://api.tavily.com/search is an explicit external transmission path. In this skill context, outbound search requests are expected, but they still pose security and privacy risk if users are not informed that queries, API credentials, and derived research interests may be transmitted to a third party.

Content

Scanner excerpt · references/api_reference.md (reported line 54)May include surrounding context.

python
import requests

response = requests.get('https://api.tavily.com/search', params={
    'query': 'AI product development news',
    'api_key': 'TAVILY_API_KEY',
    'max_results': 10,

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The translate method defaults to_lang to "zh", and the example usage translates text to Chinese automatically. This imposes a specific language choice by default without indicating user selection or opt-in, which is a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes a skill for scraping and summarizing recent AI research and product updates from well-known AI websites. This file instead implements a generic Baidu text translation client that sends arbitrary text to an external translation API, which is a different end-user capability from research collection or summarization.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The code exposes a standalone capability to transmit arbitrary text to Baidu's translation API. The manifest only justifies fetching and summarizing AI research information; it does not mention translation or third-party language-processing services as part of the skill's purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code sends arbitrary input text to Baidu's external translation API, which can expose scraped content, prompts, or user-provided data to a third party without any visible disclosure or consent mechanism. In a research-scraping skill, collected content may include sensitive or proprietary material, so silent exfiltration to an external service creates a real privacy and compliance risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script presents its banner text in Chinese (抓取AI领域最新研究信息) and later emits Chinese-only labels such as 标题, 摘要, and 来源. This imposes a specific language on users without any opt-in, fallback, or documented region-specific justification, which matches the locale-policy violation criteria.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.