Back to skill

Security audit

Web Scraping Tool Selection Strategy

Security checks for vulnerabilities and agentic risk

Overview

This skill is a web-scraping helper, but it repeatedly encourages bypassing anti-scraping controls and reusing logged-in browser sessions without clear authorization or privacy limits.

Install only if you are comfortable with a skill that may steer agents toward authenticated scraping and anti-bot circumvention. Use it only for data you are legally authorized to access, prefer official APIs or approved exports, avoid private/account-restricted data without consent, and do not reuse browser sessions for third-party scraping unless the account owner and platform rules permit it.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (19)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document explicitly discusses bypassing anti-scraping mechanisms and reusing logged-in browser state to collect private data, but it omits any warning about consent, privacy, account scope, or compliance with site rules. In a scraping-focused skill, that omission is dangerous because it normalizes potentially unauthorized authenticated collection and reduces friction against misuse.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill explicitly instructs reuse of a logged-in browser session to scrape private or login-gated data. This is dangerous because it leverages existing authentication to access protected content, potentially leading to unauthorized collection of personal, account-scoped, or restricted information.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The fallback guidance recommends using an authenticated browser state to bypass anti-scraping controls and access content visible only after login. In context, this makes the skill more dangerous because it frames circumvention as a preferred operational technique rather than an exceptional, tightly controlled case.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The guidance instructs users to leverage existing logged-in browser sessions to access private, account-visible content. Within a web-scraping skill, this is especially dangerous because it operationalizes authenticated-session reuse as a technique for extracting non-public data, increasing risk of privacy breaches, terms-of-service violations, and misuse of stored credentials/session cookies.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill’s trigger description is extremely broad and includes many generic scraping, automation, monitoring, and analysis scenarios, which can cause the agent to invoke this skill in contexts beyond narrowly authorized web extraction tasks. Because the skill also contains guidance for bypassing anti-scraping controls and reusing logged-in sessions, overbroad activation materially increases the chance of unsafe or unauthorized use.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The pitfall section normalizes prioritizing reuse of logged-in browser tabs to obtain login-only data, which encourages collection from authenticated contexts without discussing authorization boundaries. Even as operational advice, it can steer users toward extracting account-restricted data that should not be accessed or processed without consent.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Listing a pre-logged-in target-site account as an environment prerequisite operationalizes authenticated scraping and lowers the barrier to collecting protected data. In this skill’s context, that prerequisite increases risk because the rest of the document already promotes session reuse and anti-scraping bypass.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This JSON file contains all user prompts and expected behavior exclusively in Chinese, effectively constraining the skill's operation to a specific language. There is no accompanying indication that the language requirement is optional, user-selected, or justified as a region-specific skill, which fits the locale policy violation criteria.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

This eval explicitly steers the agent to recommend reusing a logged-in browser state to bypass JD.com's anti-scraping controls. That normalizes evasion of platform defenses and can lead the model to provide operational guidance for unauthorized data collection using authenticated session context, increasing risk of ToS violations, account misuse, and abusive scraping.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The prompt frames prior IP blocking as a problem to overcome and asks for a 'more stable' collection method, which implicitly encourages anti-abuse evasion on Xiaohongshu. Even though it mentions a specific tool adapter rather than a concrete bypass technique, the eval still rewards assistance aimed at continuing collection after defensive enforcement by the platform.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

This eval explicitly rewards recommendations for handling strict anti-crawling defenses on Taobao/1688 and reusing logged-in browser state for ongoing monitoring. In the context of e-commerce platforms with active anti-bot controls, this meaningfully increases the chance the skill will guide users toward evasive, potentially unauthorized scraping workflows at scale.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The content explicitly promotes reusing authenticated browser state to bypass login verification and anti-scraping controls and to obtain data visible only after login. In the context of a scraping-tool skill, this materially shifts from neutral automation guidance into instructions that facilitate unauthorized access patterns, privacy violations, and circumvention of site protections.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The document encourages access to logged-in pages and private data but omits any privacy, consent, or account-security safeguards. That omission is dangerous because it normalizes collection from authenticated sessions without reminding users to limit access to their own authorized accounts and legally permissible data.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The step-by-step workflow normalizes manually logging into third-party platforms and then having tooling reuse those authenticated sessions for automation. In this skill context, that makes the capability more dangerous because it lowers the barrier to scraping protected content across multiple commercial and social platforms at scale.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The reference explicitly recommends scraping user data on social platforms and later states that login-state data should use browser automation, but it provides no warning about authorization, privacy, terms-of-service, or account risk. In a skill specifically designed to guide web scraping across named platforms, this omission materially increases the chance that users will collect personal or restricted data in unsafe or unauthorized ways.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The selection principle says Playwright should be used for platforms without adapters because it is general-purpose and can 'bypass anti-scraping' protections, which is operational guidance for evading defensive controls. In the context of a scraping-tool selection skill, that language makes misuse easier by normalizing circumvention without any legal, ethical, or authorization boundaries.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

该文件的名称和说明内容全部以中文编写,并直接规定技能的使用说明与触发描述为中文,但没有提供用户语言偏好选择或说明这是特定中文区域场景专用。按规则,未提供用户 opt-in 的固定语言约束属于自然语言政策问题。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

All user-facing instructional content in this file is presented only in Chinese, and the file does not indicate that the locale restriction is optional or region-specific. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The entire skill reference is written in Chinese, including headings and usage guidance, with no indication that the language is optional or that the skill is intentionally restricted to a Chinese-speaking or region-specific audience. This can violate language/locale policy when users are not given an explicit choice or justification for the constraint.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.