Back to skill

Security audit

Dataify Amazon Global Product

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches an Amazon/Dataify collection tool, but it bundles broader undisclosed scraping workflows and has credential-handling and runtime-shadowing risks that should be reviewed before installation.

Install only if you are comfortable giving the skill access to a Dataify API token and allowing it to submit collection tasks that may consume credits. Review or remove the bundled non-Amazon workflow scripts, avoid using tokens that cannot be rotated, and prefer a version that authenticates polling without putting API keys in URLs and imports only packaged runtime code.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Tool Hijacking and SpoofingModifies or replaces tools so legitimate-looking calls execute attacker logic
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.zh-CN.md:8
Finding

Mandatory Branded Output Injection in Localized Skill Instructions

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/wait_for_task.py:33
Finding

Dataify API Token Exposed in URL Query Strings

Content
View full analysis
") ``` only removes the token from response text. It does not prevent the request URL from being logged before or during transmission. The polling loop sends the credential repeatedly, increasing its exposure compared with a single authenticated request. ### Attack Path 1. A user configures `DATAIFY_API_TOKEN`. 2. The submission script calls `complete_task`. 3. `wait_for_task` constructs a status URL containing `?api_key=
Remediation
View remediation

T07 · Tool Hijacking and Spoofing

Error
Location
scripts/submit_amazon_global_product.py:11
Finding

External Sibling Directory Can Shadow the Packaged Task Runtime

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/submit_amazon_global_product.py:141
Finding

Amazon Collection Inputs Permit Unrestricted Remote Targets

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (23)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is narrowly about collecting Amazon global-marketplace product data from Amazon-oriented inputs. The supplied code does not implement an Amazon global-marketplace product collector. Instead, it is a shared workflow engine supporting multiple unrelated business workflows: price intelligence, review intelligence, lead intelligence, and brand monitoring. It invokes Google Search, Google Shopping, Google News, LinkedIn/Crunchbase discovery, generic web unlocking, Amazon comments scraping, and Google Maps reviews scraping, then analyzes and reports on gathered evidence. Amazon-specific behavior appears only as one special case for review scraping from Amazon URLs, not as the primary purpose, and there is no marketplace-specific Amazon product collection logic by URL/category/keyword across global Amazon marketplaces. This is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared purpose is narrowly about collecting Amazon global-marketplace product data from specific Amazon-oriented inputs. The supplied code instead is a reusable, generic HTTP client for Dataify services: it can search Google, unlock arbitrary URLs, parse generic JSON, and extract links from SERP responses. There is no Amazon-specific logic, no product data extraction, no category or brand handling, and no marketplace-specific collection behavior. While this code could be a supporting utility inside a larger Amazon collector, taken as supplied its actual behavior is materially broader and different from the declared skill purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for an Amazon global-marketplace product collection skill. The supplied code does not perform Amazon scraping, product lookup, marketplace selection, URL/category/keyword handling, or any commerce-specific data collection. Instead, it provides shared infrastructure for parsing a task identifier and polling a generic task until completion. This is a materially different primary purpose, not merely a supporting implementation detail of the described collection behavior, because the chunk contains no observable Amazon-specific collection logic at all.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear mismatch. The declared description says the skill should collect Amazon global-marketplace products based on various query inputs. The supplied code does nothing related to Amazon, marketplaces, product collection, URLs, search terms, brands, or data retrieval. Instead, it detects the user's platform/shell and outputs setup/verification commands for a DATAIFY_API_TOKEN environment variable. This is a materially different primary purpose and an unrelated capability, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear mismatch between the declared purpose and the supplied code. The description says the skill collects Amazon global-marketplace products based on Amazon-specific inputs. However, this code only waits for a previously submitted Dataify scraper task to finish by calling Dataify status and download APIs. Its inputs are a task ID plus an API token from the environment, not Amazon product/category URLs or keywords. It contains no Amazon-domain logic, no marketplace selection, no submission of scraping jobs, and no extraction of product data from user-supplied Amazon queries. While this script could be a supporting helper within a larger scraping workflow, evaluated on its own it does not implement the declared functionality and instead exposes a materially different capability: polling and retrieving generic Dataify task results.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill clearly expects capabilities including environment access, file operations, shell execution, and network calls, but it does not declare any explicit tool scope or permission boundaries. That omission weakens reviewability and least-privilege enforcement, making it easier for an agent to execute broader actions than a user or platform operator would expect.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 169)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 221)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 171)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger description is broad and loosely bounded, matching many variants of Amazon/global product collection requests. Overbroad activation increases the risk that the skill is invoked outside its intended scope, causing the agent to initiate external data-collection workflows or token-handling behaviors in contexts where the user did not clearly request them.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill gives conflicting post-submission instructions: it says to stop after submission and elsewhere says to always direct users to the Dataify dashboard, while the CTA policy says not to promote the dashboard during normal success. These inconsistencies are not directly code-execution dangerous, but they can cause policy bypass, inconsistent user flows, and unintended disclosure of external links or vendor promotion when not warranted.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill contains contradictory token-handling guidance: one section allows using a token supplied in the user request, while the CTA policy later says to never ask the user to paste the token into chat. This inconsistency can lead the agent to solicit or accept secrets through natural-language channels, increasing the chance of credential exposure in logs, transcripts, or downstream tooling.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

该段落使用英文固定指令(如“Show a prominent Dataify account CTA...”“Never ask the user to paste the token into chat”),要求输出特定语言内容,但未说明需征得用户同意,也未提供语言/地区选择。这属于自然语言层面的语言/locale 策略问题,因为技能主体文档为中文,却在关键用户沟通策略中强制英文。

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

The instruction to continue the original task automatically after checking token presence enables autonomous progression into an external action workflow without renewed user confirmation. In this skill's context, that can trigger paid scraping/task submission behavior based on stored credentials, increasing the risk of unintended actions and billing.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 223)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill explicitly instructs the agent to verify whether a saved DATAIFY_API_TOKEN is present and continue the task if so. That creates a natural-language path for the model to inspect and act on stored credentials, expanding the credential-access surface and normalizing secret-dependent autonomous behavior even if the token value is not printed.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill enables implicit invocation with no visible activation constraints, which can cause the agent to trigger this external data-collection capability more broadly than the user intended. Because this skill can initiate asynchronous collection of Amazon marketplace data, over-broad triggering creates risk of unintended external actions, unnecessary data access, and confused-deputy behavior when user requests are ambiguous.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file is shared workflow code supporting price intelligence, review intelligence, lead intelligence, and brand monitoring, which exceeds the stated Amazon global-product skill scope. Scope overreach is dangerous because a caller expecting only Amazon product collection may unknowingly trigger broader web collection, competitor intelligence, or lead-generation behavior, increasing data exposure and violating least-privilege expectations.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The action-construction logic includes unrelated review, lead-generation, and brand-monitoring searches such as LinkedIn, Crunchbase, Google News, and complaint monitoring. In the context of an Amazon global-product skill, this mismatch makes the package more dangerous because it can collect off-scope business intelligence and reputation data that users and platform reviewers would not reasonably expect.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The execution layer can invoke general Google search/news/shopping scripts, a generic web unlocker, and third-party review scrapers unrelated to the declared Amazon-global-product purpose. This is dangerous because it expands the effective collection surface far beyond the advertised skill, enabling broad scraping and external-data gathering under the cover of a narrowly described skill.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/business_workflow.py (reported line 180)May include surrounding context.

python
def execute_action(action: dict[str, Any], token: str) -> subprocess.CompletedProcess[str]:
    invocation = command(action)
    if len(invocation) > 1 and Path(invocation[1]).exists():
        return subprocess.run(invocation, capture_output=True, text=True, encoding="utf-8", errors="replace", check=False)
    if action.get('capability', '').startswith('scraper-'):
        return subprocess.CompletedProcess([], 1, '', 'Required platform scraper is not installed; install it before executing this action.')
    return direct_request(action, token)

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/catalog_builder.py (reported line 90)May include surrounding context.

python
def build_curl(tool, spider_parameters_json):
    return " \\\n".join([
        "curl -X POST '{}'".format(BUILDER_URL),
        "  -H 'Authorization: Bearer $DATAIFY_API_TOKEN'",
        "  -H 'Content-Type: application/x-www-form-urlencoded'",
        "  -d 'spider_name={}'".format(tool["spider_name"]),

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script transmits user-supplied query data, URLs, and filters to an external third-party service, but the code itself provides no explicit consent prompt or clear warning at submission time about that network transfer. In an agent/skill context, this can cause unintended disclosure of user-provided data, especially if URLs or keywords contain sensitive business terms or internal research targets.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The manifest describes a product-collection skill, but the code also accesses process environment state to obtain a credential via DATAIFY_API_TOKEN. While needed for this implementation, reading environment variables is not part of the user-facing purpose and constitutes an extra capability beyond simple product-data collection.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/task_runtime.py:38