Back to skill

Security audit

Dataify Amazon Product List

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly performs the advertised Dataify Amazon collection, but it ships broader under-disclosed scraping helpers and has credential and import-handling risks that warrant review before installation.

Review this skill before installing. It needs a Dataify account token and will submit paid or credit-consuming Dataify jobs to collect and retrieve results. The advertised Amazon-list workflow is understandable, but the package contains broader scraping and business-intelligence helper code, and the current token-in-query and external import-path behavior should be fixed or accepted explicitly before use.

Vulnerability Patterns
  • Tool Hijacking and SpoofingModifies or replaces tools so legitimate-looking calls execute attacker logic
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/wait_for_task.py:35
Finding

Long-Lived API Token Exposed in HTTP Query Strings

Content
View full analysis

Vulnerability Details

File Location: scripts/wait_for_task.py, lines 35–37, with credential-bearing calls at lines 105–109 and 130–134
Vulnerability Type: Sensitive credential exposure through request URLs
Risk Level: Medium

Vulnerable Code

python
def request_json(endpoint, params, api_key, timeout):
    url = endpoint + "?" + urllib.parse.urlencode(params)
    request = urllib.request.Request(url, method="GET")

The API token is supplied as an api_key query parameter when polling task status:

python
payload = request_json(
    STATUS_ENDPOINT,
    {"api_key": api_key, "task_id": task_id},
    api_key,
    request_timeout,
)

It is supplied in the same manner when downloading results:

python
return request_json(
    DOWNLOAD_ENDPOINT,
    {"api_key": api_key, "task_id": task_id, "type": "json"},
    api_key,
    request_timeout,
)

Technical Analysis

request_json() serializes all parameters into the URL of a GET request. Consequently, the long-lived DATAIFY_API_TOKEN becomes part of every status and download URL.

HTTPS protects the request from passive network interception while in transit, but it does not prevent the complete URL from being recorded by the Dataify service, reverse proxies, gateways, request tracing systems, application-performance monitoring tools, browser-like network instrumentation, or exception diagnostics. Query strings are commonly retained in access logs.

The response-body redaction performed later in request_json() does not protect the outgoing request URL. The behavior is necessary only to authenticate task operations; placing the credential in the query string is not necessary if the service supports authorization headers or authenticated POST bodies.

Attack Path

  1. A user configures a valid DATAIFY_API_TOKEN and submits a collection task.
  2. The Skill begins polling the status endpoint.
  3. `reques ...[truncated 831 chars]
Remediation
View remediation

Remediation Suggestions

  1. Replace the api_key query parameter with an authorization header:

    python
    def request_json(endpoint, params, api_key, timeout):
        url = endpoint + "?" + urllib.parse.urlencode(params)
        request = urllib.request.Request(
            url,
            headers={"Authorization": "Bearer {}".format(api_key)},
            method="GET",
        )
    
  2. Keep only non-sensitive values such as task_id and result type in the query string.

  3. If the API does not support bearer headers, use an authenticated POST body and ensure request-body logging is disabled or redacted.

  4. Configure server, proxy, and telemetry systems to redact api_key, token, and authorization values.

  5. Avoid including full credential-bearing URLs in errors, diagnostics, or progress output.

  6. Rotate tokens that may already have appeared in request logs.

  7. Prefer narrowly scoped, revocable tokens with account-level rate and spending limits.

T07 · Tool Hijacking and Spoofing

Warning
Location
scripts/submit_amazon_product_list.py:10
Finding

Untrusted External Directory Precedes Bundled Runtime During Import

Content
View full analysis

Vulnerability Details

File Location: scripts/submit_amazon_product_list.py, lines 10–14
Vulnerability Type: Python module search-path hijacking
Risk Level: Medium

Vulnerable Code

python
TASK_RUNTIME_DIR = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..", "dataify-task-operations", "scripts"))
if TASK_RUNTIME_DIR not in sys.path:
    sys.path.insert(0, TASK_RUNTIME_DIR)
from task_runtime import complete_task, extract_task_id

Technical Analysis

The submission script computes a path outside the audited project directory and inserts it at index zero of sys.path. Python therefore searches this external sibling directory before the Skill’s bundled script directory when resolving task_runtime.

The code does not verify that the external directory exists, is controlled by a trusted administrator, has safe permissions, or contains an approved version of task_runtime.py. If an attacker can create or modify the sibling path, a malicious module placed there will be imported instead of the bundled implementation.

Python executes top-level module code immediately during import. A substituted module can therefore execute arbitrary Python code before normal task submission begins. Because the process inherits the user’s environment, malicious import-time code can independently read DATAIFY_API_TOKEN and other environment variables.

This external override is not required for the declared Amazon product-list functionality because the package already contains scripts/task_runtime.py.

Attack Path

  1. An attacker obtains write access to the expected sibling location dataify-task-operations/scripts, or creates it where deployment permissions allow.
  2. The attacker places a malicious task_runtime.py in that directory.
  3. A user invokes the documented submit_amazon_product_list.py command.
  4. The script prepends the attacker-controlled directory to sys.path.
  5. Pytho ...[truncated 949 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove the external path insertion and import the bundled runtime through a package-relative import.

  2. Package the scripts as a Python module, for example:

    python
    from .task_runtime import complete_task, extract_task_id
    
  3. If direct script execution must remain supported, load the bundled file using an explicit path anchored inside the audited Skill directory rather than changing global import precedence.

  4. If an external runtime is operationally required, require an explicit trusted configuration path rather than silently selecting a sibling directory.

  5. Validate that any external runtime path:

    • Resolves to the expected trusted directory.
    • Is not writable by untrusted users.
    • Has trusted ownership and restrictive permissions.
    • Contains a runtime matching an approved cryptographic digest or signature.
  6. Do not insert optional or unverified directories at the beginning of sys.path.

  7. Add a regression test that creates a competing external task_runtime.py and confirms that the bundled trusted module is still loaded.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (28)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is narrowly scoped to collecting Amazon product-list records by keyword and marketplace domain, explicitly excluding single-ASIN/product-detail/review use cases. The supplied code does something materially different: it is a shared workflow engine for several business-intelligence modes. It constructs actions for Google Shopping/Search/News, arbitrary source URLs, LinkedIn/Crunchbase lead discovery, brand/news monitoring, and review analysis. It also includes explicit support for Amazon review scraping (scraper-amazon-comment) and Google Maps review scraping, which directly conflicts with the declared 'do not use for reviews.' Additionally, it adds detail-page actions discovered from search results and processes/report outputs, which goes beyond a simple Amazon list collector. Because the primary purpose, accessed resources, and supported capabilities differ substantially from the declaration, this is a strong mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
89% confidence
Finding

The declared description is narrowly scoped to collecting Amazon product-list records by keyword and marketplace domain. However, the supplied code chunk does not implement Amazon-specific collection logic, keyword/domain handling, or restrictions to product-list searches. Instead, it is a reusable submission wrapper that reads generic tool definitions from a catalog, accepts arbitrary parameters, validates them according to metadata, and submits jobs to a remote builder endpoint for whichever tool_sign is chosen. This is a materially broader and different primary purpose than the declared Amazon-specific scraper behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description is narrowly scoped to collecting Amazon product-list records by keyword and marketplace domain. However, the supplied code does not implement Amazon-specific collection logic, marketplace-domain filtering, product-list extraction, or ASIN/list-result handling. Instead, it is a reusable, generic client for Dataify services: one function queries Google search results, another fetches/unlocks arbitrary webpages, and helper functions parse JSON and extract URLs/content from responses. This is a materially broader and different capability set than the declared Amazon list-record collector, so the description does not accurately represent the code's actual behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared purpose is a domain-specific Amazon product-list collector driven by keyword and marketplace domain. The supplied code does not perform Amazon search, scraping, product-list retrieval, or marketplace-specific handling. Instead, it is generic infrastructure for parsing a task identifier and polling for completion of an already-submitted task via wait_for_task. This is a materially different primary purpose and represents undeclared generic task orchestration capabilities unrelated to the stated Amazon collection behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is about Amazon product list collection by keyword and domain. The supplied code is a standalone utility for reporting how to set the DATAIFY_API_TOKEN environment variable on different platforms/shells and whether it is currently configured. It performs no network access, no Amazon querying, no product list retrieval, and no marketplace-specific logic. This is a clear material mismatch in primary purpose and behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description is for an Amazon product-list collection skill driven by keyword and marketplace domain. The supplied code instead implements a generic Dataify task waiter/downloader utility. Its primary function is to poll /task_status, handle processing/failure states, and fetch JSON from /download once a task succeeds. There is no logic for Amazon, keywords, marketplace domains, list-result scraping, or task submission. This is a materially different purpose rather than a mere supporting detail, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The file is a generic multi-workflow orchestrator supporting price, review, lead, and brand intelligence, which exceeds the declared scope of an Amazon product-list skill. In a least-privilege setting, this scope mismatch is dangerous because it enables unrelated collection behaviors and expands the reachable attack surface, data access patterns, and downstream execution paths beyond what users and reviewers would expect.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill contains review scraping logic, including Amazon comments and Google Maps reviews, which is substantially broader than Amazon product-list collection. This is dangerous because it allows collection of different data types and targets than users expect, increasing the chance of overcollection, policy violations, and unauthorized use of scraper capabilities through a misleadingly scoped skill.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

Including lead-intelligence logic that searches LinkedIn and Crunchbase is unrelated to Amazon product-list collection and materially broadens the skill's operational scope. This is dangerous because a user invoking an Amazon listing skill could trigger external company profiling workflows, creating unexpected data collection, policy, and abuse risks outside the advertised function.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill explicitly instructs use of environment variables, local scripts, network requests, and task polling, but it declares no tool scope or permission boundary. That omission can cause an agent framework to expose broader capabilities than users expect, increasing the chance of unintended shell, file, or network actions during execution.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
76% confidence
Finding

The skill authorizes the agent to verify environment state and continue the original task automatically once a token is detected, without renewed user confirmation. In a skill that can invoke network, shell, and remote paid task submission, this increases the risk of unintended external actions or charges after a setup step, especially if the user only intended to configure credentials.

Content

Scanner excerpt · SKILL.md (reported line 133)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

L003 将触发条件写为“当用户说到或请求 Amazon 产品列表采集工具、Amazon product list collection/tool/scraping,或按 keyword and domain 采集 Amazon product list 时触发”。其中包含宽泛的自然语言表述,如“tool/scraping”与“说到或请求”,但没有明确的触发短语列表、排除条件或负例,容易与普通讨论 Amazon 抓取相关的话语重叠,导致误触发。

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file first instructs that after a successful submission the user should be told to go to Dataify to view results, including '始终引导用户前往' the dashboard. But the later 'Account CTA policy' says that during normal submission and successful completion, the skill should not promote registration or the Dashboard. These instructions actively conflict, creating intent ambiguity for the skill's user-facing behavior.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 131)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 113)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 115)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Price-intelligence and Google Shopping collection go beyond a narrow Amazon product-list collector and permit broader market-intelligence behavior. In context, this mismatch is risky because it weakens least-privilege guarantees and lets the skill reach non-Amazon sources and business-analysis workflows not disclosed by its name and description.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Brand-monitoring and news collection are outside the stated Amazon product-list purpose and indicate the skill can perform broader reputation or monitoring tasks than advertised. Even if not directly exploitable as code execution, this scope expansion undermines trust boundaries and can be abused to gather unrelated intelligence under misleading packaging.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/business_workflow.py (reported line 180)May include surrounding context.

python
def execute_action(action: dict[str, Any], token: str) -> subprocess.CompletedProcess[str]:
    invocation = command(action)
    if len(invocation) > 1 and Path(invocation[1]).exists():
        return subprocess.run(invocation, capture_output=True, text=True, encoding="utf-8", errors="replace", check=False)
    if action.get('capability', '').startswith('scraper-'):
        return subprocess.CompletedProcess([], 1, '', 'Required platform scraper is not installed; install it before executing this action.')
    return direct_request(action, token)

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/catalog_builder.py (reported line 90)May include surrounding context.

python
def build_curl(tool, spider_parameters_json):
    return " \\\n".join([
        "curl -X POST '{}'".format(BUILDER_URL),
        "  -H 'Authorization: Bearer $DATAIFY_API_TOKEN'",
        "  -H 'Content-Type: application/x-www-form-urlencoded'",
        "  -d 'spider_name={}'".format(tool["spider_name"]),

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The search function performs general Google SERP queries rather than limiting discovery to Amazon marketplace product-list sources. In a skill advertised as Amazon product-list collection, this creates unnecessary open-ended reconnaissance capability that can be used to discover and pivot to arbitrary domains, undermining scope restrictions.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill is scoped to collecting Amazon product-list records, but this client exposes a generic web-unlocking function that can fetch arbitrary public URLs with JS rendering and redirect following. That materially broadens capability beyond the declared purpose and can be repurposed for unrestricted scraping or access to non-Amazon targets, increasing abuse potential and weakening least-privilege boundaries.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest describes a skill to collect Amazon product-list records by keyword and marketplace domain, which aligns with submitting a collection task. However, this script also waits for task completion and outputs the final result payload via complete_task, expanding behavior from simple submission into result retrieval/monitoring that is not stated in the description.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code treats only the Chinese strings "处理中", "成功", and "失败" as valid status values. This imposes a specific language expectation in the skill logic without user opt-in or documented locale justification, which matches the policy's language/locale violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The instruction 'Always call it API TOKEN in user-facing instructions' mandates a specific language choice for all user interactions. This is a natural-language policy concern because it removes flexibility to match the user's preferred terminology or locale and does not offer opt-in or alternatives.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/task_runtime.py:38