Back to skill

Security audit

Dataify Amazon Product

Security checks for vulnerabilities and agentic risk

Overview

The skill has a coherent Dataify collection purpose, but it runs an unaudited sibling Python module with access to the user's API token and has inconsistent localized instructions.

Review before installing. This skill will use a Dataify API token and make external Dataify requests that may consume credits. The main concern is that its script imports result-monitoring code from a sibling directory outside the reviewed artifact and passes that code the API token, so only use it in an environment where that dependency is trusted and controlled.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.zh-CN.md:8
Finding

Localized Instructions Override the Declared Result-Delivery Workflow

Content
View full analysis

Vulnerability Details

File Location: SKILL.zh-CN.md, lines 8–9 and 190–191
Vulnerability Type: Conflicting localized Skill instructions
Risk Level: Medium

Vulnerable Instruction Snippet

The following is an English translation of the relevant localized instructions:

text
Submit the Amazon product collection task through Dataify Builder and stop
after submission. Do not download the result file. After successful
submission, provide the task_id and direct the user to Dataify to view the
result.

After successfully creating a task, always direct the user to Dataify.

Technical Analysis

The localized instructions materially conflict with the primary workflow in SKILL.md, which states that the Skill should monitor the submitted task and return the final collected result by default. They also conflict with the account CTA policy later in the same localized file, which prohibits promoting the dashboard during successful processing and completion.

When the localized Skill is loaded, these instructions alter the Agent's expected session behavior by requiring it to stop after task submission and redirect the user to an external dashboard. This is not necessary to submit or monitor the collection task because the included script supports waiting for and returning the final result.

The issue does not grant operating-system privileges, but it changes the Agent's output and completion objective based on which localized documentation is selected.

Attack Path

  1. The Agent selects or loads SKILL.zh-CN.md.
  2. The Agent submits a paid Dataify collection task.
  3. The localized instructions require the Agent to stop after receiving the task ID.
  4. The Agent does not invoke the documented final-result retrieval workflow.
  5. The user is redirected to the Dataify dashboard instead of receiving the result promised by the primary Skill definition.

Impact Assessment

The affected scope is t ...[truncated 413 chars]

Remediation
View remediation

Remediation Suggestions

  • Make the localized workflow semantically identical to SKILL.md.
  • Remove instructions requiring the Agent to stop after submission or always redirect users to the dashboard.
  • Return the final collected result by default and retain submission-only behavior only when the user explicitly requests it.
  • Display account or dashboard links only when authentication fails, credits are insufficient, or the user explicitly asks for account-management information.
  • Add a documentation consistency test that compares security-sensitive workflow statements across localized Skill files.

T08 · Insecure Dependencies

Error
Location
scripts/submit_amazon_product.py:12
Finding

Unaudited Sibling-Directory Module Is Given Import Precedence and Receives the API Token

Content
View full analysis

Vulnerability Details

File Location: scripts/submit_amazon_product.py, lines 12–15 and 318–326
Vulnerability Type: Unsafe external dependency loading
Risk Level: High

Vulnerable Code Snippet

python
TASK_RUNTIME_DIR = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..", "dataify-task-operations", "scripts"))
if TASK_RUNTIME_DIR not in sys.path:
    sys.path.insert(0, TASK_RUNTIME_DIR)
from task_runtime import complete_task

The imported function subsequently receives the API token:

python
if not args.no_wait:
    try:
        final_result = complete_task(task_id, api_token, args.wait_timeout)
    except RuntimeError as exc:
        print(str(exc), file=sys.stderr)
        return 1
    print(json.dumps(final_result, ensure_ascii=False, indent=2))

Technical Analysis

The script constructs a path outside the audited project and inserts it at index zero of sys.path. Python therefore gives that directory precedence when resolving task_runtime.

Importing a Python module executes its top-level code immediately. The script does not verify the external module's ownership, permissions, version, cryptographic digest, or provenance. Furthermore, the imported complete_task function receives the Dataify API token and task ID.

Consequently, anyone able to create or replace the expected sibling module can execute arbitrary Python code under the invoking user's account. The malicious module does not need to wait for complete_task to be called because module-level code runs during import.

This dependency is not included in the supplied project, so its implementation and treatment of the credential could not be audited.

Attack Path

  1. An attacker gains write access to the expected sibling path: dataify-task-operations/scripts/task_runtime.py.
  2. The attacker creates or replaces task_runtime.py with malicious Python code.
  3. A user runs `scrip ...[truncated 1012 chars]
Remediation
View remediation

Remediation Suggestions

  • Do not modify sys.path to load code from a predictable sibling directory.
  • Package task_runtime as an explicit, version-pinned dependency obtained from a trusted source, or include its audited implementation within the Skill package.
  • Use dependency lock files and cryptographic hashes where the deployment mechanism supports them.
  • Ensure dependency directories are not writable by less-trusted users or processes.
  • Verify the dependency's provenance and integrity before execution.
  • Minimize credential exposure by keeping token-based network operations in the audited script rather than passing the token to an externally resolved module.
  • Add a startup check that rejects unexpected module paths and records the resolved dependency path without logging credential values.

T09 · Insecure Skill Coding Practices

Note
Location
SKILL.zh-CN.md:24
Finding

Localized Token Policy Permits Credentials from Conversation Input

Content
View full analysis

Vulnerability Details

File Location: SKILL.zh-CN.md, lines 24 and 204–211
Vulnerability Type: Inconsistent credential-handling instructions
Risk Level: Low

Vulnerable Instruction Snippet

The following is an English translation of the relevant localized instruction:

text
If the user provides a token in the request, use that token for the current run.

The same file later states:

text
Never ask the user to paste the token into chat.

Technical Analysis

The localized policy explicitly permits a token supplied through conversational input to be consumed during the run. This conflicts with the later policy directing users not to paste tokens into chat and with the safer environment-variable workflow used by the script.

A secret included in a request can remain in conversation history, Agent context, telemetry, debugging output, or other transcript storage even if the token is not subsequently printed. Environment variables are the intended credential channel and avoid placing the secret directly in conversational content.

The instruction does not itself solicit the credential, and exploitation depends on a user already including it in the request. The finding is therefore lower risk than direct token logging or hardcoding.

Attack Path

  1. A user includes a Dataify API token in a request, possibly based on the localized policy's statement that such a token can be used.
  2. The Agent processes the token from conversational context.
  3. The token remains present in the request or transcript even after the task completes.
  4. Any party or system with access to stored conversation content may recover and reuse the token.

Impact Assessment

Exposure can permit unauthorized use of the affected Dataify account within the permissions and credit limits associated with the token. Potential consequences include unauthorized task submission, credit consumption, and access to API ope ...[truncated 160 chars]

Remediation
View remediation

Remediation Suggestions

  • Remove the instruction permitting tokens supplied through chat or request text.
  • Require DATAIFY_API_TOKEN to be provided through a session-scoped environment variable or an approved secret manager.
  • Verify only whether the environment variable is present; never print, echo, summarize, or interpolate its value into user-facing output.
  • If a user sends a token in conversation, warn that it may have been exposed and recommend rotating it through Dataify's API-key management interface.
  • Align credential-handling requirements across all localized documentation.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (9)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill instructs use of environment variables and outbound network access to submit jobs to a third-party API, but it declares no tool scope or permission boundaries. This creates an authorization and transparency gap: an agent may perform networked actions and inspect token presence without explicit least-privilege constraints, increasing the chance of unintended external calls or secret-handling mistakes.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 154)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 206)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 156)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 208)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger description is broad enough that routine Amazon-related requests could invoke this skill unexpectedly, causing unintended external API use, task submission, and potential credit consumption. In an agent setting, over-broad activation can route user data or URLs to third-party services without sufficiently specific user intent.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The main workflow and notes say that after successful submission the skill should tell users to go to Dataify to view or manage results, and to always guide them there on success (L008, L052, L192). However, the Account CTA policy says that during normal submission and successful completion, the skill should not promote registration or the Dashboard (L210), which directly conflicts with the stated success behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill enables implicit invocation without any visible constraint on when it may activate. Because this skill can trigger external data collection workflows and instructs the agent to wait for asynchronous completion and return results, an unrelated or ambiguously phrased user request could cause unintended third-party requests, data retrieval, cost incurrence, or disclosure of collected content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

该文件是 zh-CN 技能说明,但在 Account CTA policy 段落中直接以英文规定对用户显示的文案和交互方式,且未说明应基于用户语言偏好本地化输出。此类硬编码语言要求可能导致在未获用户选择的情况下强制使用特定语言。

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.