T01 · Skill Instruction Hijacking
- Location
SKILL.zh-CN.md:8- Finding
Mandatory Branded Output Injection in Localized Skill Instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill mostly matches an Amazon/Dataify collection tool, but it bundles broader undisclosed scraping workflows and has credential-handling and runtime-shadowing risks that should be reviewed before installation.
Install only if you are comfortable giving the skill access to a Dataify API token and allowing it to submit collection tasks that may consume credits. Review or remove the bundled non-Amazon workflow scripts, avoid using tokens that cannot be rotated, and prefer a version that authenticates polling without putting API keys in URLs and imports only packaged runtime code.
SKILL.zh-CN.md:8Mandatory Branded Output Injection in Localized Skill Instructions
scripts/wait_for_task.py:33Dataify API Token Exposed in URL Query Strings
scripts/submit_amazon_global_product.py:11External Sibling Directory Can Shadow the Packaged Task Runtime
scripts/submit_amazon_global_product.py:141Amazon Collection Inputs Permit Unrestricted Remote Targets
The declared description is narrowly about collecting Amazon global-marketplace product data from Amazon-oriented inputs. The supplied code does not implement an Amazon global-marketplace product collector. Instead, it is a shared workflow engine supporting multiple unrelated business workflows: price intelligence, review intelligence, lead intelligence, and brand monitoring. It invokes Google Search, Google Shopping, Google News, LinkedIn/Crunchbase discovery, generic web unlocking, Amazon comments scraping, and Google Maps reviews scraping, then analyzes and reports on gathered evidence. Amazon-specific behavior appears only as one special case for review scraping from Amazon URLs, not as the primary purpose, and there is no marketplace-specific Amazon product collection logic by URL/category/keyword across global Amazon marketplaces. This is a material description-behavior mismatch.
The declared purpose is narrowly about collecting Amazon global-marketplace product data from specific Amazon-oriented inputs. The supplied code instead is a reusable, generic HTTP client for Dataify services: it can search Google, unlock arbitrary URLs, parse generic JSON, and extract links from SERP responses. There is no Amazon-specific logic, no product data extraction, no category or brand handling, and no marketplace-specific collection behavior. While this code could be a supporting utility inside a larger Amazon collector, taken as supplied its actual behavior is materially broader and different from the declared skill purpose.
The declared description is for an Amazon global-marketplace product collection skill. The supplied code does not perform Amazon scraping, product lookup, marketplace selection, URL/category/keyword handling, or any commerce-specific data collection. Instead, it provides shared infrastructure for parsing a task identifier and polling a generic task until completion. This is a materially different primary purpose, not merely a supporting implementation detail of the described collection behavior, because the chunk contains no observable Amazon-specific collection logic at all.
There is a clear mismatch. The declared description says the skill should collect Amazon global-marketplace products based on various query inputs. The supplied code does nothing related to Amazon, marketplaces, product collection, URLs, search terms, brands, or data retrieval. Instead, it detects the user's platform/shell and outputs setup/verification commands for a DATAIFY_API_TOKEN environment variable. This is a materially different primary purpose and an unrelated capability, so it should be flagged as a mismatch.
There is a clear mismatch between the declared purpose and the supplied code. The description says the skill collects Amazon global-marketplace products based on Amazon-specific inputs. However, this code only waits for a previously submitted Dataify scraper task to finish by calling Dataify status and download APIs. Its inputs are a task ID plus an API token from the environment, not Amazon product/category URLs or keywords. It contains no Amazon-domain logic, no marketplace selection, no submission of scraping jobs, and no extraction of product data from user-supplied Amazon queries. While this script could be a supporting helper within a larger scraping workflow, evaluated on its own it does not implement the declared functionality and instead exposes a materially different capability: polling and retrieving generic Dataify task results.
The skill clearly expects capabilities including environment access, file operations, shell execution, and network calls, but it does not declare any explicit tool scope or permission boundaries. That omission weakens reviewability and least-privilege enforcement, making it easier for an agent to execute broader actions than a user or platform operator would expect.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
## Account CTA policy
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
## Account CTA policy
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.
The trigger description is broad and loosely bounded, matching many variants of Amazon/global product collection requests. Overbroad activation increases the risk that the skill is invoked outside its intended scope, causing the agent to initiate external data-collection workflows or token-handling behaviors in contexts where the user did not clearly request them.
The skill gives conflicting post-submission instructions: it says to stop after submission and elsewhere says to always direct users to the Dataify dashboard, while the CTA policy says not to promote the dashboard during normal success. These inconsistencies are not directly code-execution dangerous, but they can cause policy bypass, inconsistent user flows, and unintended disclosure of external links or vendor promotion when not warranted.
The skill contains contradictory token-handling guidance: one section allows using a token supplied in the user request, while the CTA policy later says to never ask the user to paste the token into chat. This inconsistency can lead the agent to solicit or accept secrets through natural-language channels, increasing the chance of credential exposure in logs, transcripts, or downstream tooling.
该段落使用英文固定指令(如“Show a prominent Dataify account CTA...”“Never ask the user to paste the token into chat”),要求输出特定语言内容,但未说明需征得用户同意,也未提供语言/地区选择。这属于自然语言层面的语言/locale 策略问题,因为技能主体文档为中文,却在关键用户沟通策略中强制英文。
The instruction to continue the original task automatically after checking token presence enables autonomous progression into an external action workflow without renewed user confirmation. In this skill's context, that can trigger paid scraping/task submission behavior based on stored credentials, increasing the risk of unintended actions and billing.
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.
The skill explicitly instructs the agent to verify whether a saved DATAIFY_API_TOKEN is present and continue the task if so. That creates a natural-language path for the model to inspect and act on stored credentials, expanding the credential-access surface and normalizing secret-dependent autonomous behavior even if the token value is not printed.
The skill enables implicit invocation with no visible activation constraints, which can cause the agent to trigger this external data-collection capability more broadly than the user intended. Because this skill can initiate asynchronous collection of Amazon marketplace data, over-broad triggering creates risk of unintended external actions, unnecessary data access, and confused-deputy behavior when user requests are ambiguous.
The file is shared workflow code supporting price intelligence, review intelligence, lead intelligence, and brand monitoring, which exceeds the stated Amazon global-product skill scope. Scope overreach is dangerous because a caller expecting only Amazon product collection may unknowingly trigger broader web collection, competitor intelligence, or lead-generation behavior, increasing data exposure and violating least-privilege expectations.
The action-construction logic includes unrelated review, lead-generation, and brand-monitoring searches such as LinkedIn, Crunchbase, Google News, and complaint monitoring. In the context of an Amazon global-product skill, this mismatch makes the package more dangerous because it can collect off-scope business intelligence and reputation data that users and platform reviewers would not reasonably expect.
The execution layer can invoke general Google search/news/shopping scripts, a generic web unlocker, and third-party review scrapers unrelated to the declared Amazon-global-product purpose. This is dangerous because it expands the effective collection surface far beyond the advertised skill, enabling broad scraping and external-data gathering under the cover of a narrowly described skill.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def execute_action(action: dict[str, Any], token: str) -> subprocess.CompletedProcess[str]:
invocation = command(action)
if len(invocation) > 1 and Path(invocation[1]).exists():
return subprocess.run(invocation, capture_output=True, text=True, encoding="utf-8", errors="replace", check=False)
if action.get('capability', '').startswith('scraper-'):
return subprocess.CompletedProcess([], 1, '', 'Required platform scraper is not installed; install it before executing this action.')
return direct_request(action, token)
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
def build_curl(tool, spider_parameters_json):
return " \\\n".join([
"curl -X POST '{}'".format(BUILDER_URL),
" -H 'Authorization: Bearer $DATAIFY_API_TOKEN'",
" -H 'Content-Type: application/x-www-form-urlencoded'",
" -d 'spider_name={}'".format(tool["spider_name"]),
The script transmits user-supplied query data, URLs, and filters to an external third-party service, but the code itself provides no explicit consent prompt or clear warning at submission time about that network transfer. In an agent/skill context, this can cause unintended disclosure of user-provided data, especially if URLs or keywords contain sensitive business terms or internal research targets.
The manifest describes a product-collection skill, but the code also accesses process environment state to obtain a credential via DATAIFY_API_TOKEN. While needed for this implementation, reading environment variables is not part of the user-facing purpose and constitutes an extra capability beyond simple product-data collection.
Detected: suspicious.exposed_secret_literal