Back to skill

Security audit

Dataify ChatGPT Answers

Security checks for vulnerabilities and agentic risk

Overview

This skill is purpose-aligned for Dataify ChatGPT collection, but it can automatically send prompts or ChatGPT URLs to a third-party paid API with broad implicit invocation enabled.

Review before installing. Use this only when you intentionally want Dataify to process the ChatGPT URL or prompt, and avoid confidential prompts or private conversation links unless that transfer is acceptable. Prefer session-scoped token setup when possible, and monitor credit usage because successful collection is billed.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
Findings (15)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

Based on the supplied chunk alone, the code does not implement the declared ChatGPT-answer collection behavior. It merely sets up Python path access to an external operations directory, configures UTF-8 output, and calls run_catalog_builder(...). That suggests a generic request/catalog-building entrypoint, not a clear extractor for ChatGPT responses, markdown/raw variants, citations, links, prompts, or send times. Because the visible behavior is materially different and insufficient to support the specific declared purpose, this is a mismatch.

Content

No source excerpt is available for this finding.

YARA rule 'backdoor_persistence': Backdoor persistence with malicious payloads (shell commands, SSH key injection, hidden root users) [malware]

High
Category
YARA Match
Confidence
75% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · SKILL.md (reported line 92)May include surrounding context.

�在当前终端临时设置。

Windows PowerShell,当前用户永久设置:

powershell
[Environment]::SetEnvironmentVariable("DATAIFY_API_TOKEN", "your_token_here", "User")

然后重新打开 PowerShell。如果当前会话也要立即生效,再执行:

powershell
$env:DATAIFY_API_TOKEN = "your_token_here"

macOS 或 Linux,bash 永久设置:

bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.bashrc
source ~/.bashrc

macOS 或 Linux,zsh 永久设置:

bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.zshrc
source ~/.zshrc

脚本用法

Python:

bash
python scripts/build-dataify-request.py --tool-sign <selected_tool_sign> --values-file values.json

PowerShell:

powershell
& ".\scripts\build-dataify-request.ps1" -ToolSign "<selected_tool_sign>" -ValuesFile ".\values.json"

values.json 可以是单个对象,也可以是对象数组。例如:

json

YARA rule 'backdoor_persistence': Backdoor persistence with malicious payloads (shell commands, SSH key injection, hidden root users) [malware]

High
Category
YARA Match
Confidence
75% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 91)May include surrounding context.

�在当前终端临时设置。

Windows PowerShell,当前用户永久设置:

powershell
[Environment]::SetEnvironmentVariable("DATAIFY_API_TOKEN", "your_token_here", "User")

然后重新打开 PowerShell。如果当前会话也要立即生效,再执行:

powershell
$env:DATAIFY_API_TOKEN = "your_token_here"

macOS 或 Linux,bash 永久设置:

bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.bashrc
source ~/.bashrc

macOS 或 Linux,zsh 永久设置:

bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.zshrc
source ~/.zshrc

脚本用法

Python:

bash
python scripts/build-dataify-request.py --tool-sign <selected_tool_sign> --values-file values.json

PowerShell:

powershell
& ".\scripts\build-dataify-request.ps1" -ToolSign "<selected_tool_sign>" -ValuesFile ".\values.json"

values.json 可以是单个对象,也可以是对象数组。例如:

json

Vague Triggers

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

Enabling implicit invocation without strong activation constraints allows the agent to call this skill automatically based on loose semantic matches rather than explicit user consent. In this context, the skill can collect ChatGPT conversation content, prompts, citations, links, and timestamps, so unintended invocation could expose sensitive conversation data or cause unauthorized external processing.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill asks for prompts and ChatGPT URLs, then sends them to a third-party API service, but the description does not prominently disclose that external transmission. Users may provide sensitive prompts, links, or conversation identifiers without informed consent, creating privacy and data-handling risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The workflow instructs the agent to ask the user to choose from tool names presented only in Chinese, even though the rest of the skill is in English and no opt-in or language preference is requested. This creates a locale/language policy issue because it imposes a specific language on users without explicit choice or justification.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The skill explicitly instructs construction of a curl request that transmits user inputs and an authorization token to an external endpoint. External transmission is expected for this skill's function, but it is still a real security concern because prompts and URLs may contain sensitive content and are sent off-platform to a third party.

Content

Scanner excerpt · SKILL.md (reported line 125)May include surrounding context.

Required output shape

Generate a curl command in this form:

bash
curl -X POST 'https://scraperapi.dataify.com/builder?platform=1' \

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

第 24 行写明“先让用户从下面的中文工具列表中明确选择一个工具”,这构成了对交互语言/locale 的固定要求。文件其余部分也未提供改用其他语言的选项或说明该限制的必要性,因此符合语言/locale policy violation。

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill instructs users to persist an API token in shell startup files and user-level environment storage without warning about plaintext secret exposure, shell history leakage, local account compromise, or shared-machine risk. While this is common operational guidance, omitting safer handling advice increases the chance that long-lived credentials are stored insecurely and later exfiltrated.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
91% confidence
Finding

The documented workflow sends user-supplied prompts, ChatGPT URLs, answer content, and an authorization token to a third-party external service. External transmission is expected for this skill's purpose, but it still creates a real data exposure boundary and should be treated as a privacy and secret-handling risk, especially if prompts or harvested content contain sensitive data.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 127)May include surrounding context.

最终 curl 命令应为:

bash
curl -X POST 'https://scraperapi.dataify.com/builder?platform=1' \
  -H "Authorization: Bearer $DATAIFY_API_TOKEN" \
  -H 'Content-Type: application/x-www-form-urlencoded' \
  -d 'spider_name=chatgpt.com' \

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 173)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 172)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 175)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 174)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The display name, short description, and default prompt are broad enough that the skill could be selected for ordinary requests about ChatGPT answers or generic data collection, even when the user did not intend to invoke this specific capability. Because the skill performs external collection and asynchronous task handling, overbroad matching increases the chance of unintended data retrieval, surprise tool use, or scope creep beyond the user's request.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.