Back to skill

Security audit

Dataify Crunchbase Builder

Security checks for vulnerabilities and agentic risk

Overview

The skill describes a Crunchbase Dataify integration, but it includes broader scraping workflows and unsafe credential handling that users should review before installing.

Install only if you are comfortable sending target company URLs and parameters to Dataify and using a DATAIFY_API_TOKEN from the agent environment. Prefer a session-scoped token, rotate it if it may have been logged, and avoid using the included broad business_workflow.py or arbitrary task-ID waiting unless the package is narrowed and the import and token-handling issues are fixed.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/wait_for_task.py:35
Finding

API Token Exposed in Task Polling and Download Query Strings

Content
View full analysis
") try: return json.loads(text) except ValueError: raise RuntimeError("Dataify returned a non-JSON response") ``` The credential-bearing call sites are: ```python payload = request_json( STATUS_ENDPOINT, {"api_key": api_key, "task_id": task_id}, api_key, request_timeout, ) ``` ```python return request_json( ...[truncated 2541 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Error
Location
scripts/build-dataify-request.py:5
Finding

Unsafe Module Resolution Allows a Sibling Skill to Override Bundled Code

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (33)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The skill can monitor arbitrary Dataify task IDs and download completed JSON outputs, which is broader than scraping known Crunchbase company URLs. This matters because a user or another component could use the skill as a generic accessor for other remote scraping jobs and datasets under the cover of a benign-sounding Crunchbase-only label.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill can monitor arbitrary Dataify task IDs and download completed JSON outputs, which is broader than scraping known Crunchbase company URLs. This matters because a user or another component could use the skill as a generic accessor for other remote scraping jobs and datasets under the cover of a benign-sounding Crunchbase-only label.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The skill can monitor arbitrary Dataify task IDs and download completed JSON outputs, which is broader than scraping known Crunchbase company URLs. This matters because a user or another component could use the skill as a generic accessor for other remote scraping jobs and datasets under the cover of a benign-sounding Crunchbase-only label.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill can monitor arbitrary Dataify task IDs and download completed JSON outputs, which is broader than scraping known Crunchbase company URLs. This matters because a user or another component could use the skill as a generic accessor for other remote scraping jobs and datasets under the cover of a benign-sounding Crunchbase-only label.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill can monitor arbitrary Dataify task IDs and download completed JSON outputs, which is broader than scraping known Crunchbase company URLs. This matters because a user or another component could use the skill as a generic accessor for other remote scraping jobs and datasets under the cover of a benign-sounding Crunchbase-only label.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill can monitor arbitrary Dataify task IDs and download completed JSON outputs, which is broader than scraping known Crunchbase company URLs. This matters because a user or another component could use the skill as a generic accessor for other remote scraping jobs and datasets under the cover of a benign-sounding Crunchbase-only label.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill can monitor arbitrary Dataify task IDs and download completed JSON outputs, which is broader than scraping known Crunchbase company URLs. This matters because a user or another component could use the skill as a generic accessor for other remote scraping jobs and datasets under the cover of a benign-sounding Crunchbase-only label.

Content

No source excerpt is available for this finding.

YARA rule 'backdoor_persistence': Backdoor persistence with malicious payloads (shell commands, SSH key injection, hidden root users) [malware]

High
Category
YARA Match
Confidence
75% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · SKILL.md (reported line 59)May include surrounding context.

量,而不是只在当前终端临时设置。

Windows PowerShell,当前用户永久设置:

powershell
[Environment]::SetEnvironmentVariable("DATAIFY_API_TOKEN", "your_token_here", "User")

然后重新打开 PowerShell。如果当前会话也要立即生效,再执行:

powershell
$env:DATAIFY_API_TOKEN = "your_token_here"

macOS 或 Linux,bash 永久设置:

bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.bashrc
source ~/.bashrc

macOS 或 Linux,zsh 永久设置:

bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.zshrc
source ~/.zshrc

脚本用法

Python:

bash
python scripts/build-dataify-request.py --tool-sign <selected_tool_sign> --values-file values.json

PowerShell:

powershell
& ".\scripts\build-dataify-request.ps1" -ToolSign "<selected_tool_sign>" -ValuesFile ".\values.json"

values.json 可以是单个对象,也可以是对象数组。

输出格式

最终 curl 命令应为

YARA rule 'backdoor_persistence': Backdoor persistence with malicious payloads (shell commands, SSH key injection, hidden root users) [malware]

High
Category
YARA Match
Confidence
75% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 49)May include surrounding context.

量,而不是只在当前终端临时设置。

Windows PowerShell,当前用户永久设置:

powershell
[Environment]::SetEnvironmentVariable("DATAIFY_API_TOKEN", "your_token_here", "User")

然后重新打开 PowerShell。如果当前会话也要立即生效,再执行:

powershell
$env:DATAIFY_API_TOKEN = "your_token_here"

macOS 或 Linux,bash 永久设置:

bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.bashrc
source ~/.bashrc

macOS 或 Linux,zsh 永久设置:

bash
echo 'export DATAIFY_API_TOKEN="your_token_here"' >> ~/.zshrc
source ~/.zshrc

脚本用法

Python:

bash
python scripts/build-dataify-request.py --tool-sign <selected_tool_sign> --values-file values.json

PowerShell:

powershell
& ".\scripts\build-dataify-request.ps1" -ToolSign "<selected_tool_sign>" -ValuesFile ".\values.json"

values.json 可以是单个对象,也可以是对象数组。

输出格式

最终 curl 命令应为

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This file is a shared multi-workflow orchestrator for price, review, lead, and brand intelligence, which materially exceeds the advertised skill scope of collecting Crunchbase company profiles from known company URLs. Scope expansion is dangerous because it enables broader data collection and search behavior than the user or platform reviewer would expect, undermining least-privilege and trust boundaries.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The code generates generic Google, Google News, Google Shopping, LinkedIn, and Crunchbase searches, and also accepts arbitrary source URLs for scraping. For a skill that claims to process known company URLs for Crunchbase profiles, this broad discovery behavior can collect unrelated external data and cause unexpected outbound requests far beyond the declared purpose.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The fallback path sends user-supplied queries and URLs to generic third-party search and web-unlocking APIs rather than a Crunchbase-specific collector. This creates undisclosed external data sharing and permits broad web access inconsistent with the skill's stated role, making misuse and overcollection more likely.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares no explicit tool scope or permissions even though its instructions clearly rely on environment access, file reads, shell execution, and outbound network use. In an agent environment, missing scope boundaries can let the skill exercise broader capabilities than users would reasonably expect, increasing the chance of unintended data access or external transmission.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The instructions explicitly require the user to choose from a Chinese list, which imposes a specific language on the interaction. The file does not offer an opt-in, alternative language, or a documented region-specific justification for this locale constraint.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

The skill explicitly transmits user-supplied parameters and task metadata to an external third-party endpoint. External transmission is expected for this type of integration, but it remains a real security/privacy concern because sensitive URLs, business targets, and returned data are sent outside the local trust boundary.

Content

Scanner excerpt · SKILL.md (reported line 41)May include surrounding context.

md
13. Set `spider_name` to `crunchbase.com`.
14. Set `spider_id` to the selected tool's `tool_sign`.
15. Always include `spider_errors=true` and `file_name={{TasksID}}`.
16. Return a curl command for `https://scraperapi.dataify.com/builder`.

## Set DATAIFY_API_TOKEN

External Transmission

Medium
Category
Data Exfiltration
Confidence
85% confidence
Finding

The skill explicitly instructs the agent to construct and submit authenticated requests to an external service using company URLs and user-supplied parameters. Any such outbound transmission can expose sensitive targets, query contents, or operational metadata to a third party, and the risk is heightened because the workflow is designed to automate request creation for an external scraping platform.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 3)May include surrounding context.

md
---
name: "dataify-crunchbase-company-by-url"
description: "为 crunchbase.com 上以 crunchbase_company_by-url 为根的 scraper 系列准备 Dataify builder 请求。当需要处理成功的 Dataify scraper detail 条目 crunchbase_company_by-url、让用户选择可用工具、读取已保存的 getToolParams 选项,并使用 DATAIFY_API_TOKEN 生成 scraperapi.dataify.com/builder curl 请求时,使用此 skill。"
---

# Dataify Builder Skill 中文版

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

该文件包含自然语言指令“先让用户从下面的中文工具列表中明确选择一个工具”,对交互语言作了强制性限定。文档未说明这是用户可选项,也未说明该技能仅适用于中文环境,因此构成语言/locale 策略风险。

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 134)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 112)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 136)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 114)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts get 50 free credits, enough for about 6,000 trial results, valid for 7 days, and only successful requests are billed. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The UI metadata and default prompt broaden the skill from a narrowly scoped 'company-by-URL' Crunchbase lookup into a more generic 'Crunchbase Builder' and 'Dataify collection' capability. This scope expansion can cause an agent or user to invoke the skill for unsupported collection tasks, increasing the chance of unauthorized data gathering, policy bypass, or misuse beyond the intended input constraints.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill is described as accepting one or more known company URLs, but the tool parameters expose an additional keyword-based Crunchbase lookup. This creates a scope mismatch that can enable broader search behavior than the user or policy expects, increasing the risk of unauthorized data collection, misuse for general research, or bypass of workflow restrictions tied to URL-only input.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

A keyword search function is not aligned with the stated purpose of processing known company URLs and materially broadens the operational scope of the skill. In context, this makes the skill more dangerous because it can be repurposed for general company discovery or enrichment workflows that the description explicitly says should not be performed.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring explicitly states that the file is shared infrastructure for multiple Dataify business skills, which conflicts with the manifest's narrow single-skill purpose. Such mismatch is a security concern because reviewers and users may approve or invoke the skill under false assumptions about its capabilities.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/task_runtime.py:38