Back to skill

Security audit

Dataify YouTube Video Post

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its YouTube/Dataify collection purpose, but it loads unaudited helper code from outside the package and passes it the API token.

Review before installing. This skill sends YouTube collection targets and a Dataify bearer token to Dataify, and its helper script depends on task_runtime code located outside this package. Install only if you trust the surrounding Dataify task-operations runtime and are comfortable with implicit, paid or credit-consuming collection jobs using your DATAIFY_API_TOKEN.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Tool Hijacking and SpoofingModifies or replaces tools so legitimate-looking calls execute attacker logic
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.zh-CN.md:316
Finding

Localized Skill Instructions Hijack Successful Responses with Mandatory Third-Party Promotion

Content
View full analysis
Remediation
View remediation

T07 · Tool Hijacking and Spoofing

Error
Location
scripts/submit_dataify_youtube_video_post.py:11
Finding

Unaudited Sibling Module Is Imported from a Precedence-Modified Path and Receives the API Token

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (16)

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest says the skill must not be used to download media files, but the body explicitly permits media-download handling by recommending extended timeouts for such operations. This creates policy drift that can enable prohibited collection behavior, higher-cost actions, or handling of content types outside the stated safety and compliance boundary.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill is presented as a read-only collector for YouTube video-post records, yet later instructions discuss media-download usage. This contradiction makes the skill more dangerous because operators and reviewers may trust the narrower description while the implementation guidance expands into a more sensitive capability set.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill operationally requires environment-variable access and outbound network access, but it does not declare tool scope or permissions. That mismatch can cause an agent platform to run the skill with broader-than-expected privileges or without proper review, reducing visibility into secret handling and external data exfiltration paths.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 171)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 330)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 173)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 332)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

该描述将触发条件扩展到“通过关键词/URL/标签/探索抓取 YouTube 视频等含义时触发”,属于按语义泛化匹配,而不是限定的明确短语列表。虽然主题域是 YouTube 抓取,但缺少排除条件或负例,容易覆盖大量普通的 YouTube 数据收集请求并导致非预期调用此技能。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
80% confidence
Finding

整个技能文件以中文撰写,并多处直接规定面向用户的固定表述方式,但没有说明仅适用于中文用户场景,也没有提供语言切换或用户自选机制。对于通用技能,这会构成默认强制特定语言输出的自然语言政策问题。

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The skill explicitly encourages direct execution of collection requests under 'low-risk' conditions without a clear disclosure that user-supplied inputs will be sent to an external third-party service. That can cause unintended data exfiltration of URLs, keywords, or other sensitive research targets, especially when users may not realize a network submission occurs immediately.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The earlier workflow and success-handling instructions say to tell users to visit the Dataify dashboard after successful submission or task creation, including for viewing/managing results. The later Account CTA policy explicitly says that during normal submission, processing, and successful completion, the skill should not promote registration or the Dashboard, which is a direct documentation-level contradiction about post-success behavior.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill enables implicit invocation with no visible activation constraints, so an agent may trigger it broadly based on vague user intent. Because this skill can launch asynchronous third-party data collection jobs, over-broad triggering can cause unintended external requests, unnecessary data processing, cost exposure, and disclosure of user-supplied targets to the external service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This markdown file includes parameter values such as 最新, 热门, and 最早 while other parts of the document are written in English. The mixed-language requirement could force a specific locale or confuse users unless the skill explicitly states that these values are locale-bound or offers a language/locale choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The duration and upload_date option tables mix Chinese labels like 4 分钟以内, 上一小时, and 今年 with English values such as Last hour and This year. Because the file does not explain the required locale behavior, it risks imposing a language constraint without user choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The script hard-codes Chinese values such as "最新", "热门", and "最早" for accepted parameters and maps mixed English inputs to Chinese duration values. This imposes a specific language/locale behavior in the skill interface without offering a language choice or explaining why a Chinese locale is required.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The function submits collected parameters and a bearer token to an external Dataify endpoint via HTTP POST. While this is the script's purpose, the file lacks a clear user-facing disclosure near the network operation or in comments/docstrings describing that user-provided search terms, URLs, and other parameters are transmitted to a third-party service.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.