Back to skill

Security audit

Dataify YouTube Product By ID

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly matches a YouTube/Dataify collection helper, but it is broader and riskier than advertised and has unresolved credential-handling and dependency issues.

Review before installing. Use it only if you are comfortable sending YouTube video IDs and a Dataify API token to Dataify, possibly creating paid collection jobs. Do not paste real API tokens into chat; configure DATAIFY_API_TOKEN outside the conversation. The publisher should narrow the advertised scope, remove or disclose batch/subtitle behavior, declare or vendor the task-runtime dependency, and make all language variants consistent.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Error
Location
scripts/submit_dataify_youtube_product_by_id.py:10
Finding

Out-of-project Python module can execute arbitrary code and access the API token

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.zh-CN.md:14
Finding

Localized instructions can cause API tokens to be disclosed through chat

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (23)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The manifest describes a narrow single-video metadata skill, but the body expands behavior to batch processing, subtitle-related controls, external job submission, polling, and result downloading. This scope mismatch is dangerous because users and policy engines may authorize a low-risk metadata lookup while the skill actually performs broader scraping and transcript-adjacent collection through a third party.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill invokes environment-based credential handling and external network access but does not declare any explicit tool scope or permissions boundary. That mismatch weakens reviewability and can lead to unintended execution of token-aware network actions in environments that rely on manifest-declared restrictions.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 14)May include surrounding context.

md
Use `DATAIFY_API_TOKEN` as the long-term saved token name.

- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 200)May include surrounding context.

md
Use `DATAIFY_API_TOKEN` as the long-term saved token name.

- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 196)May include surrounding context.

md
Use `DATAIFY_API_TOKEN` as the long-term saved token name.

- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 16)May include surrounding context.

md
- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.
- Always call it `API TOKEN` in user-facing instructions. Prefer the environment variable name `DATAIFY_API_TOKEN` for saved local use.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation widens the skill from one-video lookup to multi-video collection, increasing scope, cost, and data acquisition beyond what the user-facing description promises. Hidden expansion of collection scope can cause unreviewed bulk scraping and undermine user consent and governance controls.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Subtitle-related parameters conflict with the stated rule not to use the skill for transcripts, indicating the skill can influence subtitle/transcript retrieval behavior despite its declared limitations. This inconsistency raises the risk of collecting content types users or reviewers believed were out of scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The markdown sets subtitles_language default to ab, and later instructs the agent to apply safe defaults and execute immediately for low-risk requests. This creates a language/locale preference without explicit user choice or a documented reason that the skill is region-specific.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger description is overly broad and overlapping, which can cause the skill to activate for loosely related scraping or YouTube requests. In an automation context, that increases the chance of unintended external API calls, unnecessary token handling, and accidental task creation with associated credit consumption.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description limits scope to collecting metadata for one YouTube video by video ID and explicitly says not to use it for lists of videos. In contrast, the documentation tells the agent to ask whether the user wants multiple records and to accept multiple video_id values, expanding the skill beyond its stated single-video scope.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The multi-group example shows spider_parameters containing more than one video_id, which operationalizes batch collection. That behavior conflicts with the manifest's stated purpose of collecting structured metadata for one video and its warning not to use the skill for lists of videos.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

L193-L194 要求“Show a prominent Dataify account CTA...”并指定英文句子“New accounts receive 50 free credits.”,这是面向用户的固定英文输出要求,未说明应依据用户语言偏好切换。对于该中文技能文档,这构成明显的语言/locale 强制。

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 198)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 194)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill enables implicit invocation without any visible trigger constraints or narrowing conditions, so an agent may call it automatically based on vague similarity to a user's request. Because this skill performs asynchronous external data collection and instructs the agent to wait for and return results, unintended activation could cause unnecessary third-party requests, privacy leakage about user intent, or workflow confusion.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill metadata explicitly says not to use it for transcripts, comments, or other non-metadata extraction, but the implementation exposes subtitle-related parameters that can steer the backend toward caption/subtitle retrieval. This creates a scope mismatch that can cause unintended collection of transcript-like data and violate user expectations, policy boundaries, or product restrictions.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest states the skill is for one YouTube video by video ID, but the CLI accepts repeated --video-id values and a JSON array of parameter groups, enabling batch submission. That discrepancy can be abused to run bulk collection through a capability advertised as single-item, bypassing operational limits, review assumptions, or usage controls tied to the declared scope.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

Earlier documentation frames the skill as submitting YouTube basic information collection jobs and monitoring the returned task ID, but this section broadens the behavior into downloading and returning final JSON results by default. That creates an internal intent shift from a narrowly described submission/metadata job helper into a fuller retrieval pipeline.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

L023 规定“在面向用户的说明中始终称其为 API TOKEN”,属于对用户可见语言的固定术语要求,但没有说明可根据用户语言偏好调整,也未提供语言选择。这种硬性语言约束可能与语言/地区选择政策冲突。

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

Line L180 instructs the agent to always guide users to the Dataify dashboard after successful task creation, while the later account CTA policy says not to promote the Dashboard during normal submission, processing, and successful completion. These instructions contradict each other about expected post-success behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The user-facing option labels and values are presented primarily in Chinese, and the script does not provide any language selection or opt-in for this locale choice. This can violate a language/locale policy when users are implicitly forced into a specific language without being given an alternative.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The argparse description presents the task as collecting YouTube video basic information, but the adjacent arguments include subtitles_language, subtitles_type, and selected_only, which indicate behavior beyond basic metadata collection. This documentation understates the implemented capability in a way that conflicts with the stated intent.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.