Back to skill

Security audit

Dataify YouTube Video By URL

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly aimed at Dataify YouTube video collection, but it has unsafe credential-handling ambiguity and loads an unaudited external Python module that can receive the user's API token.

Review before installing. Use this only if you trust both the Dataify service and the local skill package environment. Do not paste API tokens into chat; configure DATAIFY_API_TOKEN locally and rotate any token previously shared in conversation. Be aware that running the script may execute code from an external sibling directory if present, submit paid or quota-consuming Dataify jobs, and retrieve results automatically unless no-wait behavior is requested.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Error
Location
scripts/submit_dataify_youtube_video_by_url.py:10
Finding

Unverified Executable Dependency Loaded from Outside the Skill Package

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.zh-CN.md:14
Finding

Localized Workflow Encourages API Token Disclosure Through Chat

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (17)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The stated purpose is a narrow media-file collection by known YouTube URL, but the skill also supports subtitle-related configuration and retrieval of broader Dataify task results. This mismatch can mislead operators about what data leaves the system and what content is returned, increasing the chance of unintended collection scope and policy bypass.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill contains conflicting credential-handling instructions: earlier sections tell the agent to request the API token from the user, while the CTA policy says never to ask the user to paste the token into chat. In practice, this ambiguity can lead agents to solicit secrets through chat, increasing the chance of credential disclosure in conversation logs or downstream systems.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill documents use of environment variables and outbound network access, but it does not declare an explicit tool scope such as permissions or allowed-tools. That weakens sandboxing and review controls because the agent may use capabilities broader than what a caller or platform expects, especially when handling tokens and remote job submission.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 14)May include surrounding context.

md
Use `DATAIFY_API_TOKEN` as the long-term saved token name.

- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 210)May include surrounding context.

md
Use `DATAIFY_API_TOKEN` as the long-term saved token name.

- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 16)May include surrounding context.

md
- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.
- Always call it `API TOKEN` in user-facing instructions. Prefer the environment variable name `DATAIFY_API_TOKEN` for saved local use.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The default completion behavior does more than submit a task: it automatically monitors the job and downloads the final JSON result. That expands the action surface from a simple request trigger into continued remote polling and data retrieval, which can increase data exposure, costs, and user surprise if executed automatically.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill allows immediate execution for some requests despite performing media downloads and remote task submission that may consume credits, bandwidth, storage, and potentially retrieve copyrighted or sensitive material. Without an upfront warning about system and data impact, users may trigger consequential actions unintentionally.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Automatically continuing the original task once the environment token is detected reduces friction, but it can also resume a networked, potentially paid action without a fresh confirmation at the moment credentials become available. In this skill's context, that can lead to unintended remote submissions or downloads using stored credentials.

Content

Scanner excerpt · SKILL.md (reported line 214)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest description says the skill is for collecting a YouTube video media file from a known URL, and explicitly excludes unrelated content domains. However, this file also claims the skill is used to receive task status, configure DATAIFY_API_TOKEN, and troubleshoot Dataify Builder requests, which are operational/account-support functions rather than video-file collection itself.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The instruction '在面向用户的说明中始终称其为 API TOKEN' and the overall document being a zh-CN skill file indicate the skill is prescribing a specific language/terminology for user-facing responses without any user opt-in or locale choice. The policy for this audit requires flagging natural-language locale or language constraints when the skill forces a specific language absent an explicit choice or justification.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The stated purpose is to collect/download a YouTube video file by URL, but these instructions add promotional account-creation, API-key management, OS-specific environment setup, and insufficient-credit handling behaviors. Those are support and account-management capabilities that exceed the narrowly described collection scope.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 212)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 208)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill enables implicit invocation without defining any trigger constraints, exclusions, or narrower guardrails. Because this skill can initiate collection of external YouTube media by URL and its default prompt instructs the agent to submit and wait for asynchronous tasks, a model may invoke it too broadly from ambiguous user requests, causing unintended external actions, unnecessary data collection, or policy-violating automation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill manifest explicitly says the tool must not be used for transcripts, yet the implementation exposes a subtitles_language parameter and forwards it to the backend as part of spider_universal. That creates a capability/policy mismatch: users or downstream agents can request subtitle-related data through a skill that is supposed to be limited to media-file retrieval, which may expand data collection beyond the declared scope and bypass higher-level safety or compliance controls.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
75% confidence
Finding

Lines L202-L204 say to execute immediately only for low-risk, read-only requests and to ask for confirmation for a media download. However, lines L184-L192 define media-oriented collection and result retrieval as the default completion behavior, creating an internal contradiction in the documented intent for when the action should proceed automatically.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.