Back to skill

Security audit

Dataify YouTube Audio By URL

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly for Dataify YouTube audio collection, but it needs review because it can use a saved token automatically, submit third-party jobs that may consume credits, expose subtitle options despite excluding transcripts, and load unaudited sibling code with the API token.

Install only if you trust Dataify and the required dataify-task-operations runtime in the same environment. Use an environment variable or secret manager for DATAIFY_API_TOKEN, do not paste tokens into chat, confirm each submission before it runs, verify whether credits may be consumed, and leave subtitle options disabled unless subtitle collection is intended.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
scripts/submit_dataify_youtube_audio_by_url.py:10
Finding

Untrusted Sibling Runtime Is Imported and Given the API Token

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
SKILL.zh-CN.md:14
Finding

Conflicting Token-Handling Instructions May Cause API Tokens to Be Disclosed in Chat

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (20)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The manifest says the skill is only for YouTube audio and not for transcripts/subtitles, but the documented parameters allow subtitle-related collection via subtitles_language and is_subtitles. This creates a deceptive scope expansion where users or calling systems may authorize the skill for a narrower purpose than it actually performs, increasing privacy, compliance, and misuse risk.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill explicitly instructs the agent to use environment-stored credentials and make external network requests, but it declares no tool scope or permissions boundaries. That mismatch weakens governance and can let a host agent execute sensitive env/network actions without clear least-privilege constraints or user visibility.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 14)May include surrounding context.

md
Use `DATAIFY_API_TOKEN` as the long-term saved token name.

- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 214)May include surrounding context.

md
Use `DATAIFY_API_TOKEN` as the long-term saved token name.

- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 16)May include surrounding context.

md
- If `DATAIFY_API_TOKEN` is saved locally, use it without asking the user to re-enter the token.
- If the user does not have an API TOKEN, tell them they can register or log in at [Dataify](https://dashboard.dataify.com/login?utm_source=skill) to get one.
- If the user wants to save it, give the appropriate command for their shell and ask them to run it; do not silently persist tokens without confirmation.
- Do not call the Builder endpoint without a token.
- Always call it `API TOKEN` in user-facing instructions. Prefer the environment variable name `DATAIFY_API_TOKEN` for saved local use.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation exposes subtitle/transcript-related parameters even though the manifest excludes transcripts. This inconsistency broadens the effective capability of the skill and can bypass user expectations or platform policy decisions based on the declared narrower scope.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The default behavior goes beyond submitting an audio-collection job and instructs the agent to continue monitoring tasks and retrieve final JSON results automatically. This expands the skill's operational scope and data access beyond the simple action suggested by its name and description, which can surprise users and increase unintended data handling.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger description is overly broad and includes vague phrases that overlap with common user language, increasing the chance the skill activates in contexts the user did not intend. Because this skill can submit external collection jobs and consume stored credentials, over-triggering can cause unintended third-party requests, data transfer, and account-credit usage.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill describes transmitting user-supplied URLs and using a locally stored API token to make external network requests, but it does not clearly warn users that their input will be sent to a third-party service. In a skill that automates external submissions, missing disclosure weakens informed consent and can lead to unintended sharing of user data or use of stored credentials.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 212)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 208)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Earlier sections explicitly say that if the user provides a token in the request, the skill should use that token, which implies accepting a token through the conversation. Line L208 then says 'Never ask the user to paste the token into chat,' which directly conflicts with the documented token-handling workflow and changes the intended interaction model.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
84% confidence
Finding

The instruction to continue the original task automatically after merely verifying that DATAIFY_API_TOKEN is present encourages the agent to proceed with an external action without renewed user confirmation. In this skill's context, that can trigger third-party requests and consume account credits using stored credentials, which makes autonomous continuation materially riskier than in a read-only workflow.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 210)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Multiple earlier lines instruct the skill to tell users to visit Dataify to view or manage results after task creation, including mandatory wording after success. Line L213 instead states that during normal submission, processing, and successful completion, the skill should not promote registration or the Dashboard, which is an active contradiction in the documented intent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill enables allow_implicit_invocation without any visible trigger constraints, exclusions, or user-confirmation requirements. Because this skill can initiate external collection of YouTube audio from a URL and waits on an asynchronous task by default, a model may invoke it opportunistically from ambiguous user requests, causing unintended external actions, unnecessary data processing, or policy-bypassing tool use.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill manifest says it should only download or collect YouTube audio and explicitly excludes transcripts, yet the script exposes subtitle-related parameters (subtitles_language and is_subtitles). That creates a capability mismatch that can cause collection of transcript-like data outside the declared scope, undermining user consent and policy boundaries.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This code presents multiple option labels, including audio format, subtitle language, and boolean selections, in Chinese string literals. Because the script does not offer a locale selection or document that it is intentionally region-specific, it effectively imposes a specific language on users, which is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The submitted spider_universal payload includes subtitles_language and is_subtitles, meaning the backend task can be instructed to fetch subtitles despite the skill being described as audio-only. This is dangerous because it silently expands data collection beyond the approved scope and may lead to unauthorized transcript acquisition.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The policy labels the action as 'read-only' and 'low-cost' even though the skill submits external collection jobs that may consume paid credits and trigger processing on a third-party service. Misclassifying a paid external submission as read-only can reduce user scrutiny and encourage automatic execution of billable actions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

“在面向用户的说明中始终称其为 API TOKEN”规定了固定语言表述,但没有说明这是用户可选择的术语偏好,也未提供地域或合规上的必要性。该类强制语言/术语要求属于自然语言策略约束,应提供用户选择或明确 justification。

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.