Back to skill

Security audit

Dataify YouTube Profiles

Security checks for vulnerabilities and agentic risk

Overview

The skill fits its YouTube profile collection purpose, but it needs review because it loads an unpinned external Python helper and passes the Dataify API token to it.

Review before installing. Use this only if you trust the Dataify runtime dependency expected at `dataify-task-operations/scripts`, configure the API token through an environment variable rather than chat, and prefer explicit invocation or no-wait mode for tasks that may consume credits.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Error
Location
scripts/submit_dataify_youtube_profiles.py:11
Finding

Untrusted Code Import from an External Sibling Directory

Content
View full analysis

Vulnerability Details

File Location: scripts/submit_dataify_youtube_profiles.py, lines 11–14
Vulnerability Type: Untrusted dependency loading and Python import-path manipulation
Risk Level: High

Vulnerable Code:

python
TASK_RUNTIME_DIR = os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..", "dataify-task-operations", "scripts"))
if TASK_RUNTIME_DIR not in sys.path:
    sys.path.insert(0, TASK_RUNTIME_DIR)
from task_runtime import complete_task

Technical Analysis

The script constructs a path outside the audited Skill package, places that directory at the beginning of sys.path, and imports task_runtime from it. The external component is not included in the project, version-pinned, or integrity-verified.

Inserting the directory at index zero gives modules in that location precedence during Python module resolution. Consequently, a malicious or compromised task_runtime.py placed in the expected sibling directory will execute immediately during import. Python module-level code runs before main() and therefore before the script performs its normal validation or network workflow.

The imported complete_task function is later called with both the returned task identifier and the Dataify API token:

python
final_result = complete_task(task_id, api_token, args.wait_timeout)

This directly exposes the credential to code whose implementation and network behavior cannot be verified from the audited project.

Attack Path

  1. An attacker gains write access to the expected sibling directory, or supplies a compromised dataify-task-operations component.
  2. The attacker creates or modifies dataify-task-operations/scripts/task_runtime.py.
  3. A user or agent invokes scripts/submit_dataify_youtube_profiles.py.
  4. The script prepends the external directory to sys.path.
  5. Python imports and executes the attacker-controlled module-level code.
  6. The maliciou ...[truncated 761 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove the sys.path.insert(0, TASK_RUNTIME_DIR) behavior.
  2. Bundle the required result-monitoring implementation inside the Skill as a normal package using explicit relative imports.
  3. Alternatively, obtain the dependency from a trusted package repository and pin an exact version and cryptographic hash.
  4. Ensure the dependency location is not writable by untrusted users or processes.
  5. Verify the dependency's integrity before importing it if an external runtime is unavoidable.
  6. Avoid passing the long-lived API token to a broad helper module. Expose only the narrowly scoped operation required for task monitoring.
  7. Document every network destination used during result monitoring and restrict outbound requests to an explicit allowlist.
  8. Add installation or startup checks that fail closed if the trusted dependency is absent or fails integrity verification.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.zh-CN.md:21
Finding

Chinese Skill Instructions Permit API Tokens to Be Supplied Through Chat

Content
View full analysis

Vulnerability Details

File Location: SKILL.zh-CN.md, lines 21 and 47
Vulnerability Type: Unsafe credential-handling instructions
Risk Level: Medium

Relevant Instructions, translated into English:

text
- If the user provides a token in the request, use that token.
text
6. Resolve the Dataify token from explicit user input or a saved DATAIFY_API_TOKEN.

Technical Analysis

These instructions explicitly authorize the agent to consume an API token supplied in a user request. Credentials entered into a conversation can be retained in conversation history, application telemetry, debugging records, agent traces, or other logs outside the intended local execution environment.

This behavior conflicts with the safer policy later in the same file, which states that users should never be asked to paste tokens into chat and that the agent should verify only whether DATAIFY_API_TOKEN exists without printing its value. It also conflicts with the English Skill instructions, which direct the implementation to use the environment variable.

Although the bundled Python script reads the token only from DATAIFY_API_TOKEN, an agent following the alternate-language instructions could still handle a token from the conversation and place it into the process environment or another execution channel. The contradictory instructions make credential handling dependent on which section the agent follows.

Attack Path

  1. A user invokes the Skill under the alternate-language instructions.
  2. The agent follows the instruction permitting a token supplied in the request.
  3. The user includes a valid Dataify API token in the conversation.
  4. The token becomes part of the conversation context and may be stored in histories, telemetry, or operational traces.
  5. A party with access to those records obtains the token.
  6. The exposed token can be reused against the Dataify service according to the permissions a ...[truncated 569 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove every instruction that permits tokens to be provided through user messages or explicit conversational input.
  2. Require DATAIFY_API_TOKEN to be supplied through a session-scoped environment variable or an approved secret manager.
  3. Make the English and alternate-language Skill documents use an identical credential policy.
  4. State unambiguously that the agent must never request, repeat, display, summarize, or persist the token.
  5. When checking configuration, verify only whether the environment variable is present and non-empty.
  6. Avoid including credentials in command-line arguments because they may appear in shell history or process listings.
  7. Add automated documentation checks that reject contradictory token-handling instructions across translated files.
  8. If a user posts a token in chat, instruct them to revoke or rotate it and configure the replacement through the secure environment mechanism.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (9)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill instructs the agent to use environment variables and make authenticated network requests, but it does not declare any explicit tool scope such as permissions or allowed-tools. This creates a least-privilege failure: an agent runtime may permit broader env/network access than the user expects, increasing the chance of unintended secret access or outbound requests to external services.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill's default behavior expands beyond submitting a profile collection job into monitoring tasks and downloading final results, and it even references extended timeouts for media downloads despite the manifest saying not to use the skill for media downloads. This scope creep can cause unintended paid actions, larger-than-expected data retrieval, and behavior inconsistent with the user's and platform's expectations.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 181)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 221)May include surrounding context.

md
## Account CTA policy

- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 183)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 223)May include surrounding context.

md
- Show a prominent Dataify account CTA only when the API token is missing, rejected/invalid, or the account has insufficient credits.
- For a missing token, offer https://dashboard.dataify.com/login?utm_source=skill and state: New accounts receive 50 free credits. Never ask the user to paste the token into chat.
- Detect the current operating system and shell. Show only the matching session-scoped setup command first (`export` for macOS/Linux shells, `$env:` for Windows PowerShell, or `set` for Windows Command Prompt). Show other platforms or persistent setup only when detection is ambiguous or the user asks.
- After the user says the token is configured, verify only whether `DATAIFY_API_TOKEN` is present; never print its value. If verification succeeds, continue the original task without asking the user to repeat it.
- Explain that persistent shell changes may require a new terminal or restarting the agent application. Do not recommend a project `.env` unless the execution path explicitly loads it, and ensure `.env` is ignored by version control.
- For an invalid token, direct the user to API-key management without implying that a new registration is required. For insufficient credits, direct the user to balance or recharge management.
- During normal submission, processing, and successful completion, do not promote registration or the Dashboard. Never expose the token or include it in CTA attribution parameters.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The account CTA policy says that during normal submission, processing, and successful completion, the skill should not promote registration or the Dashboard. Earlier workflow and safety instructions explicitly require telling the user to visit the Dataify dashboard after successful task creation to view or manage results, so the documentation gives conflicting behavioral directions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill enables implicit invocation without any visible trigger constraints, which allows the agent to activate this external data-collection capability more broadly than necessary. In a skill that performs asynchronous collection of YouTube profile/channel data, this increases the chance of unintended tool use, unnecessary outbound requests, and data collection actions occurring without sufficiently explicit user intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The file is primarily written in Chinese, but the entire 'Account CTA policy' section is written in English imperative instructions. This creates a language/locale inconsistency and effectively forces a different language for part of the skill behavior without any user opt-in or justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.