Back to skill

Security audit

Content Collector

Security checks for vulnerabilities and agentic risk

Overview

This content-collection skill has a coherent purpose, but it uses browser cookies and local execution paths in ways that need careful review before installation.

Install only if you are comfortable with this skill saving collected content locally, syncing to an Obsidian vault, downloading article images, and running video-transcription scripts. Before using video features, review or disable browser-cookie access, avoid processing untrusted local media filenames, and prefer pinned dependencies and explicit consent before Obsidian writes or authenticated media retrieval.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/bilibili_extract.py:23
Finding

Cross-Origin Disclosure of Bilibili Session Cookies

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/video_transcribe.sh:182
Finding

Automatic Access to Chrome Browser Cookies Without Explicit Consent

Content
View full analysis
/dev/null || true) ``` ```bash # Download subtitle local temp_base="$OUTDIR/subtitle_temp" local download_cmd="yt-dlp --write-sub --sub-lang $sub_lang --skip-download -o $temp_base" if [ "$platform" = "bilibili" ]; then download_cmd="$download_cmd --cookies-from-browser chrome" fi $download_cmd "$url" 2>&1 | grep -E '(Downloading|Writing|Error)' >&2 || true ``` ```bash # Build yt-dlp command based on platform local ytdlp_cmd="yt-dlp -x --audio-format mp3 --audio-quality 5 -o $OUTDIR/${platform}_${video_id}.%(ext)s --no-playlist" case "$platform" in bilibili) ytdlp_cmd="$ytdlp_cmd --cookies-from-browser chrome" ;; youtube) # No special options needed ;; xiaohongshu|douyin) # Try with cookie first ytdlp_cmd="$ytdlp_cmd --cookies-from-browser chrome" ;; esac ``` ### Technical Analysis The script automatically invokes `yt-dlp --cookies-from-browser chrome` for Bilibili, Xiaohongshu, and Douyin operations. This causes a third-party executable to access Chrome's browser cookie database and associated credential-decryption facilities. The declared functionality is video download and transcription. Public content ordinarily can and should be attempted without authentication first. Automatically accessing a user's primary browser profile breaks least-privilege boundaries, particularly for Xiaohongshu and Douyin, where authenticated access is attempted before the existing unauthenticated fallback. There is no explicit co ...[truncated 1102 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/video_transcribe.sh:548
Finding

Python Code Injection Through Crafted Local Media Filenames

Content
View full analysis
&2 python3 "$MERGER_SCRIPT" "$transcript_json" -o "$merged_json" 2>&1 | grep -v "^$" >&2 || true if [ -f "$merged_json" ]; then # Generate merged plain text python3 -c " import json with open('$merged_json') as f: sentences = json.load(f) with open('$merged_txt', 'w') as f: f.write(' '.join(s['text'] for s in sentences)) " 2>/dev/null local merged_count=$(python3 -c "import json; print(len(json.load(open('$merged_json'))))" 2>/dev/null || echo "0") local raw_count=$(python3 -c "import json; print(len(json.load(open('$transcript_json'))))" 2>/dev/null || echo "0") ``` ### Technical Analysis The basename of a user-selected local media file is incorporated into `merged_json`, `merged_txt`, and `transcript_json`. Those path values are then interpolated directly into Python source passed to `python3 -c`. Shell quoting does not make the generated Python source safe. A filename containing a single quote and Python syntax can terminate or alter the Python string expression. The same unsafe pattern appears in the merged-text generator and the segment-count commands. For example, a filename crafted so ...[truncated 1396 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
scripts/video_transcribe.sh:488
Finding

Unpinned Dependencies and Runtime Package Resolution

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
Findings (49)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill includes transcription, subtitle/media downloading, local media processing, temporary artifact creation, and possible cookie-backed access, none of which are clearly represented in the declared scope. Hidden media download and authenticated retrieval capabilities materially expand attack surface, especially if arbitrary URLs or local files are accepted.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The skill includes transcription, subtitle/media downloading, local media processing, temporary artifact creation, and possible cookie-backed access, none of which are clearly represented in the declared scope. Hidden media download and authenticated retrieval capabilities materially expand attack surface, especially if arbitrary URLs or local files are accepted.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill includes transcription, subtitle/media downloading, local media processing, temporary artifact creation, and possible cookie-backed access, none of which are clearly represented in the declared scope. Hidden media download and authenticated retrieval capabilities materially expand attack surface, especially if arbitrary URLs or local files are accepted.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill includes transcription, subtitle/media downloading, local media processing, temporary artifact creation, and possible cookie-backed access, none of which are clearly represented in the declared scope. Hidden media download and authenticated retrieval capabilities materially expand attack surface, especially if arbitrary URLs or local files are accepted.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 41)May include surrounding context.

md
| B站 | videos | `video_transcribe.sh` 本地转录 |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 42)May include surrounding context.

md
| B站 | videos | `video_transcribe.sh` 本地转录 |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 56)May include surrounding context.

md
| B站 | videos | `video_transcribe.sh` 本地转录 |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 146)May include surrounding context.

md
| B站 | videos | `video_transcribe.sh` 本地转录 |

YARA rule 'info_stealer': Information stealer patterns (credential harvesting, browser data theft) [malware]

High
Category
YARA Match
Confidence
75% confidence
Finding

The YARA match is driven by repeated --cookies-from-browser chrome usage, which resembles credential-harvesting patterns. This does not appear to be classic infostealer malware because the code uses cookies to authenticate media access rather than exfiltrating them directly, but it still performs sensitive credential access and therefore represents a real security concern in an agent-executed skill.

Content

Scanner excerpt · scripts/video_transcribe.sh (reported line 188)May include surrounding context.

sh
1"
    local url="$2"
    local output_file="$3"

    # Only for Bilibili and YouTube
    if [ "$platform" != "bilibili" ] && [ "$platform" != "youtube" ]; then
        return 1
    fi

    echo "Checking for native subtitles..." >&2

    # Build yt-dlp command for subtitle detection
    local ytdlp_cmd="yt-dlp --list-subs"
    if [ "$platform" = "bilibili" ]; then
        ytdlp_cmd="$ytdlp_cmd --cookies-from-browser chrome"
    fi

    # List available subtitles
    local sub_list=$($ytdlp_cmd "$url" 2>/dev/null || true)

    # Check for Chinese/AI subtitles
    local sub_lang=""
    if echo "$sub_list" | grep -qE '(zh-Hans|zh-CN|ai-zh|zh)'; then
        # Priority: ai-zh > zh-CN > zh-Hans > zh
        if echo "$sub_list" | grep -q 'ai-zh'; then
            sub_lang="ai-zh"
        elif echo "$sub_list" | grep -q 'zh-CN'; then
            sub_lang="zh-CN"
        elif echo "$sub_list" | grep -q 'zh-Hans'; then
            sub_lang="zh-Hans"
        elif echo "$sub_list" | grep -q 'zh'

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README says the skill 'automatically detects collection requests' and the metadata includes broad trigger phrases like '收藏', '记录下来', 'save', 'bookmark', and memory-retrieval phrases. In an agent setting, ambiguous activation can cause the skill to run on conversational text the user did not intend as a persistence action, leading to unintended saving of URLs, messages, or external content into local storage.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The usage and feature descriptions emphasize automatic saving and optional Obsidian sync, but they do not clearly warn users that external content, transcripts, extracted images, and search metadata may be written to local directories and potentially mirrored into another knowledge base. This can expose sensitive browsing history, copyrighted material, or private links if the user does not realize persistence and sync are occurring.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
87% confidence
Finding

The skill invokes filesystem, network, environment-backed, and shell-script behaviors but declares no explicit tool scope or allowed-tools boundary. That weakens least-privilege enforcement and makes it easier for the skill to access more capabilities than users or the orchestrator may expect.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill mandates persistent storage in the workspace and synchronized writing into an Obsidian vault without an explicit warning or consent checkpoint. Users may not realize content, metadata, and potentially sensitive URLs will be retained in multiple locations, increasing privacy and data-retention risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill requires downloading and caching external images locally for WeChat content without warning the user. This creates additional persistent copies of third-party content and may store sensitive or copyrighted material unexpectedly, while also expanding disk usage and forensic footprint.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill instructs writing to a handoff file for another workflow (temp/handoffs/collector-to-writing.md), which is outside the core collection store. Cross-workflow writes create an integrity boundary issue: one skill can silently influence another skill's inputs or downstream behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The default translation/styling policy for English content changes user content automatically without opt-in. While not a classic security flaw, it can cause integrity and consent issues by transforming saved material in a way users did not request, especially for quote accuracy or evidentiary use.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The eval cases describe broad activation and behavior expectations for a content-collection skill without clearly constraining when the skill should or should not take over. In an agent setting, ambiguous trigger boundaries can cause the skill to activate on loosely related prompts, leading to unintended web fetching, local script execution, file writes, or retrieval over a user's stored content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

All user-facing prompt examples and expected behavior are written exclusively in Chinese, with no indication that the skill supports other languages or that Chinese is an intentional, justified locale restriction. This can constitute a language/locale policy violation when a skill effectively assumes or enforces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The file’s operational headings, plan summary, assertions, and explanatory text are written in Chinese throughout, with no indication that language selection is optional or that the skill is intentionally limited to a Chinese-language context. Under the policy for natural-language violations, forcing a specific language without user opt-in is in scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The instruction You MUST generate all content in Chinese forces a specific language for all generated content. The file does not provide user opt-in, fallback behavior, or a documented region-specific justification, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The spec mandates downloading remote images with curl and copying them into local collection and Obsidian paths without requiring user confirmation, provenance checks, copyright handling, or clear disclosure that local files will be created. In a content-collection skill, this can cause unintended local file modifications, retention of copyrighted material, and storage of untrusted content from external sites, making the behavior materially risky even if not overtly malicious.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The instructions direct the agent to write into a local Obsidian vault and append to Daily Notes without requiring explicit user confirmation at the time of modification. Because this skill handles content collection and claims the downgrade should be transparent to the user, it can silently modify local files in a way users may not expect, increasing the risk of unauthorized persistence or accidental data pollution.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This markdown file contains all user-facing instructional content in Chinese and does not indicate that the user can choose another language. Under the policy rule for natural-language violations, forcing a specific language without opt-in is a reportable issue unless the locale constraint is clearly documented and justified.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger section says the skill applies to '所有内容类型' as long as there are at least 200 characters, which is a very broad activation condition without clear constraints or exclusions. This lack of specificity increases the chance of unintended invocation across ordinary content-processing workflows.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The language requirement states that all keywords must be generated in Chinese, which imposes a fixed locale behavior. The file does not offer user opt-in or explain a justified region-specific requirement, so this is a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.