Back to skill

Security audit

Idea Vault

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly matches a local note vault, but it has unsafe attachment handling and under-disclosed network and persistence behavior that should be reviewed before install.

Install only if you are comfortable with the skill reading recent chat content, making outbound requests, and writing a persistent local vault. Before broad use, the publisher should sanitize attachment filenames, restrict attachment URL hosts and response sizes, allowlist the TranscriptAPI origin, document or remove channel-monitor, narrow the trigger, and pin dependencies.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/idea_vault.py:569
Finding

Unrestricted Attachment Downloads Enable Server-Side Request Forgery

Content
View full analysis
None: ensure_dir(os.path.dirname(out_path)) headers = {"User-Agent": "Mozilla/5.0"} r = requests.get(url, headers=headers, timeout=60) r.raise_for_status() with open(out_path, "wb") as f: f.write(r.content) ``` The function is invoked with URLs originating from captured message attachments: ```python assets = capture.get("assets") or [] if assets: mm = dt.datetime.now(dt.timezone.utc).strftime("%m") assets_dir = os.path.join(vault_dir, "assets", year, mm) ensure_dir(assets_dir) for a in assets: url_a = a.get("url") if not url_a: continue fn = a.get("filename") or slugify(url_a.split("/")[-1]) out_path = os.path.join(assets_dir, fn) try: download_asset(url_a, out_path) assets_out.append({"url": url_a, "path": out_path, "filename": fn}) except Exception as e: assets_out.append({"url": url_a, "error": str(e), "filename": fn}) ``` ### Technical Analysis Attachment URLs extracted from chat input are passed directly to `requests.get`. The implementation does not restrict destination hosts, resolve and validate IP addresses, reject non-public networks, constrain redirects, or limit the downloaded response size. Although `requests` only supports particular URL schemes, an attacker can still supply an HTTP or HTTPS URL pointing to: - Loopback services such as `127.0.0.1` or `::1` - RFC1918 private networks - Link-local addresses and cloud metadata endpoints - Internal DNS names - Public endpoints that redirect to restricted destinations Redirects are followed by default. Consequently, validating only the ini ...[truncated 1562 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/idea_vault.py:1489
Finding

Unsanitized Attachment Filenames Permit Path Traversal and Arbitrary File Overwrite

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/idea_vault.py:34
Finding

Unrestricted Transcript API Base URL Can Exfiltrate the API Credential

Content
View full analysis
Tuple[dict, List[TranscriptSeg], Optional[dict]]: """Fetch transcript from TranscriptAPI. Returns: (raw_json, segs, metadata) """ url = TRANSCRIPTAPI_BASE.rstrip('/') + '/youtube/transcript' headers = {"Authorization": f"Bearer {api_key}", "User-Agent": "OpenClaw-IdeaVault/1.0"} params = { "video_url": video_url_or_id, "format": "json", "include_timestamp": "true", "include_metadata": "true", } r = requests.get(url, headers=headers, params=params, timeout=40) ``` The channel-monitor endpoint has the same behavior: ```python def transcriptapi_channel_latest(channel: str, api_key: str) -> dict: """Fetch latest channel uploads (up to 15) via TranscriptAPI's RSS-backed endpoint.""" url = TRANSCRIPTAPI_BASE.rstrip('/') + '/youtube/channel/latest' headers = {"Authorization": f"Bearer {api_key}", "User-Agent": "OpenClaw-IdeaVault/1.0"} r = requests.get(url, headers=headers, params={"channel": channel}, timeout=40) ``` ### Technical Analysis The TranscriptAPI bearer key is intentionally transmitted to the declared external service under the default configuration, which is necessary for authenticated transcript retrieval. However, the destination origin can be replaced by the `IDEA_VAULT_TRANSCRIPTAPI_BASE` environment variable without validation. If the environment is misconfigured or influenced by an attacker, the Authorization header is sent to an arbitrary HT ...[truncated 1489 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
requirements.txt:1
Finding

Dependency Specification Accepts Unreviewed Future Releases

Content
View full analysis
=2.31.0 ``` The documented setup installs directly from this specification: ```bash python3 -m pip install -r requirements.txt ``` ### Technical Analysis The dependency name is legitimate, and the audit found no typo-squatted package, malicious alternate package index, or known malicious component. However, the lower-bound-only constraint accepts any future `requests` release selected by the configured package index. There is no lock file, upper bound, or hash verification. Consequently, installations performed at different times may receive materially different dependency code that was not part of this audit. Installation also relies on the user’s package-index configuration and transitive dependency resolution. This is a supply-chain hardening weakness rather than evidence that the current dependency is malicious. ### Attack Path 1. A user follows the documented setup instructions. 2. `pip` resolves `requests>=2.31.0` and its transitive dependencies using the configured package index. 3. A later, compromised, or otherwise unreviewed compatible release is selected. 4. Package installation executes the selected package’s build or installation behavior and makes its runtime code available to the Skill. 5. The Skill subsequently imports and executes that dependency. ### Impact Assessment If an accepted dependency release or configured package source is compromised, package code can execute with the permissions of the installing user and later with the permissions of the Skill process. This can affect local files, environment-held credentials, and network-accessible resources. No current compromise was established. The immediate impact is reduced build reproducibility and an unnecessarily broad future supply-chain trust boundary. ]]>
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
Findings (37)

Tainted flow: 'url' from os.environ.get (line 451, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 244)May include surrounding context.

python
"User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120 Safari/537.36",
        "Accept-Language": f"{lang},en;q=0.9",
    }
    r = requests.get(url, headers=headers, timeout=30)
    r.raise_for_status()
    html_text = r.text

Tainted flow: 'url' from os.environ.get (line 303, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 267)May include surrounding context.

python
url = f"https://www.youtube.com/oembed?url=https://www.youtube.com/watch?v={video_id}&format=json"
    headers = {"User-Agent": "Mozilla/5.0", "Accept": "application/json"}
    try:
        r = requests.get(url, headers=headers, timeout=20)
        if r.status_code != 200:
            return None
        return r.json()

Tainted flow: 'url' from os.environ.get (line 303, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 311)May include surrounding context.

python
"User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120 Safari/537.36",
        "Accept-Language": "en,en;q=0.9",
    }
    r = requests.get(url, headers=headers, timeout=30)
    r.raise_for_status()
    try:
        return r.json()

Tainted flow: 'url' from os.environ.get (line 303, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
98% confidence
Finding

The TranscriptAPI base URL is taken from IDEA_VAULT_TRANSCRIPTAPI_BASE and used directly for outbound requests with an Authorization bearer token. If an attacker can influence environment configuration, they can redirect transcript requests and exfiltrate the API key and requested video metadata to an arbitrary host.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 428)May include surrounding context.

python
"include_timestamp": "true",
        "include_metadata": "true",
    }
    r = requests.get(url, headers=headers, params=params, timeout=40)
    if r.status_code != 200:
        raise RuntimeError(f"TranscriptAPI error {r.status_code}: {r.text[:200]}")
    data = r.json()

Tainted flow: 'url' from os.environ.get (line 303, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
98% confidence
Finding

This channel-monitor request also uses the environment-controlled TranscriptAPI base and sends the bearer token to that destination. A malicious or compromised deployment configuration could turn this into credential leakage and unauthorized external polling against attacker infrastructure.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 453)May include surrounding context.

python
"""Fetch latest channel uploads (up to 15) via TranscriptAPI's RSS-backed endpoint."""
    url = TRANSCRIPTAPI_BASE.rstrip('/') + '/youtube/channel/latest'
    headers = {"Authorization": f"Bearer {api_key}", "User-Agent": "OpenClaw-IdeaVault/1.0"}
    r = requests.get(url, headers=headers, params={"channel": channel}, timeout=40)
    if r.status_code != 200:
        raise RuntimeError(f"TranscriptAPI channel/latest error {r.status_code}: {r.text[:200]}")
    return r.json()

Tainted flow: 'url' from os.environ.get (line 303, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
99% confidence
Finding

The skill downloads attachment URLs from message data without validating destination host, scheme, or content type. If attachment metadata can be attacker-controlled or spoofed, this creates an SSRF primitive and can be used to make the agent fetch internal resources or large malicious payloads.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 569)May include surrounding context.

python
def download_asset(url: str, out_path: str) -> None:
    ensure_dir(os.path.dirname(out_path))
    headers = {"User-Agent": "Mozilla/5.0"}
    r = requests.get(url, headers=headers, timeout=60)
    r.raise_for_status()
    with open(out_path, "wb") as f:
        f.write(r.content)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The public description frames the skill as a simple vault for saving and querying notes, but the body reveals additional behaviors including transcript retrieval, external fetching, local asset storage, annotation/editing, and other richer operations. This mismatch undermines informed consent and can cause users or platform policy checks to underestimate the skill's access to external content, local storage, and persistent modification of saved data.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 28)May include surrounding context.

Environment setup (recommended):

bash
cp .env.example .env
# then set IDEA_VAULT_TRANSCRIPTAPI_KEY and IDEA_VAULT_DIR in your shell/env manager

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 78)May include surrounding context.

md
python3 ./scripts/idea_vault.py extract --user-id <author.id> --fallback-messages 30 < messages.json > capture.json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 86)May include surrounding context.

md
python3 ./scripts/idea_vault.py extract --user-id <author.id> --fallback-messages 30 < messages.json > capture.json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 92)May include surrounding context.

md
python3 ./scripts/idea_vault.py extract --user-id <author.id> --fallback-messages 30 < messages.json > capture.json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 98)May include surrounding context.

md
python3 ./scripts/idea_vault.py extract --user-id <author.id> --fallback-messages 30 < messages.json > capture.json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 99)May include surrounding context.

md
python3 ./scripts/idea_vault.py extract --user-id <author.id> --fallback-messages 30 < messages.json > capture.json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 100)May include surrounding context.

md
python3 ./scripts/idea_vault.py extract --user-id <author.id> --fallback-messages 30 < messages.json > capture.json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 106)May include surrounding context.

md
python3 ./scripts/idea_vault.py extract --user-id <author.id> --fallback-messages 30 < messages.json > capture.json

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill declares substantial capabilities—environment access, file read/write, network, and shell-driven setup/CLI execution—without an explicit tool scope or permission boundary in the manifest. That makes the runtime trust model opaque and increases the chance the skill is invoked with broader authority than users or operators expect, especially given it processes chat content and external URLs.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Allowing activation on the plain word 'vault' is overly broad and can trigger the skill during ordinary conversation rather than a deliberate command. In this skill's context, accidental activation is more dangerous because it may read recent chat messages, extract content, make external requests for URLs, and persist notes or links without clear user intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Multiple request paths hard-code English preferences, including a default lang='en', Accept-Language headers, and English-only subtitle selection, which imposes a specific language choice regardless of user preference. This is a natural-language policy issue because the skill does not offer a locale choice or document that it is intentionally English-only.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Transcript fetching sends user-supplied YouTube links and related metadata to external services, including TranscriptAPI fallback, without any explicit disclosure or consent mechanism in the code. This is a privacy and transparency problem in an agent workflow because users may not expect their saved links to be transmitted to third parties.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 515)May include surrounding context.

python
# If supported, provide a JS runtime (helps with newer YouTube extraction flows)
    try:
        help_out = subprocess.run([ytdlp, '--help'], capture_output=True, text=True).stdout
        if '--js-runtimes' in (help_out or ''):
            node_path = shutil.which('node')
            if node_path:

Tainted flow: 'ytdlp' from os.environ.get (line 495, credential/environment) → subprocess.run (code execution)

Medium
Category
Data Flow
Confidence
97% confidence
Finding

This line executes the yt-dlp binary selected from an environment override before any other safeguards, so a hostile IDEA_VAULT_YTDLP_BIN immediately yields arbitrary code execution. The danger is amplified because it may run automatically during transcript fetching for user-submitted URLs.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 515)May include surrounding context.

python
# If supported, provide a JS runtime (helps with newer YouTube extraction flows)
    try:
        help_out = subprocess.run([ytdlp, '--help'], capture_output=True, text=True).stdout
        if '--js-runtimes' in (help_out or ''):
            node_path = shutil.which('node')
            if node_path:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 528)May include surrounding context.

python
cmd.append(video_url)

    p = subprocess.run(cmd, cwd=out_dir, capture_output=True, text=True)
    if p.returncode != 0:
        tail = (p.stderr or p.stdout or '').strip()[-900:]
        raise RuntimeError(f"yt-dlp failed: {tail}")

Tainted flow: 'cmd' from os.environ.get (line 503, credential/environment) → subprocess.run (code execution)

Medium
Category
Data Flow
Confidence
97% confidence
Finding

The yt-dlp executable path can come from IDEA_VAULT_YTDLP_BIN, and the script executes it directly. If an attacker can control that environment variable or PATH in the agent runtime, they can cause arbitrary code execution under the agent's privileges.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 528)May include surrounding context.

python
cmd.append(video_url)

    p = subprocess.run(cmd, cwd=out_dir, capture_output=True, text=True)
    if p.returncode != 0:
        tail = (p.stderr or p.stdout or '').strip()[-900:]
        raise RuntimeError(f"yt-dlp failed: {tail}")

Tainted flow: 'transcript_json' from os.environ.get (line 850, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 877)May include surrounding context.

python
if segs:
                    transcript_available = True
                    write_transcript_txt(segs, transcript_txt)
                    with open(transcript_json, "w", encoding="utf-8") as f:
                        json.dump(cj, f)

        # Fallback 1: TranscriptAPI (reliable in datacenter environments)

Tainted flow: 'transcript_json' from os.environ.get (line 850, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/idea_vault.py (reported line 888)May include surrounding context.

python
if segs:
                    transcript_available = True
                    write_transcript_txt(segs, transcript_txt)
                    with open(transcript_json, "w", encoding="utf-8") as f:
                        json.dump(cj, f)

        # Fallback 1: TranscriptAPI (reliable in datacenter environments)

Static analysis

No suspicious patterns detected.