Back to skill

Security audit

Tavily Extract

Security checks for vulnerabilities and agentic risk

Overview

The skill’s main extraction function is coherent, but its automatic authentication path reads local MCP tokens and can run an unpinned npm package, so it needs review before use.

Install only if you are comfortable with Tavily receiving the URLs, query text, and extracted-page request context. Prefer setting a Tavily API key explicitly instead of relying on cached MCP tokens, avoid internal or authenticated URLs, and treat the first-run OAuth path as higher risk until `mcp-remote` is pinned or otherwise integrity-controlled.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Error
Location
scripts/extract.sh:96
Finding

Unpinned npm Package Is Downloaded and Executed Automatically

Content
View full analysis

Vulnerability Details

File Location: scripts/extract.sh:96
Vulnerability Type: Unpinned dynamically retrieved dependency
Risk Level: High

Vulnerable Code:

bash
npx -y mcp-remote https://mcp.tavily.com/mcp </dev/null >/dev/null 2>&1 &

Technical Analysis

When no Tavily credential is available, the script invokes npx -y mcp-remote. The -y option suppresses the package-installation confirmation, while the absence of an exact package version or integrity constraint causes npm to resolve mutable package content at execution time.

Consequently, the code executed during a future invocation is not necessarily the code that was available when the skill was audited. If the npm package, its maintainer account, the relevant registry infrastructure, or a transitive dependency is compromised, attacker-controlled JavaScript can execute locally. This is a supply-chain risk rather than evidence that the current package is malicious.

Attack Path

  1. An attacker compromises the mcp-remote npm package, its publication account, registry resolution, or one of its dependencies.
  2. The attacker publishes a malicious version that is selected by the unversioned package reference.
  3. A user invokes scripts/extract.sh without TAVILY_API_KEY and without a usable cached Tavily token.
  4. Execution enters the OAuth branch at lines 92–119.
  5. npx -y downloads the currently resolved package without interactive approval or an integrity check.
  6. The package executes with the same operating-system identity and environment as the user running the skill.
  7. Malicious package code can access resources available to that user before continuing, disrupting, or impersonating the expected OAuth process.

Impact Assessment

Successful exploitation permits arbitrary code execution with the privileges of the invoking user. The affected scope can include readable local files, environment variables, ...[truncated 324 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace the unversioned package invocation with an exact, reviewed version rather than allowing npm to resolve the latest release.
  2. Commit and enforce a lockfile with integrity metadata for the package and all transitive dependencies.
  3. Prefer installing dependencies during a controlled build or installation phase instead of downloading executable code during normal skill execution.
  4. Consider bundling a reviewed OAuth client implementation with the skill so that runtime behavior cannot change independently of the audited artifact.
  5. Verify package provenance and signatures where supported, and use a trusted registry configured to reject unexpected sources.
  6. Remove automatic confirmation via -y, or require explicit informed user consent before any runtime package installation.
  7. Run the OAuth helper with minimal filesystem, environment, credential, and network access where sandboxing is available.
  8. Fail closed and display a clear manual authentication procedure if the pinned and integrity-verified helper is unavailable.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared purpose is simple remote content extraction, but the skill also describes searching ~/.mcp-auth/, reading cached OAuth tokens, decoding JWT metadata, and launching browser-based authentication. That undocumented access to local auth state materially expands the trust boundary and could expose credentials or surprise users who did not consent to local token discovery.

Content

No source excerpt is available for this finding.

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

The documentation instructs users to place an API key in ~/.claude/settings.json, which is a sensitive agent configuration location. Encouraging direct use of that file without declaring config-directory access or discussing security implications normalizes interaction with high-value local secrets and increases the blast radius if the skill or surrounding tooling is compromised.

Content

Scanner excerpt · SKILL.md (reported line 20)May include surrounding context.

Alternative: API Key

If you prefer using an API key, get one at https://tavily.com and add to ~/.claude/settings.json:

json
{
  "env": {

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill documents shell execution (./scripts/extract.sh) but declares no permissions or allowed-tools, so consumers are not given an explicit scope boundary for what the skill may execute. In combination with authentication steps and local file access described elsewhere, this increases the chance of over-privileged or surprising behavior beyond simple URL extraction.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill description emphasizes extraction functionality but does not clearly warn that user-supplied URLs and the resulting page content are transmitted to Tavily's external service. This omission can lead users to submit sensitive internal URLs or private content without understanding the third-party data exposure.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This example shows POSTing user-provided URLs and extracted content requests to Tavily's external API, which is a real data egress path. In an extraction skill this is expected behavior, but it is still security-relevant because users may unknowingly send sensitive targets or content to a third party.

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

Basic Extraction

bash
curl --request POST \
  --url https://api.tavily.com/extract \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

The explicit Tavily endpoint confirms a third-party network destination for skill data. The risk is contextual rather than inherently malicious: in a URL-extraction skill external transmission is expected, but it remains dangerous when users may provide sensitive URLs or expect local-only processing.

Content

Scanner excerpt · SKILL.md (reported line 56)May include surrounding context.

bash
curl --request POST \
  --url https://api.tavily.com/extract \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

This instance again demonstrates outbound transmission to Tavily for multi-URL extraction and query-focused processing. The more data-rich request format increases exposure because multiple URLs and query context can reveal user interests, internal resources, or sensitive research topics.

Content

Scanner excerpt · SKILL.md (reported line 68)May include surrounding context.

bash
curl --request POST \
  --url https://api.tavily.com/extract \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
86% confidence
Finding

The API reference section identifies the remote extraction endpoint, confirming that network egress is a core behavior. Even though this is expected for the skill category, it is still a real security concern because the documentation lacks proportional warnings about external processing and data sensitivity.

Content

Scanner excerpt · SKILL.md (reported line 86)May include surrounding context.

Endpoint

text
POST https://api.tavily.com/extract

Headers

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This later example repeats the same external transmission pattern to Tavily's API, reinforcing that the skill's operation depends on sending request data off-host. While aligned with the skill's purpose, repeated undocumented egress increases the risk of accidental disclosure if users assume extraction is local.

Content

Scanner excerpt · SKILL.md (reported line 135)May include surrounding context.

Single URL Extraction

bash
curl --request POST \
  --url https://api.tavily.com/extract \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

This example sends a specific URL to the external extract endpoint, creating the same egress concern as the earlier curl examples. In context it is normal product usage, but without warnings it could facilitate disclosure of sensitive browsing targets or internal documentation URLs.

Content

Scanner excerpt · SKILL.md (reported line 136)May include surrounding context.

bash
curl --request POST \
  --url https://api.tavily.com/extract \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

This targeted extraction example sends both URLs and a topical query to Tavily, increasing the amount of potentially sensitive contextual data disclosed externally. The skill context makes the behavior expected, but the lack of transparent warning makes accidental oversharing more likely.

Content

Scanner excerpt · SKILL.md (reported line 149)May include surrounding context.

bash
curl --request POST \
  --url https://api.tavily.com/extract \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

The JavaScript-heavy-page example again relies on third-party extraction, potentially against authenticated or app-like pages that may contain more sensitive rendered content. That context can make this more dangerous than simple public web scraping if users apply it to internal dashboards or session-dependent resources.

Content

Scanner excerpt · SKILL.md (reported line 166)May include surrounding context.

bash
curl --request POST \
  --url https://api.tavily.com/extract \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

Batch extraction to Tavily increases the volume of outbound data and can amplify accidental leakage by sending many URLs in one request. In a bulk-processing context, a single misuse can expose a larger set of internal or sensitive resources to the third-party service.

Content

Scanner excerpt · SKILL.md (reported line 180)May include surrounding context.

bash
curl --request POST \
  --url https://api.tavily.com/extract \
  --header "Authorization: Bearer $TAVILY_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script reads access tokens from the local MCP auth cache and exports them into the environment without prominently disclosing that credential material will be accessed. Even if the token is only used for Tavily, silent credential harvesting behavior increases the risk of unauthorized credential use and violates least surprise.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill scans the local ~/.mcp-auth cache for OAuth tokens and repurposes any matching Tavily token for this operation. Accessing credential material outside the explicit user input is broader than a simple URL extraction function and can unintentionally use or expose tokens the user did not intend this skill to consume.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

When no token is present, the script launches an interactive OAuth flow via npx and browser-based authentication, expanding behavior beyond the stated extraction purpose. This introduces extra code execution and authentication side effects that users may not expect from a content-extraction helper.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

The script executes npx -y mcp-remote without pinning an exact package version or integrity, which creates a supply-chain risk at runtime. If the package is compromised or a breaking/malicious version is published, the skill may execute attacker-controlled code during authentication.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The script sends user-provided URLs and optional query text to Tavily's remote service without an explicit privacy warning. This can expose sensitive internal URLs, document locations, or query context to a third party, which is especially relevant because the skill's purpose is to process arbitrary user-specified web targets.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

This is a real external transmission: the script POSTs user-supplied extraction arguments and an authorization token to https://mcp.tavily.com/mcp. In the context of a URL-extraction skill, network transmission is expected, but it still presents confidentiality and privacy risk if users supply sensitive URLs or queries.

Content

Scanner excerpt · scripts/extract.sh (reported line 151)May include surrounding context.

sh
}')

# Call Tavily MCP server via HTTPS (SSE response)
RESPONSE=$(curl -s --request POST \
    --url "https://mcp.tavily.com/mcp" \
    --header "Authorization: Bearer $TAVILY_API_KEY" \
    --header 'Content-Type: application/json' \

Static analysis

No suspicious patterns detected.