Back to skill

Security audit

Discovery

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly does what it says, but it gives agents enough authority to publish analysis results publicly by default and make billing changes without a hard confirmation gate.

Install only if you are comfortable sending selected datasets to Disco's hosted service. Use private visibility for confidential data, verify the agent explicitly sets visibility instead of relying on defaults, and do not expose a payment-enabled API key to autonomous workflows unless you can enforce human approval for purchases, subscriptions, and payment-method changes.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
server.py:533
Finding

Public Analysis Defaults Can Disclose Dataset Results Without Enforced Confirmation

Content
View full analysis

Vulnerability Details

File Location: server.py, lines 533–629
Vulnerability Type: Unsafe default causing unintended public disclosure
Risk Level: High

Vulnerable Code

python
async def discovery_analyze(
    target_column: str,
    file_ref: str | dict | None = None,
    analysis_depth: int = 2,
    visibility: str = "public",
    title: str | None = None,
    description: str | None = None,
    excluded_columns: str | list | None = None,
    column_descriptions: str | dict | None = None,
    author: str | None = None,
    source_url: str | None = None,
    use_llms: bool = False,
    api_key: str | None = None,
) -> str:
python
run_payload: dict = {
    "file": {
        "key": uploaded_file["key"],
        "name": uploaded_file.get("name", "dataset"),
        "size": uploaded_file.get("size", 0),
        "fileHash": uploaded_file.get("fileHash", ""),
    },
    "columns": columns,
    "targetColumn": target_column,
    "analysisDepth": analysis_depth,
    "isPublic": visibility == "public",
    "useLlms": use_llms,
}

Technical Analysis

The discovery_analyze tool defaults visibility to "public". That value is converted directly into "isPublic": True in the analysis request. Public analyses are documented as being published to a public gallery.

SKILL.md lines 62–66 instruct the agent to ask whether the analysis should be public or private. However, this is a documentation-level instruction rather than a constraint enforced by the executable entry point. A direct MCP caller, an automated workflow, or an agent that omits the optional argument can bypass the documented selection step.

The tool's API therefore treats the absence of an explicit privacy decision as consent to public disclosure. This is unsafe for datasets whose analysis results, column metadata, summaries, or discovered patterns are confidential.

Attack Path

  1. A user uploads a confidential dataset through discovery_upload.

...[truncated 1074 chars]

Remediation
View remediation

Remediation Suggestions

  1. Change the default to visibility="private" so omission fails closed.
  2. Prefer making visibility mandatory rather than assigning any implicit publication state.
  3. Require an explicit parameter such as publish_publicly=True before creating a public analysis.
  4. For agent-driven workflows, implement a short-lived confirmation token issued only after the user is shown that the results will be publicly published.
  5. Reject public submissions if the confirmation token is missing, expired, or does not match the uploaded file and analysis parameters.
  6. Present the publication destination and data classification warning immediately before confirmation.
  7. Add tests verifying that omitted visibility never produces "isPublic": True and that direct MCP calls cannot bypass the confirmation requirement.

T09 · Insecure Skill Coding Practices

Error
Location
server.py:902
Finding

Credit Purchases and Paid Subscriptions Lack a Code-Enforced User Confirmation Gate

Content
View full analysis

Vulnerability Details

File Location: server.py, lines 902–969
Vulnerability Type: Unconfirmed financial transaction and recurring billing change
Risk Level: High

Vulnerable Code

python
@mcp.tool(
    annotations=ToolAnnotations(
        title="Purchase Credits", destructiveHint=True, idempotentHint=False
    )
)
async def discovery_purchase_credits(packs: int = 1, api_key: str | None = None) -> str:
    """Purchase Disco credit packs using a stored payment method.

    Credits cost $0.10 each, sold in packs of 100 ($10/pack). Credits are used
    for private analyses (public analyses are free). Requires a payment method
    on file — use discovery_add_payment_method first.

    Args:
        packs: Number of 100-credit packs to purchase. Default 1.
        api_key: Disco API key (disco_...). Optional if DISCOVERY_API_KEY env var is set.
    """
    resolved_key = _resolve_api_key(api_key)
    if not resolved_key:
        return json.dumps(
            {"error": "API key required. Pass api_key or set DISCOVERY_API_KEY env var."}
        )

    result = await _dashboard_request(
        "POST",
        "/api/account/credits/purchase",
        api_key=resolved_key,
        json_body={"packs": packs},
    )
    return json.dumps(result, indent=2)
python
@mcp.tool(
    annotations=ToolAnnotations(
        title="Subscribe to Plan", destructiveHint=True, idempotentHint=True
    )
)
async def discovery_subscribe(plan: str, api_key: str | None = None) -> str:
    """Subscribe to or change your Disco plan.

    Available plans:
    - "free_tier": Explorer — free, 10 credits/month
    - "tier_1": Researcher — $49/month, 500 credits/month
    - "tier_2": Team — $199/month, 2000 credits/month

    Paid plans require a payment method on file. Credits roll over on paid plans.

    Args:
        plan: Plan tier ID ("free_tier", "tier_1", or "tier_2").
        api_key: Disco API key (disco_...). Optional if DISCOVERY_API_KEY env var is 
...[truncated 2614 chars]
Remediation
View remediation

Remediation Suggestions

  1. Implement a two-phase transaction flow:
    • First return an immutable quote containing the exact amount, currency, quantity, plan, and recurrence.
    • Then require a short-lived, single-use confirmation token tied to that quote.
  2. Require explicit user approval after displaying the final charge and before issuing the purchase or subscription request.
  3. Reject direct transaction calls that do not include a valid confirmation token.
  4. Add strict server-side limits for the number of credit packs per transaction and cumulative spending over a defined period.
  5. Validate subscription plan values against an explicit allowlist before sending the request.
  6. Require renewed confirmation when changing from a free plan to a paid plan or when increasing recurring cost.
  7. Preserve destructiveHint=True, but do not rely on tool annotations as the authorization mechanism.
  8. Add audit logs for quote creation, confirmation, purchase, and subscription changes without logging API keys or payment tokens.
  9. Add tests proving that transaction endpoints are never called when confirmation is absent, expired, reused, or mismatched with the requested amount.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (46)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared purpose is data pattern discovery, but the skill also includes account creation, authentication, billing, payment-method handling, and dataset upload workflows. This broadens the trust boundary substantially beyond what a user would infer from the description, increasing the chance of surprise credential handling, financial actions, or sensitive data transfer.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The public/private choice is mentioned, but the risk is easy to miss despite public runs publishing results and always using LLMs. This can lead to irreversible disclosure of dataset-derived information or sensitive inferences if a user selects the free public option without fully understanding the privacy consequences.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill exposes credit purchase and subscription-changing endpoints that can trigger financial transactions beyond the core data-analysis purpose. In an agent context, this creates a meaningful risk of unauthorized or accidental charges if an LLM is induced to upgrade plans or buy credits without strong user approval boundaries.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

Exposing payment-method attachment in a skill that primarily performs dataset analysis expands the blast radius from analysis to financial account modification. Even if Stripe tokenization keeps raw card data away from the service, an agent could still attach a payment method token and enable later charges, which is especially risky in autonomous or prompt-influenced workflows.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill metadata says this tool is for tabular data discovery, but the server exposes a much broader surface including signup, login, payment method attachment, credit purchase, and subscription changes. This scope expansion is dangerous because an agent or user may invoke financially or identity-impacting operations that are not clearly justified by the advertised purpose, increasing the risk of surprise account actions and abuse.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The code includes tools to add payment methods, purchase credits, and subscribe to plans, which go well beyond a discovery-analysis skill. These capabilities enable direct monetary impact and can be triggered through the same agent interface, creating a real risk of unauthorized or accidental charges if an agent misinterprets user intent or is prompt-injected.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 41)May include surrounding context.

bash
# Step 1: request verification code (no password, no card)
curl -X POST https://disco.leap-labs.com/api/signup \
  -H "Content-Type: application/json" \
  -d '{"email": "you@example.com"}'

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The quickstart demonstrates sending a local dataset to a remote hosted service without an immediate, prominent warning that analysis requires upload to Leap Labs infrastructure and that some runs may publish data/results depending on the visibility setting. This creates a realistic risk of accidental disclosure of sensitive or regulated data because users often copy quickstart examples before reading later caveats.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill exposes capabilities involving environment secrets and outbound network access but does not declare any explicit tool scope or permission boundaries. In an agent ecosystem, this weakens least-privilege controls and makes it easier for the skill to access API keys and transmit user data without clear policy gating.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The integration guidance encourages sending datasets to a remote service and highlights convenience paths for URL and local-file upload, but it does not front-load a strong privacy warning about remote processing. Users may unintentionally submit regulated, proprietary, or personal data without understanding that it leaves the local environment.

Content

No source excerpt is available for this finding.

Unbounded Output

Medium
Category
Output Handling
Confidence
75% confidence
Finding

Output size or generation rate is not bounded. Unbounded output enables denial-of-service through resource exhaustion, log flooding, or context-window stuffing.

Content

Scanner excerpt · SKILL.md (reported line 18)May include surrounding context.

md
## Integration Options

- **MCP server** — remote server at `https://disco.leap-labs.com/mcp`, no install required. Best for datasets at a URL.
- **Python SDK** — `pip install discovery-engine-api`. **Use this for local files of any size.** Runs on your machine and streams files directly — no base64, no size limits.

**Quick rule:** if the data is at a URL, use `file_url` in `discovery_upload`. If it's a local file, use the Python SDK — or if Python isn't available, upload directly via the presign API and pass the result to `discovery_analyze`. Don't use `file_content` (base64) unless the file is already in memory and tiny.

Unbounded Output

Medium
Category
Output Handling
Confidence
75% confidence
Finding

Output size or generation rate is not bounded. Unbounded output enables denial-of-service through resource exhaustion, log flooding, or context-window stuffing.

Content

Scanner excerpt · server.py (reported line 368)May include surrounding context.

python
## Integration Options

- **MCP server** — remote server at `https://disco.leap-labs.com/mcp`, no install required. Best for datasets at a URL.
- **Python SDK** — `pip install discovery-engine-api`. **Use this for local files of any size.** Runs on your machine and streams files directly — no base64, no size limits.

**Quick rule:** if the data is at a URL, use `file_url` in `discovery_upload`. If it's a local file, use the Python SDK — or if Python isn't available, upload directly via the presign API and pass the result to `discovery_analyze`. Don't use `file_content` (base64) unless the file is already in memory and tiny.

Unbounded Output

Medium
Category
Output Handling
Confidence
75% confidence
Finding

Output size or generation rate is not bounded. Unbounded output enables denial-of-service through resource exhaustion, log flooding, or context-window stuffing.

Content

Scanner excerpt · SKILL.md (reported line 227)May include surrounding context.

result = await engine.discover("data.csv", target_column="outcome")

text

Handles upload, polling, and results in one call. No size limit. See the **Python SDK** section for full documentation.

---

Unbounded Output

Medium
Category
Output Handling
Confidence
75% confidence
Finding

Output size or generation rate is not bounded. Unbounded output enables denial-of-service through resource exhaustion, log flooding, or context-window stuffing.

Content

Scanner excerpt · SKILL.md (reported line 242)May include surrounding context.

result = await engine.discover("data.csv", target_column="outcome")

text

Handles upload, polling, and results in one call. No size limit. See the **Python SDK** section for full documentation.

---

Unbounded Output

Medium
Category
Output Handling
Confidence
75% confidence
Finding

Output size or generation rate is not bounded. Unbounded output enables denial-of-service through resource exhaustion, log flooding, or context-window stuffing.

Content

Scanner excerpt · SKILL.md (reported line 271)May include surrounding context.

result = await engine.discover("data.csv", target_column="outcome")

text

Handles upload, polling, and results in one call. No size limit. See the **Python SDK** section for full documentation.

---

External Transmission

Medium
Category
Data Exfiltration
Confidence
92% confidence
Finding

This flow sends dataset metadata and then the dataset itself to a remote service using a presigned upload path. In the context of a data-analysis skill this is functional behavior, but it is still a real data-exfiltration risk if users are not clearly warned or if sensitive files are uploaded by an agent automatically.

Content

Scanner excerpt · SKILL.md (reported line 252)May include surrounding context.

bash
# 1. Get a presigned upload URL
curl -X POST https://disco.leap-labs.com/api/data/upload/presign \
  -H "Authorization: Bearer disco_..." \
  -H "Content-Type: application/json" \
  -d '{"fileName": "data.csv", "contentType": "text/csv", "fileSize": 1048576}'

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
88% confidence
Finding

The documentation states that signup may directly return an API key with no verification needed in the common case. If implemented as described, unauthenticated email-only key issuance lowers assurance of account ownership and could enable unauthorized account creation or abuse tied to arbitrary email addresses.

Content

Scanner excerpt · SKILL.md (reported line 412)May include surrounding context.

-d '{"email": "agent@example.com"}'

text

Response — direct key (no verification needed):
```json
{"key": "disco_...", "key_id": "...", "organization_id": "...", "tier": "free_tier", "credits": 10}

Cloud Storage Exfiltration

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

The skill uploads local files to cloud storage as part of its normal operation. Because datasets may contain proprietary or personal information, this creates a material exfiltration pathway whose danger depends on informed consent and strict scoping; the current document emphasizes convenience more than risk.

Content

Scanner excerpt · SKILL.md (reported line 481)May include surrounding context.

python
# Upload once and get the server's parsed column list
upload = await engine.upload_file(file="data.csv", title="My dataset")
print(upload["columns"])   # [{"name": "col1", "type": "continuous", ...}, ...]
print(upload["rowCount"])  # e.g., 5000

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 728)May include surrounding context.

Or via REST:

bash
curl https://disco.leap-labs.com/api/account \
  -H "Authorization: Bearer disco_..."
# → { "stripe_publishable_key": "pk_live_...", "stripe_customer_id": "cus_...", "credits": {...}, ... }

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The API description states that public runs are free and results are published, but the top-level workflow and upload guidance do not prominently warn users about privacy consequences when uploading datasets. In an agent setting, this can lead to accidental disclosure of sensitive or proprietary data because the default path may appear convenient and low-risk.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest focuses on automated pattern discovery in tabular data, but the documented operations also include listing and deleting API keys plus retrieving user/session profile data. These identity and credential-management functions are materially broader than the stated analysis-only behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a skill for discovering statistically validated patterns in tabular data and returning structured findings. This OpenAPI spec also exposes account status, plan listing, subscription changes, payment-method attachment, and credit purchases, which are commercial account-management capabilities rather than data-discovery functionality.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The primary quick-start immediately shows uploading and analyzing a local dataset, but does not foreground that public runs publish results and that datasets are transmitted to an external service. In an agent setting, this increases the chance that users or automated systems send sensitive or regulated data off-host without informed consent or proper privacy review.

Content

No source excerpt is available for this finding.

Cloud Storage Exfiltration

Medium
Category
Data Exfiltration
Confidence
55% confidence
Finding

Data is uploaded to cloud storage (S3 / GCS / Azure Blob). This may be a legitimate backup or exfiltration to an external bucket. Manual review is recommended.

Content

Scanner excerpt · docs/python-sdk.md (reported line 171)May include surrounding context.

python
# Upload once and get the server's parsed column list
upload = await engine.upload_file(file="data.csv", title="My dataset")
# upload["file"]    -> {"key": "uploads/abc123.csv", "name": "data.csv",
#                        "size": 1048576, "fileHash": "sha256:..."}
# upload["columns"] -> [{"name": "col1", "type": "continuous", ...}, ...]

Cloud Storage Exfiltration

Medium
Category
Data Exfiltration
Confidence
55% confidence
Finding

Data is uploaded to cloud storage (S3 / GCS / Azure Blob). This may be a legitimate backup or exfiltration to an external bucket. Manual review is recommended.

Content

Scanner excerpt · docs/python-sdk.md (reported line 363)May include surrounding context.

python
# Upload once and get the server's parsed column list
upload = await engine.upload_file(file="data.csv", title="My dataset")
# upload["file"]    -> {"key": "uploads/abc123.csv", "name": "data.csv",
#                        "size": 1048576, "fileHash": "sha256:..."}
# upload["columns"] -> [{"name": "col1", "type": "continuous", ...}, ...]

Static analysis

No suspicious patterns detected.