Back to skill

Security audit

DeepRead Medical Records

Security checks for vulnerabilities and agentic risk

Overview

This skill is coherent for medical-document extraction, but it handles highly sensitive medical records by uploading them to external services without enough compliance and downstream-provider warnings.

Review this carefully before installing for real patient data. Only use it where uploading PHI to DeepRead and any BYOK model provider is authorized, covered by required agreements, and acceptable under your retention, logging, residency, and compliance requirements. Prefer synthetic or already de-identified test files first, use environment variables for keys, and harden the sample download code before production use.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:262
Finding

Unvalidated Response-Controlled File Download

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:262-272
Vulnerability Type: Unvalidated remote URL retrieval
Risk Level: Medium

python
result = requests.get(f"{BASE}/v1/pii/{redact_id}", headers=headers).json()
if result["status"] == "completed":
    report = result["report"]
    print(f"Redacted {report['total_redactions']} PII instances")
    for pii_type, info in report["pii_detected"].items():
        print(f"  {pii_type}: {info['count']} found")

    # Download redacted file
    pdf = requests.get(result["redacted_file_url"]).content
    with open("patient_record_redacted.pdf", "wb") as f:
        f.write(pdf)

Technical Analysis

The example treats redacted_file_url from the remote API response as trusted and passes it directly to requests.get(). It does not validate the URL scheme, hostname, port, redirects, content type, response size, or PDF signature. It also omits request timeouts and HTTP status validation.

If the DeepRead service, an upstream component, or its response channel is compromised, an attacker could provide a URL targeting an internal network service or an attacker-controlled host. Automatic redirect handling can also bypass a superficial hostname check unless every redirect destination is validated. Reading the entire response through .content allows an oversized response to exhaust client memory or storage.

Attack Path

  1. An attacker compromises or manipulates the API response associated with a redaction job.
  2. The response sets redacted_file_url to an attacker-selected URL, an internal service address, or an endpoint returning an oversized body.
  3. The client performs the request without validating the destination or limiting the response.
  4. The request may reach resources available from the client network, disclose request metadata, or consume excessive memory and disk space.
  5. The unverified response is saved locally with a .pdf extension r ...[truncated 682 chars]
Remediation
View remediation

Remediation Suggestions

  • Parse the returned URL and require the https scheme.
  • Maintain an explicit allowlist of trusted DeepRead download hostnames and permitted ports.
  • Disable redirects or validate the scheme, hostname, and port of every redirect destination.
  • Reject loopback, link-local, private, and reserved IP destinations after DNS resolution where appropriate.
  • Add connection and read timeouts.
  • Stream the response and enforce a maximum file-size limit instead of loading the entire body into memory.
  • Call raise_for_status() before processing the response.
  • Validate the expected content type and verify the PDF magic bytes before saving.
  • Generate a controlled output path and avoid replacing an existing file unless explicitly authorized.

Example hardening should use a pattern equivalent to:

python
response = requests.get(
    validated_url,
    timeout=(5, 30),
    allow_redirects=False,
    stream=True,
)
response.raise_for_status()

T09 · Insecure Skill Coding Practices

Note
Location
SKILL.md:151
Finding

Sample Encourages Hardcoded Production API Credentials

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:151-153
Vulnerability Type: Hardcoded secret pattern
Risk Level: Low

python
API_KEY = "sk_live_YOUR_KEY"
BASE = "https://api.deepread.tech"
headers = {"X-API-Key": API_KEY}

Technical Analysis

The included value is a placeholder rather than an exposed working credential. However, the example instructs users to replace a source-code string with a live API key even though the skill metadata and setup instructions identify DEEPREAD_API_KEY as the intended environment variable.

Embedding credentials in source files increases the chance that they will be committed to version control, included in backups, copied into support requests, or shared with other users. The sk_live_ format also suggests that users are expected to insert a production credential.

Attack Path

  1. A user replaces sk_live_YOUR_KEY with a valid DeepRead API key.
  2. The script is committed to a repository, archived, logged, or shared.
  3. An unauthorized party obtains the embedded key.
  4. The party authenticates to the DeepRead API as the affected account.
  5. The key may be used to consume account quota or access API capabilities available to that credential.

Impact Assessment

Exposure would grant the attacker the API privileges assigned to the compromised key. Potential effects include unauthorized API use, quota or billing consumption, and access to account-scoped operations supported by the service. The available documentation does not establish whether a key can retrieve existing jobs or medical records, so such access cannot be assumed.

No active secret is present in the audited project; the risk arises from the insecure credential-handling pattern promoted by the example.

Remediation
View remediation

Remediation Suggestions

Replace the source-code credential with environment-based loading and fail safely when the variable is absent:

python
import os

API_KEY = os.environ["DEEPREAD_API_KEY"]
BASE = "https://api.deepread.tech"
headers = {"X-API-Key": API_KEY}

Additionally:

  • Do not print, log, or include the API key in exception messages.
  • Keep local secret files outside version control and provide an appropriately restrictive .gitignore.
  • Prefer scoped, revocable credentials when the service supports them.
  • Document immediate key rotation if a credential is committed or otherwise disclosed.
  • Use automated secret scanning in repositories and CI pipelines.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (3)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill explicitly instructs users to upload medical records to a third-party API, but it does not prominently warn that PHI leaves the local environment for external processing. In a healthcare context, this omission is dangerous because users may assume 'HIPAA-compliant redaction' means data stays local or is safe to send without verifying legal, contractual, and regulatory requirements such as BAAs, consent, retention, and processor handling.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

The cURL example operationalizes external transmission of a patient PDF to api.deepread.tech, which is expected functionality but still represents a real security/privacy risk in this medical-data context. Because the content is specifically patient records, any misuse, misunderstanding, or insecure deployment could expose PHI to a third party without adequate notice, approval, or safeguards.

Content

Scanner excerpt · SKILL.md (reported line 229)May include surrounding context.

cURL

bash
curl -s -X POST https://api.deepread.tech/v1/process \
  -H "X-API-Key: $DEEPREAD_API_KEY" \
  -F "file=@patient_record.pdf" \
  -F 'schema={"type":"object","properties":{"patient_name":{"type":"string","description":"Patient full name"},"mrn":{"type":"string","description":"Medical record number"},"diagnoses":{"type":"array","items":{"type":"object","properties":{"code":{"type":"string","description":"ICD-10 code"},"description":{"type":"string"}}},"description":"Diagnoses"},"medications":{"type":"array","items":{"type":"object","properties":{"name":{"type":"string"},"dosage":{"type":"string"},"frequency":{"type":"string"}}},"description":"Medications"}}}'

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The BYOK section says processing may route through the user's OpenAI, Google, or OpenRouter provider, but it fails to clearly warn that medical records and other sensitive data may then be disclosed to additional third parties beyond DeepRead. This increases privacy and compliance risk because users may unknowingly expand the data-sharing chain and assume BYOK only affects billing rather than downstream PHI exposure.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.