Back to skill

Security audit

Medical Scribe (Dictation)

Security checks for vulnerabilities and agentic risk

Overview

This medical-scribe skill has a real review concern because it can send sensitive clinical dictation to external LLM providers while its own risk table says there are no external API calls.

Review carefully before installing or using this with real patient data. Use only the local/rule-based path unless your organization has approved the OpenAI or Anthropic data flow for PHI, and avoid saving generated notes to shared or insecure directories. Treat all outputs as drafts requiring clinician review, especially diagnoses, medications, allergies, and treatment plans.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/main.py:425
Finding

Undisclosed Transmission of Medical Dictation to External LLM Providers

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/main.py:457
Finding

Prompt Injection Through Untrusted Medical Dictation

Content
View full analysis
str: """Build prompt for LLM extraction.""" return f"""Extract clinical information from the following medical dictation and format as JSON with this structure: {{ "chief_complaint": "...", "history_present_illness": "...", "review_of_systems": "...", "past_medical_history": "...", "medications": ["..."], "allergies": ["..."], "social_history": "...", "family_history": "...", "vital_signs": {{ "temperature": "...", "heart_rate": "...", "blood_pressure": "...", "respiratory_rate": "...", "oxygen_saturation": "..." }}, "physical_examination": "...", "diagnostic_studies": "...", "primary_diagnosis": "...", "differential_diagnoses": ["..."], "clinical_reasoning": "...", "diagnostic_plan": "...", "therapeutic_plan": "...", "patient_education": "...", "follow_up": "..." }} Dictation text: {text} Respond ONLY with valid JSON.""" ``` ### Technical Analysis The `text` value is untrusted input obtained from a command-line string, a text file, standard input, or audio transcription. It is inserted directly into the same prompt that tells the model how to behave. A malicious or accidentally instruction-like dictation can therefore tell the model to ignore the extraction task, fabricate fields, omit warnings, or place unsupported content into diagnoses and treatment plans. Requiring valid JSON only constrains the response syntax; it does not ensure that the returned clinical content is supported by the source dictation. The response is parsed and copied into a `SOAPNote` without source-grounding checks: - No field is required to include supporting source text ...[truncated 1889 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned and Ambiguous Third-Party Dependencies

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (22)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared purpose is narrow, but the documented behavior includes audio transcription, CLI file handling, and optional use of OpenAI/Anthropic APIs without corresponding declared capabilities. This mismatch can hide materially different data flows, including transmission of sensitive medical content to external providers, causing users and reviewers to underestimate privacy, compliance, and execution risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation covers audio and transcription processing but lacks an explicit privacy warning for patient recordings, transcriptions, and generated clinical notes. In a medical context, this omission is serious because users may process PHI without adequate consent, retention controls, or understanding of whether data may be stored locally or sent to third-party APIs.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

Variant A explicitly documents use for academic writing tasks, directly contradicting the skill's declared medical-scribe function. That contradiction broadens the apparent allowed use cases and can let the skill be invoked outside its intended medical workflow, undermining controls and increasing the risk of misuse or unsafe outputs.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The LLM path transmits raw clinical dictation over the network to third-party providers, which is not inherently required just because the skill's purpose is SOAP note generation. In a medical context, sending identifiable patient narratives externally can expose PHI, violate privacy expectations, and create regulatory risk even if the feature is functionally useful.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

Sensitive medical dictation may be sent to an external LLM provider with no explicit warning, consent prompt, or privacy notice at the point of use. In healthcare workflows, lack of transparent notice materially increases the chance that clinicians unknowingly disclose PHI to third parties under unsafe or noncompliant conditions.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The full dictation text is embedded directly into the LLM prompt and sent externally, meaning raw patient narratives, medications, diagnoses, and other PHI may be disclosed wholesale to third parties. In this medical-scribe context, forwarding complete unredacted clinical content is especially dangerous because it maximizes privacy exposure and downstream misuse, retention, or breach impact.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill documents file read/write capable behavior but does not declare any explicit tool scope such as permissions or allowed-tools. This creates an authorization and review gap: operators may approve or invoke the skill without understanding that it can access local inputs and persist outputs, which is especially risky given the likely presence of PHI in medical dictation and notes.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest description ends with the standalone word "Trigger," which does not define a specific activation phrase or boundary for when the skill should run. This creates ambiguity about invocation conditions and may overlap with broad routing behavior rather than a narrowly scoped trigger.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The 'When to Use' section broadens activation beyond the medical scribe domain into generic academic writing and reproducible formatting tasks. Over-broad invocation criteria can cause the skill to be selected in inappropriate contexts, increasing the chance that sensitive medical-style processing or file-handling workflows are applied where they should not be.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation claims the tool does not provide diagnostic suggestions, yet the SOAP structure includes fields such as primary diagnosis, differential diagnoses, and clinical reasoning. This contradiction can mislead users into overtrusting the safety boundary and may result in unsupported diagnostic content being generated or accepted in a clinical workflow.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The audit metadata classifies this skill as 'Academic Writing' even though the stated skill purpose is medical dictation to SOAP notes. This misclassification can cause the skill to be evaluated, routed, or approved under the wrong safety framework, reducing scrutiny around medical-content handling and boundary controls.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The repeated trigger phrase is vague and lacks clear exclusions, which makes it easier for callers or downstream systems to over-activate the skill for loosely related tasks. In a medical context, ambiguous activation increases the chance of unintended use on non-dictation inputs or unsupported documentation requests.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The recorded evaluation includes academic-writing scenarios outside the manifest's declared purpose. This creates scope drift in validation evidence, so a deployment decision may rely on tests that do not meaningfully assess the medical documentation behavior or its safety constraints.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The broad activation description overlaps with general academic-writing requests, which materially expands the apparent operating scope beyond medical scribing. Because this is a medical-context skill, such overlap is more dangerous: it can bypass domain-specific safeguards and lead users to rely on the skill for unsupported tasks with higher trust assumptions.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code conditionally sends physician dictation text to OpenAI or Anthropic for parsing, which expands the skill's effective behavior beyond local dictation-to-note conversion into external transmission of clinical content. Because the content is medical dictation that may contain PHI, this creates a real confidentiality and compliance risk if users are not clearly informed and if provider agreements, retention settings, and data-handling controls are not enforced.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

Writing generated clinical notes to an arbitrary output path without any privacy or handling warning can lead to storage of PHI in insecure locations, shared directories, or accidentally committed files. While file output is expected functionality, the medical context makes silent persistence of sensitive content materially risky.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency list is unpinned, so builds may pull different versions over time, including newly introduced vulnerable or incompatible releases. In a medical dictation skill that likely handles sensitive clinical text, weak dependency control increases supply-chain risk and makes security posture non-reproducible.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
anthropic
dataclasses
openai
whisper

Unverifiable Dependency: anthropic has 4 known advisory(ies) (CVE-2026-34450 (Claude SDK for Python has Insecure Default File Permissions in Local Filesystem ); CVE-2026-34452 (Claude SDK for Python: Memory Tool Path Validation Race Condition Allows Sandbox); CVE-2026-34450 (The Claude SDK for Python provides access to the Claude API from Python applicat) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
93% confidence
Finding

The manifest includes anthropic without a version pin, and the package has known advisories; without a fixed version, it is impossible to verify whether deployment will resolve to a safe release. In a medical scribe context, exploitation of an affected SDK could expose sensitive local files, weaken sandboxing, or otherwise impact systems handling clinical data.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
60% confidence
Finding

Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
anthropic
dataclasses
openai
whisper

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

An unpinned openai dependency allows uncontrolled upgrades or transitive dependency changes, which can introduce vulnerable releases or breaking behavior without review. Because this skill processes physician dictation and may handle protected health information, reproducibility and strict supply-chain control are especially important.

Content

Scanner excerpt · requirements.txt (reported line 3)May include surrounding context.

text
anthropic
dataclasses
openai
whisper

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

The whisper dependency is not version-pinned, so the environment may install an unexpected release with security, integrity, or compatibility changes. For an audio transcription workflow, that creates unnecessary supply-chain exposure and can affect reliability of processing sensitive recordings.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
anthropic
dataclasses
openai
whisper

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The manifest frames the skill as converting physician verbal dictation into SOAP notes, which most directly describes note structuring from dictation content. The code additionally includes speech-to-text transcription from audio files via Whisper, expanding the behavior from note structuring into audio transcription, which is not stated in the manifest description.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.