T09 · Insecure Skill Coding Practices
- Location
scripts/main.py:425- Finding
Undisclosed Transmission of Medical Dictation to External LLM Providers
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This medical-scribe skill has a real review concern because it can send sensitive clinical dictation to external LLM providers while its own risk table says there are no external API calls.
Review carefully before installing or using this with real patient data. Use only the local/rule-based path unless your organization has approved the OpenAI or Anthropic data flow for PHI, and avoid saving generated notes to shared or insecure directories. Treat all outputs as drafts requiring clinician review, especially diagnoses, medications, allergies, and treatment plans.
scripts/main.py:425Undisclosed Transmission of Medical Dictation to External LLM Providers
scripts/main.py:457Prompt Injection Through Untrusted Medical Dictation
requirements.txt:1Unpinned and Ambiguous Third-Party Dependencies
The declared purpose is narrow, but the documented behavior includes audio transcription, CLI file handling, and optional use of OpenAI/Anthropic APIs without corresponding declared capabilities. This mismatch can hide materially different data flows, including transmission of sensitive medical content to external providers, causing users and reviewers to underestimate privacy, compliance, and execution risks.
The documentation covers audio and transcription processing but lacks an explicit privacy warning for patient recordings, transcriptions, and generated clinical notes. In a medical context, this omission is serious because users may process PHI without adequate consent, retention controls, or understanding of whether data may be stored locally or sent to third-party APIs.
Variant A explicitly documents use for academic writing tasks, directly contradicting the skill's declared medical-scribe function. That contradiction broadens the apparent allowed use cases and can let the skill be invoked outside its intended medical workflow, undermining controls and increasing the risk of misuse or unsafe outputs.
The LLM path transmits raw clinical dictation over the network to third-party providers, which is not inherently required just because the skill's purpose is SOAP note generation. In a medical context, sending identifiable patient narratives externally can expose PHI, violate privacy expectations, and create regulatory risk even if the feature is functionally useful.
Sensitive medical dictation may be sent to an external LLM provider with no explicit warning, consent prompt, or privacy notice at the point of use. In healthcare workflows, lack of transparent notice materially increases the chance that clinicians unknowingly disclose PHI to third parties under unsafe or noncompliant conditions.
The full dictation text is embedded directly into the LLM prompt and sent externally, meaning raw patient narratives, medications, diagnoses, and other PHI may be disclosed wholesale to third parties. In this medical-scribe context, forwarding complete unredacted clinical content is especially dangerous because it maximizes privacy exposure and downstream misuse, retention, or breach impact.
The skill documents file read/write capable behavior but does not declare any explicit tool scope such as permissions or allowed-tools. This creates an authorization and review gap: operators may approve or invoke the skill without understanding that it can access local inputs and persist outputs, which is especially risky given the likely presence of PHI in medical dictation and notes.
The manifest description ends with the standalone word "Trigger," which does not define a specific activation phrase or boundary for when the skill should run. This creates ambiguity about invocation conditions and may overlap with broad routing behavior rather than a narrowly scoped trigger.
The 'When to Use' section broadens activation beyond the medical scribe domain into generic academic writing and reproducible formatting tasks. Over-broad invocation criteria can cause the skill to be selected in inappropriate contexts, increasing the chance that sensitive medical-style processing or file-handling workflows are applied where they should not be.
The documentation claims the tool does not provide diagnostic suggestions, yet the SOAP structure includes fields such as primary diagnosis, differential diagnoses, and clinical reasoning. This contradiction can mislead users into overtrusting the safety boundary and may result in unsupported diagnostic content being generated or accepted in a clinical workflow.
The audit metadata classifies this skill as 'Academic Writing' even though the stated skill purpose is medical dictation to SOAP notes. This misclassification can cause the skill to be evaluated, routed, or approved under the wrong safety framework, reducing scrutiny around medical-content handling and boundary controls.
The repeated trigger phrase is vague and lacks clear exclusions, which makes it easier for callers or downstream systems to over-activate the skill for loosely related tasks. In a medical context, ambiguous activation increases the chance of unintended use on non-dictation inputs or unsupported documentation requests.
The recorded evaluation includes academic-writing scenarios outside the manifest's declared purpose. This creates scope drift in validation evidence, so a deployment decision may rely on tests that do not meaningfully assess the medical documentation behavior or its safety constraints.
The broad activation description overlaps with general academic-writing requests, which materially expands the apparent operating scope beyond medical scribing. Because this is a medical-context skill, such overlap is more dangerous: it can bypass domain-specific safeguards and lead users to rely on the skill for unsupported tasks with higher trust assumptions.
The code conditionally sends physician dictation text to OpenAI or Anthropic for parsing, which expands the skill's effective behavior beyond local dictation-to-note conversion into external transmission of clinical content. Because the content is medical dictation that may contain PHI, this creates a real confidentiality and compliance risk if users are not clearly informed and if provider agreements, retention settings, and data-handling controls are not enforced.
Writing generated clinical notes to an arbitrary output path without any privacy or handling warning can lead to storage of PHI in insecure locations, shared directories, or accidentally committed files. While file output is expected functionality, the medical context makes silent persistence of sensitive content materially risky.
The dependency list is unpinned, so builds may pull different versions over time, including newly introduced vulnerable or incompatible releases. In a medical dictation skill that likely handles sensitive clinical text, weak dependency control increases supply-chain risk and makes security posture non-reproducible.
anthropic
dataclasses
openai
whisper
The manifest includes anthropic without a version pin, and the package has known advisories; without a fixed version, it is impossible to verify whether deployment will resolve to a safe release. In a medical scribe context, exploitation of an affected SDK could expose sensitive local files, weaken sandboxing, or otherwise impact systems handling clinical data.
Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.
anthropic
dataclasses
openai
whisper
An unpinned openai dependency allows uncontrolled upgrades or transitive dependency changes, which can introduce vulnerable releases or breaking behavior without review. Because this skill processes physician dictation and may handle protected health information, reproducibility and strict supply-chain control are especially important.
anthropic
dataclasses
openai
whisper
The whisper dependency is not version-pinned, so the environment may install an unexpected release with security, integrity, or compatibility changes. For an audio transcription workflow, that creates unnecessary supply-chain exposure and can affect reliability of processing sensitive recordings.
anthropic
dataclasses
openai
whisper
The manifest frames the skill as converting physician verbal dictation into SOAP notes, which most directly describes note structuring from dictation content. The code additionally includes speech-to-text transcription from audio files via Whisper, expanding the behavior from note structuring into audio transcription, which is not stated in the manifest description.
No suspicious patterns detected.