T08 · Insecure Dependencies
- Location
requirements.txt:1- Finding
Unpinned Third-Party Dependencies Create a Supply-Chain Risk
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill appears to be a legitimate medical text-mining tool, but it handles sensitive clinical notes and writes outputs without enough privacy and filesystem safeguards.
Install and run this only in an environment approved for clinical data. Treat both source notes and generated JSON as sensitive health information, restrict output locations, avoid absolute or traversal output paths, and pin/audit dependencies before use in production or regulated workflows.
requirements.txt:1Unpinned Third-Party Dependencies Create a Supply-Chain Risk
scripts/main.py:497Arbitrary Output Path Permits Overwriting Files Accessible to the Runtime User
The skill declares executable Python scripts and explicit output-file writing behavior, but it does not define any tool scope such as allowed-tools or permissions. That omission can cause an agent runtime to grant broader filesystem capabilities than the skill actually needs, increasing the chance of unintended writes or misuse if inputs are manipulated.
This skill is designed to process unstructured clinical notes from MIMIC-IV and produce extracted outputs, but the user-facing documentation does not prominently warn that the inputs and outputs may contain sensitive medical data. In a medical-text-mining context, that increases the risk of operators handling PHI or re-identifiable data insecurely, saving raw text to disk, or sharing derived outputs without appropriate controls.
The example explicitly demonstrates loading, processing, and exporting patient-derived MIMIC-IV notes to JSON without any user-facing warning about handling sensitive clinical data, de-identification limits, or downstream disclosure risk. Even though MIMIC-IV is controlled and de-identified, extracted insights can still contain sensitive medical content and may be re-exported or combined with other data in unsafe ways, making this a real privacy hygiene issue in a medical-text-mining skill.
The code loads and processes patient note data from CSV/Parquet with no visible privacy notice, access-control checks, or handling constraints despite operating on clinical free text. In this context, silent ingestion of sensitive health records makes accidental misuse more likely, especially by users who may not appreciate that downstream extraction and reporting can preserve regulated medical information.
The script writes extracted patient insights directly to a JSON file, and those insights can include raw note-derived entities, timelines, and clinical logic tied to subject and admission identifiers. In a medical-data mining skill operating on MIMIC-IV-style notes, exporting without explicit safeguards, de-identification, consent checks, or strong warnings materially increases the risk of sensitive health information being stored insecurely or mishandled downstream.
The configuration specifies English SciSpaCy models such as "en_core_sci_lg", which imposes an English-language processing assumption. The skill does not clearly present this as a user-selectable locale or explicitly justify that it is limited to English-language clinical text.
This manifest contains operational comments and guidance in Chinese across multiple sections, while other identifiers and labels remain in English. Because the file provides no stated language preference, opt-in, or locale-specific justification, it may violate a policy requiring user-facing skill materials to avoid forcing a specific language without consent.
The dependency negspacy is declared without any version constraint, making builds non-reproducible and allowing future installs to pull in unexpected or vulnerable releases. In a medical-text processing skill, this increases supply-chain risk because behavior and security posture can change between deployments without review.
negspacy
numpy
pandas
pyarrow
numpy is unpinned, so the environment may resolve to different versions over time, including releases with known defects or advisories. This is a supply-chain hygiene issue rather than an immediate exploit by itself, but it weakens assurance over what code is actually installed.
negspacy
numpy
pandas
pyarrow
pyyaml
The manifest does not pin numpy, and numpy has known advisories across some versions, so it is impossible to verify whether deployments are affected. This is dangerous because the actual installed package may vary by time and environment, defeating security review and patch assurance.
pandas is listed without a version, which means installations are not reproducible and may silently introduce vulnerable or incompatible releases. For software mining unstructured clinical text, unexpected dependency changes can affect both security and data-processing integrity.
negspacy
numpy
pandas
pyarrow
pyyaml
scispacy
Because pandas is not version-pinned and has at least one known advisory in certain releases, the project cannot demonstrate that installed environments are safe. In data-processing pipelines, this uncertainty can propagate across research or operational systems and complicate incident response.
pyarrow is unpinned, exposing the project to accidental installation of versions with known vulnerabilities or risky parsing behavior. Given that pyarrow often handles complex file formats, uncontrolled version drift is more concerning than for a purely utility package.
negspacy
numpy
pandas
pyarrow
pyyaml
scispacy
spacy
pyarrow has advisories including issues related to loading malicious data, but the lack of version pinning means the deployment could inadvertently install an affected release. In a skill that may process large structured or semi-structured datasets, this uncertainty is more dangerous because file-parsing libraries often sit on attacker-influenced input boundaries.
pyyaml is unpinned, which is risky because some PyYAML versions have had unsafe deserialization issues. If the broader skill ever parses YAML from untrusted or semi-trusted sources, an unreviewed upgrade or downgrade could materially increase exploitability.
numpy
pandas
pyarrow
pyyaml
scispacy
spacy
tqdm
PyYAML has a history of unsafe deserialization vulnerabilities, and without pinning there is no reliable way to know whether a deployed environment uses a safe release. If any component of the skill ingests YAML configuration or data from untrusted sources, this could become a code-execution path.
scispacy is unpinned, so future installations may bring in different dependency trees and behavior without validation. While not inherently a direct exploit, this weakens supply-chain control in a package used for processing sensitive clinical text.
pandas
pyarrow
pyyaml
scispacy
spacy
tqdm
spacy is unpinned, which can result in unpredictable installs and potential exposure to newly introduced vulnerabilities or breaking changes. NLP frameworks can also pull substantial transitive dependencies, increasing overall supply-chain uncertainty.
pyarrow
pyyaml
scispacy
spacy
tqdm
tqdm is unpinned, so an install may resolve to a version with known issues or unexpected behavior. Even though it is primarily a utility library, leaving it unconstrained still contributes to preventable supply-chain exposure.
pyyaml
scispacy
spacy
tqdm
tqdm has known advisories affecting some versions, and the absence of a version pin prevents verification that the installed package is not affected. Although tqdm is not central to clinical text mining logic, unverifiable utility dependencies still increase supply-chain uncertainty and can introduce avoidable risk.
No suspicious patterns detected.