T08 · Insecure Dependencies
- Location
plot-box-violin/scripts/plot_box_violin.py:14- Finding
Execution of an Unbundled Module Through a Prepending Python Import Path
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is mostly a bioinformatics plotting toolkit, but its scripts import missing shared code from outside the reviewed artifact, which creates a material review concern.
Review this package before installing in a sensitive workspace. Its plotting purpose is understandable, but the missing external _shared/plot_style.py import should be fixed or bundled so the skill cannot load unreviewed code from a parent directory. Install dependencies in an isolated environment with reviewed versions, and avoid running it on private clinical or genomic data until the packaging issue is resolved.
plot-box-violin/scripts/plot_box_violin.py:14Execution of an Unbundled Module Through a Prepending Python Import Path
requirements.txt:1Non-Reproducible Third-Party Dependency Installation
The declared description says this skill's purpose is to route to the correct plotting sub-skill across many bioinformatics chart types. The supplied code is not a router: it is a standalone implementation for one specific plotting family (box, violin, both, raincloud). It reads CSV/TSV data, prepares groups, computes statistics (e.g. Mann-Whitney, t-test, Kruskal-Wallis, multiple-testing correction), draws plots with matplotlib, annotates significance, and saves outputs. That is a materially different primary behavior from dispatching among multiple plot sub-skills. There is no evidence of handling volcano plots, heatmaps, scatter plots, bar charts, MA plots, correlation matrices, or bubble charts in this code chunk. Therefore the description does not accurately represent the actual behavior.
The declared description says this skill routes to the appropriate plotting sub-skill across many chart types. The supplied code does not route anything; it is a concrete implementation of one specific sub-skill: heatmap generation. It parses CLI arguments, reads matrix files, filters rows by variance, scales values, performs optional hierarchical clustering, plots a heatmap with matplotlib, and saves outputs. Those behaviors are consistent with a heatmap plotting tool, but not with a dispatcher for many plotting modes. There is no evidence of handling volcano plots, scatter plots, bar charts, MA plots, correlation matrices, bubble charts, or sub-skill selection logic. Therefore the description materially overstates and mischaracterizes the code's primary purpose.
The declared purpose describes a router that selects among multiple plotting sub-skills covering volcano, heatmap, box/violin, scatter, bar, MA, correlation matrix, and bubble plots. The supplied code is not a router: it is a concrete plotting implementation for only five plot types (scatter, bar, MA, correlation matrix, bubble). It lacks volcano, box, and violin plot support, and its only heatmap-like functionality is a correlation matrix visualization rather than a general heatmap sub-skill. Additionally, the script directly reads input files and writes output figures/tables, which is operational behavior beyond a pure routing role. These are material differences in primary purpose and supported capabilities.
The declared description says this skill is a router for various publication-quality plot sub-skills (volcano, heatmap, box/violin, scatter, bar, MA, correlation matrix, bubble). The supplied code does not perform routing at all. Instead, it is a concrete plotting implementation for survival analysis, specifically Kaplan-Meier survival curves with optional confidence intervals, censoring marks, p-value annotations, median survival lines, and at-risk tables. It also performs survival-specific statistical tests and writes a statistics TSV file. Survival plotting is not mentioned in the declared purpose, and the primary function is materially different from a routing skill. Therefore this is a clear description-behavior mismatch.
The description says this skill's role is to route requests to the correct plotting sub-skill across a broad set of plot types. The code does not route anything; it is itself the concrete plotting implementation. It supports volcano, heatmap, boxplot, violin, scatter, and a 'correlation' mode that is actually just a two-variable scatter/correlation plot. It does not implement bar charts, MA plots, bubble charts, or correlation matrices as declared. Additionally, it includes statistical hypothesis testing and significance bracket annotation for two-group box/violin plots, which is a meaningful extra behavior not mentioned in the description. These differences are material enough to count as a description-behavior mismatch.
Referenced artifact was not completely inspected
**Script:** `plot-volcano/scripts/plot_volcano.py`
Referenced artifact was not completely inspected
**Script:** `plot-volcano/scripts/plot_volcano.py`
Referenced artifact was not completely inspected
**Script:** `plot-box-violin/scripts/plot_box_violin.py`
Referenced artifact was not completely inspected
**Script:** `plot-box-violin/scripts/plot_box_violin.py`
Referenced artifact was not completely inspected
**Script:** `plot-scatter-bar/scripts/plot_scatter_bar.py`
Referenced artifact was not completely inspected
**Script:** `plot-scatter-bar/scripts/plot_scatter_bar.py`
Referenced artifact was not completely inspected
**Script:** `plot-scatter-bar/scripts/plot_scatter_bar.py`
Referenced artifact was not completely inspected
**Script:** `plot-scatter-bar/scripts/plot_scatter_bar.py`
Referenced artifact was not completely inspected
**Script:** `plot-scatter-bar/scripts/plot_scatter_bar.py`
Referenced artifact was not completely inspected
**Script:** `plot-scatter-bar/scripts/plot_scatter_bar.py`
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
| "Column not found" | Check column names in CSV header match exactly (case-sensitive) |
| Wide CSV but only one group plotted | Pass `--wide-cols "col1,col2,..."` explicitly, or use `--wide-cols auto` to use every numeric column |
| Mixing wide + long flags errors out | Use either `--wide-cols` OR `--value-col`/`--group-col`, not both |
| NaN values causing issues | Script automatically removes rows with missing data; check input file for errors |
| Points overlap too much | Increase --jitter amount (e.g., 0.15) or use --point-alpha < 0.5 |
| Legend/labels cut off | Increase --fig-width or --fig-height; use tight_layout (automatic) |
| Significance brackets overlap | Reduce number of comparisons (use --stats vs_first instead of all_pairs) |
The module docstring claims support for annotations and dendrogram coloring, and the CLI exposes related options such as --col-annotation, --row-annotation, --annotation-cmaps, --row-cluster-cutoff, and --col-cluster-cutoff. However, while these values are passed into create_heatmap at L476-L479, the function never uses them, so the implemented behavior does not match the documented capabilities.
The argument parser description says the tool generates heatmaps with clustering and annotations, and multiple help strings describe column/row annotation files and dendrogram cluster coloring. In practice, the code parses these options but never loads or applies annotation overlays or dendrogram coloring, creating an active contradiction between the documented intent and real behavior.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
### Input Format
- **File types**: CSV, TSV, or plain-text table (auto-detected)
- **Columns**: Must contain numeric columns for plotting; categorical columns optional
- **Missing values**: Automatically removed (NaN, blank cells)
- **Row labels**: Optional index column for sample names
### Output Formats
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
### Input Format
- **File types**: CSV, TSV, or plain-text table (auto-detected)
- **Columns**: Must contain numeric columns for plotting; categorical columns optional
- **Missing values**: Automatically removed (NaN, blank cells)
- **Row labels**: Optional index column for sample names
### Output Formats
The code implements Kaplan-Meier estimation plus log-rank and Wilcoxon survival-group testing, which is a specialized analytical capability rather than just routing to publication-quality plot generators for the manifest's enumerated plot types. Survival analysis is not mentioned in the manifest description, so this capability is not justified by the declared purpose.
The manifest says this skill routes to the correct plot sub-skill for a listed set of plot types, implying orchestration behavior. This file instead fully implements Kaplan-Meier survival analysis and plotting itself, including statistical tests, image generation, and writing a statistics TSV, which is materially different from simple routing and also outside the listed plot types in the manifest.
The manifest says this skill routes publication-quality plotting for volcano plots, heatmaps, box/violin plots, scatter plots, bar charts, MA plots, correlation matrices, and bubble charts. The code only supports volcano, heatmap, boxplot, violin, scatter, and correlation choices, with no implementation for bar charts, MA plots, correlation matrices beyond simple pairwise scatter correlation, or bubble charts.
Given the manifest's explicit mention of 'correlation matrices', presenting 'correlation' as a plot-type option implies a dedicated correlation visualization. In the implementation, the 'correlation' branch is handled by make_scatter with x/y columns, which produces an annotated scatter plot, not a matrix across multiple variables.
The dependency is specified with only a minimum version (numpy>=1.22), which makes builds non-reproducible and allows future installs to pull in unexpected versions. That increases supply-chain risk and makes it hard to verify whether the deployed package includes security fixes or introduces regressions.
numpy>=1.22
pandas>=1.4
matplotlib>=3.5
scipy>=1.8
No suspicious patterns detected.