Back to skill

Security audit

Swmm Calibration

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a disclosed SWMM calibration tool, but it has a real path-containment flaw that can write calibration artifacts outside the intended run directory.

Install only if you trust the calibration inputs and keep runs in a restricted workspace. Treat parameter-set JSON and validation trial names as trusted data until the skill validates trial names; avoid values containing slashes, backslashes, absolute paths, '.', or '..'. Run it with minimal filesystem privileges and review generated candidate artifacts before accepting any INP changes.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/swmm_calibrate.py:470
Finding

Path Traversal Through Unvalidated Trial Names

Content
View full analysis
Remediation
View remediation
str: if not isinstance(name, str) or not SAFE_TRIAL_NAME.fullmatch(name): raise ValueError( "Trial name may contain only letters, digits, underscores, and hyphens" ) return name ``` 2. Explicitly reject empty values, absolute paths, path separators, `.` and `..`. Apply the validation uniformly to candidate JSON names and `--trial-name`. 3. Add a containment check as a defense-in-depth measure before any directory creation or file write: ```python root = run_root.resolve() safe_name = validate_trial_name(trial["name"]) trial_dir = (root / safe_name).resolve() if not trial_dir.is_relative_to(root): raise ValueError("Trial directory escapes the configured run root") ``` For Python versions without `Path.is_relative_to()`, use `relative_to()` inside a `try` block and reject `ValueError`. 4. Reject duplicate trial names before processing a batch so one candidate cannot overwrite another candidate's artifacts. 5. Where feasible, create each trial directory with a trusted generated identifier and retain the user-provided name only as metadata. 6. Run calibration under a dedicated, unprivileged account with write access restricted to the configured run directory. 7. Add regression tests covering `../target`, `../../target`, absolute paths, platform-specific separators, empty names, duplicate names, and valid identifiers. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (7)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a full calibration and validation scaffold with optimization and Bayesian calibration capabilities. The supplied code chunk does not implement any of those behaviors. Instead, it is a narrow preprocessing/helper script that edits SWMM input files based on a patch map and parameter values. While such patching could be a supporting detail within a larger calibration system, this chunk by itself has a materially different primary purpose and lacks the declared core capabilities. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code chunk is only an observation-series reader/parser. It handles file preprocessing, delimiter detection, timestamp/flow column inference, datetime and numeric parsing, and emits summary metadata. That is at most a supporting utility for a larger calibration workflow, but it does not itself implement the declared skill’s core capabilities. The declared description presents a full SWMM calibration/validation system, whereas this code’s primary purpose is basic observed-data ingestion. This is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

Most of the declared description matches the code well: it compares simulated vs observed series, evaluates explicit parameter sets, ranks candidates by an objective, performs bounded random/LHS/adaptive search, and includes SCE-UA and DREAM-ZS branches with KGE-specific handling and convergence-related outputs. However, there is a material mismatch because the description explicitly says dedicated sensitivity-analysis methods now live on another skill, yet this code still exposes a first-class sensitivity command that evaluates parameter sets and emits sensitivity-mode outputs. That is directly inconsistent with the declared boundary of the skill. There is also an undeclared in-process gap-fill validation interface (prepare_calibration_inputs plus gap-signal emitters), though this is secondary and more of an auxiliary runtime integration detail than the primary mismatch. The main reason to flag is the explicit description-behavior inconsistency around sensitivity support.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill documentation describes capabilities that read, write, and invoke shell-backed tooling, but it does not declare any explicit tool scope or permission boundaries. In an agentic environment, missing scope metadata can allow the runtime or calling agent to over-grant filesystem and command execution privileges, increasing the blast radius if the workflow is misused or if downstream scripts are compromised.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/swmm_calibrate.py (reported line 393)May include surrounding context.

python
run_dir.mkdir(parents=True, exist_ok=True)
    rpt = run_dir / rpt_name
    out = run_dir / out_name
    proc = subprocess.run(["swmm5", str(inp), str(rpt), str(out)], capture_output=True, text=True)
    (run_dir / "stdout.txt").write_text(proc.stdout, encoding="utf-8", errors="ignore")
    (run_dir / "stderr.txt").write_text(proc.stderr, encoding="utf-8", errors="ignore")
    return proc.returncode, rpt, out

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest explicitly states that dedicated sensitivity-analysis methods now live on the swmm-uncertainty skill, implying this skill's scope is calibration/validation/search only. However, this file exposes a sensitivity subcommand that loads parameter sets, runs multiple SWMM trials, ranks them, and emits a sensitivity-mode payload, which is functionality outside the declared current scope.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This Python code invokes an external executable via subprocess.run(...) and writes its stdout/stderr to disk, but the execution path itself does not present a user-facing warning, prompt, or log message before doing so. While running SWMM is central to the script's purpose, the rule asks to flag safety-relevant code operations lacking disclosure when no visible warning is present in the code file.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.