T09 · Insecure Skill Coding Practices
- Location
scripts/build_report_v2.py:75- Finding
Arbitrary Code Execution Through Executable Report Data Module
- Content
View full analysis
Vulnerability Details
File Location:
scripts/build_report_v2.py, lines 75–86
Vulnerability Type: Unsafe dynamic import of an untrusted data file
Risk Level: HighVulnerable Code
python def load_data(path: Path | None = None) -> Any: """Load v2_data.py from the working directory or an explicit path.""" source = path or (Path.cwd() / "v2_data.py") if not source.is_file(): raise FileNotFoundError(f"Data module not found: {source}") spec = importlib.util.spec_from_file_location("briefing_v2_data", source) if spec is None or spec.loader is None: raise RuntimeError(f"Cannot load data module: {source}") module = importlib.util.module_from_spec(spec) spec.loader.exec_module(module) return moduleTechnical Analysis
The report renderer treats
v2_data.pyas a structured report-data source, but loads it through Python's import system. The call tospec.loader.exec_module(module)executes every top-level statement in the selected file.The renderer accepts an explicit data path through its second command-line argument; otherwise, it loads
v2_data.pyfrom the current working directory. Therefore, anyone able to supply or modify that file can place arbitrary Python statements in it.The subsequent
validate()operation does not mitigate this issue because module execution occurs first. HTML escaping and HTTP(S)-only link validation also protect only the generated document and do not restrict Python behavior during data loading.For example, a malicious data file could contain top-level code equivalent to:
python import os os.system("attacker-controlled command")The malicious statement would execute as soon as
load_data()imports the file, before the report data is validated or rendered.Attack Path
- An attacker creates or modifies a
v2_data.pyfile in the report-generation working directory, or convinces the operator to provide an attacker-controlled file as the ...[truncated 1462 chars]
- An attacker creates or modifies a
- Remediation
View remediation
Remediation Suggestions
-
Replace executable Python data with a non-executable format.
Use JSON as the primary data contract and load it withjson.load()rather than importing a Python module. -
Apply strict schema validation before rendering.
Define allowed top-level fields, nested structures, scalar types, required fields, permitted status values, and date formats. Reject unknown or incorrectly typed fields rather than silently accepting them. -
Enforce resource limits.
Set reasonable limits for file size, nesting depth, list length, and string length to reduce denial-of-service risks from oversized input. -
Retain existing output protections.
Continue escaping all rendered text and limiting links to absolute HTTP(S) URLs. These controls remain necessary even after the code-execution flaw is removed. -
If legacy Python files must be supported, never import them.
Parse the source with Python'sastmodule and permit only simple assignments whose values passast.literal_eval(). Reject imports, calls, attribute access, comprehensions, and all other executable syntax. This should be a temporary compatibility mechanism rather than the preferred design. -
Restrict input provenance and filesystem behavior.
Require an explicitly approved data path, avoid implicitly trusting the current working directory, and document that report inputs must come from a trusted workflow.
A safer loading pattern would be:
python import json from pathlib import Path from typing import Any def load_data(path: Path) -> dict[str, Any]: if not path.is_file(): raise FileNotFoundError(f"Data file not found: {path}") if path.stat().st_size > 10 * 1024 * 1024: raise ValueError("Data file exceeds the permitted size") with path.open("r", encoding="utf-8") as handle: data = json.load(handle) if not isinstance(data, dict): raise ValueError("Report data must be a JSON object") ...[truncated 44 chars]-
