Back to skill

Security audit

Template SDS Generator

Security checks for vulnerabilities and agentic risk

Overview

This looks like a real SDS generator, but it needs Review because its automatic dependency bootstrap is incomplete and under-controlled, and some generated or cached document data can create downstream safety and privacy risks.

Install only after the publisher ships a complete hashed lock file or a prebuilt runtime, and run it in an isolated workspace or container with only the intended SDS inputs. Replace config/fixed_company.yml before producing real documents, treat generated CSV files as untrusted unless formula-sanitized, and disable or redirect OCR caching when source SDS files contain confidential supplier or formulation data.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
src/sds_generator/outputs/source_map_csv.py:39
Finding

Spreadsheet Formula Injection in Provenance CSV Output

Content
View full analysis

Vulnerability Details

File Location: src/sds_generator/outputs/source_map_csv.py:39-44, src/sds_generator/outputs/source_map_csv.py:99-128, and src/sds_generator/outputs/source_map_csv.py:172-185
Vulnerability Type: CSV/spreadsheet formula injection caused by insufficient output neutralization
Risk Level: Medium

Vulnerable Code

python
def _serialize_cell(value: Any) -> str:
    if value is None:
        return ""
    if isinstance(value, list | dict):
        return json.dumps(value, ensure_ascii=False, sort_keys=True)
    return str(value)

Untrusted source-document values and excerpts are placed into CSV records without formula neutralization:

python
records.append(
    {
        "field_path": field_path,
        "field_priority": str(entry.get("priority", "")),
        "final_value": _serialize_cell(structured_field.value),
        "display_value": _serialize_cell(structured_field.display_value),
        "status": structured_field.status.value,
        "origin_kind": origin_kind.value if origin_kind is not None else "",
        "selected": bool(selected_row.selected) if selected_row is not None else False,
        "selection_reason_code": selection_reason_code(
            field_path=field_path,
            status=structured_field.status,
            origin_kind=origin_kind,
            rows=raw_rows,
            notes=notes,
            evidence_required=bool(entry.get("evidence_required")),
        ),
        "selection_reason": selection_reason_text(
            field_path=field_path,
            status=structured_field.status,
            origin_kind=origin_kind,
            selected_row=selected_row,
            notes=notes,
            evidence_required=bool(entry.get("evidence_required")),
        ),
        "source_file": selected_row.source_file if selected_row is not None else "",
        "source_profile": selected_row.source_profile.value if selected_row is not None else "",
        "source_authority": selecte
...[truncated 3847 chars]
Remediation
View remediation

Remediation Suggestions

  1. Introduce a dedicated serializer that neutralizes every string cell beginning with a dangerous formula prefix after accounting for leading whitespace, tabs, and control characters.
python
FORMULA_PREFIXES = ("=", "+", "-", "@")

def _csv_safe_cell(value: Any) -> str:
    serialized = _serialize_cell(value)
    inspected = serialized.lstrip(" \t\r\n")
    if inspected.startswith(FORMULA_PREFIXES):
        return "'" + serialized
    return serialized
  1. Apply the safe serializer to every CSV field, not only known evidence fields. This should include filenames, selection reasons, excerpts, normalized values, and derived values.

  2. Keep the original unmodified value in JSON if exact machine-readable provenance is required, while treating CSV as a spreadsheet-facing representation.

  3. Consider generating an XLSX file with cells explicitly typed as text if spreadsheet compatibility is a core requirement.

  4. Add regression tests for values beginning with:

    • =
    • +
    • -
    • @
    • Leading spaces followed by a formula marker
    • Tabs or carriage returns followed by a formula marker
  5. Document that previously generated CSV artifacts should be treated as untrusted when opened in formula-capable applications.

T08 · Insecure Dependencies

Warning
Location
scripts/runtime_common.py:90
Finding

Incomplete and Non-Reproducible Automatic Dependency Bootstrap

Content
View full analysis

Vulnerability Details

File Location: requirements.txt:1, scripts/runtime_common.py:18, scripts/runtime_common.py:90-96, and scripts/bootstrap_runtime.py:8-15
Vulnerability Type: Unsafe and incomplete dependency bootstrap using a missing lock file and an automatic package-manager upgrade
Risk Level: Medium

Vulnerable Code

The shipped requirements file references a lock file that is absent from the audited project:

text
-r requirements.lock

The runtime points directly to the missing file:

python
REQUIREMENTS_LOCK = PROJECT_ROOT / "requirements.lock"

The installer upgrades pip from the configured package index before attempting to process the absent lock file:

python
def install_requirements(command: list[str]) -> None:
    upgrade = run_python(command, ["-m", "pip", "install", "--upgrade", "pip"], check=False)
    if upgrade.returncode != 0:
        raise RuntimeError(upgrade.stderr.strip() or upgrade.stdout.strip() or "Failed to upgrade pip.")
    install = run_python(command, ["-m", "pip", "install", "-r", str(REQUIREMENTS_LOCK)], check=False)
    if install.returncode != 0:
        raise RuntimeError(install.stderr.strip() or install.stdout.strip() or "Failed to install requirements.")

The bootstrap invokes this installation path automatically when the local runtime is incomplete:

python
def main() -> int:
    create_or_repair_venv()
    runtime = venv_python()
    ensure_pip(runtime)
    if not runtime_complete(runtime):
        install_requirements(runtime)
    ensure_base_template(runtime)
    return 0

Technical Analysis

The canonical launcher automatically creates a virtual environment and installs dependencies. However, the repository does not contain the requirements.lock file referenced by both requirements.txt and runtime_common.py.

Before discovering that dependency installation cannot complete, the code executes:

text
python -m pip install --upgrade pip

This contact ...[truncated 2366 chars]

Remediation
View remediation

Remediation Suggestions

  1. Ship the referenced requirements.lock file as part of the Skill package.

  2. Lock all direct and transitive dependencies to exact versions and include cryptographic hashes for every accepted distribution artifact.

  3. Install locked dependencies with hash enforcement:

text
python -m pip install --require-hashes -r requirements.lock
  1. Validate the lock file before performing any network-backed or environment-modifying operation:
python
if not REQUIREMENTS_LOCK.is_file():
    raise RuntimeError(f"Required dependency lock file is missing: {REQUIREMENTS_LOCK}")
  1. Remove the unconditional automatic pip self-upgrade. If a minimum pip version is required, validate the installed version and provide an explicit, separately authorized upgrade procedure.

  2. Pin or explicitly configure an approved package index rather than silently inheriting arbitrary user or system pip configuration. For high-assurance deployments, use a controlled internal mirror.

  3. Consider installing from a prebuilt, signed wheelhouse in offline mode:

text
python -m pip install --no-index --find-links wheelhouse --require-hashes -r requirements.lock
  1. Verify dependency artifacts in CI and ensure release packaging tests fail if the lock file or required wheels are absent.

  2. Update the runtime documentation so its self-bootstrap claim accurately reflects the files distributed with the package.

Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (34)

Known Vulnerable Dependency: lxml==5.3.0 — 2 advisory(ies): CVE-2026-41066 (lxml: Default configuration of iterparse() and ETCompatXMLParser() allows XXE to); CVE-2026-41066 (lxml is a library for processing XML and HTML in the Python language. Prior to 6)

High
Category
Supply Chain
Confidence
92% confidence
Finding

The project pins lxml==5.3.0, and the supplied advisory indicates that affected versions may permit XXE through unsafe default XML parsing behavior. This skill processes DOCX/XML-derived content, so a vulnerable XML library is directly relevant and could enable file disclosure, SSRF, or parser abuse if untrusted documents are handled.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: Pillow==11.1.0 — 16 advisory(ies): CVE-2026-55379 (Pillow `BdfFontFile`: `Image.new()` called without `_decompression_bomb_check()`); CVE-2026-55798 (Pillow: WindowsViewer.get_command() OS command injection via unescaped shell pat); CVE-2026-54060 (Pillow: `FontFile.compile()`: `Image.new()` called without `_decompression_bomb_) +13 more

High
Category
Supply Chain
Confidence
88% confidence
Finding

The project pins Pillow==11.1.0, which the finding reports as having multiple high-severity advisories, including image parsing and possible command-injection related issues. Because this skill ingests document assets and may process embedded images from untrusted SDS/MSDS source files, a vulnerable imaging library increases the risk of denial of service or code-execution-adjacent exploitation paths depending on which Pillow features are exercised.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: pytest==8.3.4 — 2 advisory(ies): CVE-2025-71176 (pytest has vulnerable tmpdir handling); CVE-2025-71176 (pytest has vulnerable tmpdir handling)

High
Category
Supply Chain
Confidence
80% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · src/sds_generator/render/pdf_export.py (reported line 54)May include surrounding context.

python
directory.mkdir(parents=True, exist_ok=True)

                command = _libreoffice_command(binary, docx_path, output_dir, profile_dir)
                env = os.environ.copy()
                env.update(
                    {
                        "HOME": str(temp_root),

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The skill invokes Python launchers, creates a virtual environment, installs dependencies at runtime, reads and writes files, and may call external binaries like LibreOffice or Tesseract, but it does not declare any explicit tool scope such as permissions or allowed-tools. This creates a trust gap where an agent may grant broader shell, filesystem, and environment access than users expect, increasing the chance of unintended command execution, data exposure, or unsafe file modification.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/run_openclaw_skill.py (reported line 13)May include surrounding context.

python
def _bootstrap() -> list[str]:
    bootstrap_script = Path(__file__).with_name("bootstrap_runtime.py")
    result = subprocess.run(
        [sys.executable, str(bootstrap_script)],
        capture_output=True,
        text=True,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/run_openclaw_skill.py (reported line 34)May include surrounding context.

python
ensure_base_template(runtime)
    generator_script = Path(__file__).with_name("generate_sds.py")
    command = runtime + [str(generator_script), *args]
    completed = subprocess.run(command, check=False)
    return completed.returncode

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/runtime_common.py (reported line 51)May include surrounding context.

python
def run_python(command: list[str], args: list[str], *, check: bool = False) -> subprocess.CompletedProcess[str]:
    return subprocess.run(command + args, cwd=PROJECT_ROOT, capture_output=True, text=True, check=check)


def python_exists(command: list[str]) -> bool:

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code automatically creates a virtual environment, provisions pip, upgrades pip, and installs packages from requirements.lock at runtime. That behavior expands the trust boundary to package indexes and the lockfile contents; if dependencies or indexes are compromised, executing this helper can result in arbitrary code execution during installation, which is more dangerous than the skill's document-generation purpose suggests.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The helper performs environment creation and package installation without any visible user-facing warning or approval gate. Silent dependency installation can surprise operators, violate least astonishment, and expose systems to unintended network access and execution of package install hooks, especially in constrained or high-trust environments.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/runtime_doctor.py (reported line 19)May include surrounding context.

python
def _binary_version(binary: str, args: list[str]) -> str:
    if not shutil.which(binary):
        return "missing"
    result = subprocess.run([binary] + args, capture_output=True, text=True, check=False)
    return (result.stdout or result.stderr).splitlines()[0].strip() if result.returncode == 0 else "present"

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This function writes a JSON manifest containing potentially sensitive fields such as settings, source file paths, source authority, revision dates, and output locations. The code performs the file write silently, with no confirmation prompt, logging, or explanatory comment/docstring indicating that this metadata will be persisted.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · src/sds_generator/parsing/ocr_backends.py (reported line 45)May include surrounding context.

python
def version(self) -> str | None:
        if not self.is_available():
            return None
        completed = subprocess.run(
            ["tesseract", "--version"],
            check=False,
            capture_output=True,

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · src/sds_generator/parsing/ocr_backends.py (reported line 63)May include surrounding context.

python
image_width: int | None = None,
        image_height: int | None = None,
    ) -> OCRPageResult:
        completed = subprocess.run(
            ["tesseract", str(image_path), "stdout", "--psm", "6", "tsv"],
            check=False,
            capture_output=True,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The code persistently stores full OCR output for PDF pages to disk in a cache directory without any indication of consent, retention controls, access restrictions, or data minimization. Because SDS/MSDS workflows may process proprietary formulations, supplier details, or other sensitive document contents, this creates a confidentiality risk through unintended local persistence and later disclosure to other users, processes, backups, or shared environments.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The code forcibly replaces supplier/emergency contact fields with values from config/fixed_company.yml regardless of source-document content. In an SDS/MSDS generator, silently overriding supplier identity and emergency contacts can produce inaccurate safety documents, misattribute product responsibility, and direct responders or customers to the wrong party, which is especially sensitive in hazardous-material workflows.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · src/sds_generator/render/pdf_export.py (reported line 63)May include surrounding context.

python
"XDG_RUNTIME_DIR": str(runtime_dir),
                    }
                )
                result = subprocess.run(command, capture_output=True, text=True, check=False, env=env)

            pdf_path = output_dir / f"{docx_path.stem}.pdf"
            if result.returncode == 0 and pdf_path.exists():

Unverifiable Dependency: setuptools has 10 known advisory(ies) (CVE-2013-1633 (Setuptools vulnerable to Man-in-the-middle attacks); CVE-2025-47273 (setuptools has a path traversal vulnerability in PackageIndex.download that lead); CVE-2024-6345 (setuptools vulnerable to Command Injection via package URL) +7 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Unverifiable Dependency: wheel has 4 known advisory(ies) (CVE-2026-24049 (Wheel Affected by Arbitrary File Permission Modification via Path Traversal in w); CVE-2022-40898 (pypa/wheel vulnerable to Regular Expression denial of service (ReDoS)); CVE-2022-40898 (An issue discovered in Python Packaging Authority (PyPA) Wheel 0.37.1 and earlie) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The manifest emphasizes generation from an input DOCX template, but this code can execute an auxiliary script to synthesize a base template when the expected template file is absent. That adds a code-execution/build capability beyond straightforward SDS package generation from provided inputs.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

ensure_base_template invokes build_base_template.py when the template is missing, which is intended to create assets/templates/sds_base.docx, but there is no confirmation, log, or warning in this file indicating that a file write will occur. File creation is a state-changing operation that should be disclosed when it happens automatically.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · src/sds_generator/content_policy.py (reported line 46)May include surrounding context.

python
def _field_value(final_document: FinalSDSDocument, field_path: str) -> Any:
    section_key, field_name = field_path.split(".", 1)
    return getattr(getattr(final_document, section_key), field_name)


def _set_field_value(final_document: FinalSDSDocument, field_path: str, value: Any) -> None:

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · src/sds_generator/content_policy.py (reported line 51)May include surrounding context.

python
def _set_field_value(final_document: FinalSDSDocument, field_path: str, value: Any) -> None:
    section_key, field_name = field_path.split(".", 1)
    setattr(getattr(final_document, section_key), field_name, value)


def apply_content_policy_defaults(final_document: FinalSDSDocument) -> None:

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · src/sds_generator/outputs/structured_json.py (reported line 207)May include surrounding context.

python
field_path = f"{section_key}.{field_name}"
            fields[field_name] = _build_structured_field(
                field_path=field_path,
                value=getattr(section_model, field_name),
                rows_by_field=rows_by_field,
                review_flags_by_field=review_flags_by_field,
                review_notes_by_field=review_notes_by_field,

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · src/sds_generator/parsing/layout_cleanup.py (reported line 52)May include surrounding context.

python
def _get_attr(item: Any, name: str, default: Any = None) -> Any:
    return getattr(item, name, default)


def _get_coord(item: Any, coord: str, default: float = 0.0) -> float:

Static analysis

No suspicious patterns detected.