T09 · Insecure Skill Coding Practices
- Location
scripts/script.sh:89- Finding
Arbitrary Python Code Execution Through Unsafe Filename Interpolation
- Content
View full analysis
Vulnerability Details
File Location:
scripts/script.sh:89-126,scripts/script.sh:247-258,scripts/script.sh:502-510, andscripts/script.sh:566-578
Vulnerability Type: Python source-code injection through unquoted shell heredocs
Risk Level: HighVulnerable Code
The PDF extraction fallback directly inserts the user-supplied PDF path into Python source:
bash _extract_with_python() { local file="$1" python3 <<PYEOF import sys text = "" # Try PyPDF2 first try: from PyPDF2 import PdfReader reader = PdfReader("$file") for page in reader.pages: t = page.extract_text() if t: text += t + "\n\n" if text.strip(): print(text) sys.exit(0) except ImportError: pass except Exception as e: print(f" PyPDF2 error: {e}", file=sys.stderr) # Try pdfminer.six try: from pdfminer.high_level import extract_text as pdfminer_extract text = pdfminer_extract("$file") if text.strip(): print(text) sys.exit(0) except ImportError: pass except Exception as e: print(f" pdfminer error: {e}", file=sys.stderr) # Basic fallback: try to read raw text streams from PDF try: import re with open("$file", "rb") as f: raw = f.read()The page-count fallback contains the same unsafe interpolation:
bash if command -v python3 &>/dev/null; then python3 <<PYEOF try: from PyPDF2 import PdfReader r = PdfReader("$file") print(len(r.pages)) except Exception: # Fallback: count /Type /Page occurrences import re with open("$file", "rb") as f: data = f.read() count = len(re.findall(rb'/Type\s*/Page[^s]', data)) print(count if count > 0 else "?") PYEOFThe metadata fallback also inserts the PDF path into executable Python source:
bash python3 <<PYEO ...[truncated 3853 chars]- Remediation
View remediation
Remediation Suggestions
Never interpolate paths or other external values into generated Python source. Pass each value as a command-line argument or environment variable, and quote the heredoc delimiter to disable shell expansion.
For example, replace the extraction pattern with:
bash python3 - "$file" <<'PYEOF' import sys file_path = sys.argv[1] from PyPDF2 import PdfReader reader = PdfReader(file_path) PYEOFApply the same design to the page-count and metadata routines.
For JSON export, pass every path and metadata value separately:
bash python3 - "$latest" "$base" "$json_out" <<'PYEOF' import datetime import json import sys latest, base, json_out = sys.argv[1:4] with open(latest, encoding="utf-8") as source_file: content = source_file.read() data = { "source": base + ".pdf", "format": "markdown", "content": content, "exported": datetime.datetime.now().isoformat(), } with open(json_out, "w", encoding="utf-8") as output_file: json.dump(data, output_file, indent=2, ensure_ascii=False) PYEOFAdditional hardening should include:
- Quote every heredoc delimiter used for static Python code.
- Treat filenames, configured directories, and derived basenames as untrusted data.
- Add regression tests using filenames containing quotes, backslashes, spaces, Unicode characters, shell metacharacters, and line breaks.
- Validate that all supported commands process unusual filenames without changing generated program syntax.
- Run PDF parsing in a restricted process or sandbox when processing files from untrusted sources.
- Review future embedded-language invocations for the same source-generation pattern.
