T09 · Insecure Skill Coding Practices
- Location
scripts/parse_docx.py:133- Finding
Persistent Plaintext Copy of Confidential Client Documents
- Content
View full analysis
Vulnerability Details
File Location:
scripts/parse_docx.py, lines 133–153
Vulnerability Type: Insecure storage and handling of sensitive temporary data
Risk Level: MediumVulnerable Code
python def parse_doc_with_antiword(file_path): """Try to convert .doc file using antiword""" try: # Try antiword first output_file = file_path + '.txt' logger.info(f"Attempting to convert .doc to .txt using antiword") import subprocess result = subprocess.run( ['antiword', file_path], capture_output=True, text=True, timeout=30 ) if result.returncode == 0: # Save converted text with open(output_file, 'w', encoding='utf-8') as f: f.write(result.stdout) logger.info(f"Successfully converted .doc to {output_file}") return parse_txt(output_file)Technical Analysis
When a legacy
.docfile is processed, its complete extracted content is written to a predictable path formed by appending.txtto the original path. The resulting plaintext file is created using default process permissions and is not deleted after parsing.The project is explicitly intended to process client documents that may contain confidential company information, personal data, and business metrics. Creating an additional persistent plaintext copy unnecessarily expands the sensitive-data footprint. The predictable filename also makes the copy easy to locate.
The
subprocess.runinvocation itself does not present command injection in this code because it uses an argument list and does not enable a shell. The vulnerability is the subsequent insecure storage and retention of the conversion output.Attack Path
- A user provides a confidential legacy Word document such as
client-audit.doc. - The parser invokes
antiwordand captures the document’s complete plaintext content. - Th ...[truncated 1114 chars]
- A user provides a confidential legacy Word document such as
- Remediation
View remediation
Remediation Suggestions
- Avoid creating an intermediate file. Pass the captured output directly to the existing structure extractor:
python if result.returncode == 0: logger.info("Successfully extracted text from .doc file") return extract_structure(result.stdout, file_path)-
If disk-backed conversion is operationally necessary, create the file with
tempfile.NamedTemporaryFilein a protected temporary directory and enforce owner-only permissions such as0600. -
Delete every temporary artifact in a
finallyblock so cleanup occurs after successful parsing and after exceptions:
python import os import tempfile temp_path = None try: with tempfile.NamedTemporaryFile( mode='w', encoding='utf-8', suffix='.txt', delete=False ) as temp_file: temp_path = temp_file.name temp_file.write(result.stdout) os.chmod(temp_path, 0o600) return parse_txt(temp_path) finally: if temp_path and os.path.exists(temp_path): os.remove(temp_path)-
Do not place temporary plaintext beside the source document or use a filename derived predictably from it.
-
Document the data-retention policy and verify through automated tests that no converted plaintext remains after success, parser errors, timeouts, or interrupted processing.
