T09 · Insecure Skill Coding Practices
- Location
scripts/paddle_ocr.py:98- Finding
Predictable Temporary File Exposes Sensitive Invoice Images
- Content
View full analysis
Vulnerability Details
File Location:
scripts/paddle_ocr.py, lines 98-109
Vulnerability Type: Predictable temporary file and unsafe shared temporary-directory usage
Risk Level: Mediumpython # Write to a temporary file tmp_path = f"/tmp/paddle_ocr_page_{os.getpid()}.jpg" pix.save(tmp_path) doc.close() try: result = ocr_image(tmp_path) finally: try: os.remove(tmp_path) except OSError: passTechnical Analysis
The OCR adapter renders invoice or train-ticket pages to a predictable filename in the shared
/tmpdirectory. The filename contains only the process ID, which can often be observed or estimated by another local process.The destination is not created and reserved atomically through Python's
tempfilefacilities. The code also does not verify that the path is not an existing file or symbolic link and does not explicitly enforce owner-only permissions.Rendered pages can contain sensitive information, including company names, tax identifiers, invoice numbers, financial values, travel records, and partially masked identity information. Depending on the process umask and image library behavior, another local user may be able to read the temporary image or pre-position a symbolic link at the expected path. If
pix.save()follows that link, the victim process could overwrite a file writable under its own privileges.Although the file is normally deleted in a
finallyblock after OCR, it remains present while OCR is running. Abrupt process termination before cleanup may also leave the rendered image on disk.Attack Path
- An attacker obtains local access to the same multi-user system or container namespace.
- The attacker monitors process identifiers or predicts the path
/tmp/paddle_ocr_page_<PID>.jpg. - The attacker watches for the file to appear and attempts to read or copy it while OCR processing is active.
- A ...[truncated 951 chars]
- Remediation
View remediation
Remediation Suggestions
- Replace the predictable path with
tempfile.NamedTemporaryFileor a privatetempfile.TemporaryDirectory. - Create temporary files atomically and use owner-only permissions.
- Keep rendering, OCR, document closure, and deletion inside structured
try/finallyblocks. - Avoid shared deterministic paths derived from process IDs, page numbers, or input names.
- Ensure cleanup executes for OCR errors as well as successful processing.
- Consider processing image bytes in memory if the OCR library supports byte arrays or image objects.
Example hardening pattern:
python import os import tempfile doc = pymupdf.open(pdf_path) tmp_path = None try: page = doc[page_num] pix = page.get_pixmap(matrix=pymupdf.Matrix(scale, scale)) with tempfile.NamedTemporaryFile( prefix="paddle_ocr_", suffix=".jpg", delete=False ) as tmp: tmp_path = tmp.name os.chmod(tmp_path, 0o600) pix.save(tmp_path) result = ocr_image(tmp_path) finally: doc.close() if tmp_path: try: os.remove(tmp_path) except FileNotFoundError: passA private temporary directory created with
tempfile.TemporaryDirectory()is preferable when multiple intermediate files are required.- Replace the predictable path with
