T09 · Insecure Skill Coding Practices
- Location
step1_extract_resumes.py:23- Finding
Unbounded ZIP Extraction Enables Resource-Exhaustion Attacks
- Content
View full analysis
Vulnerability Details
File Location:
step1_extract_resumes.py, lines 23-27
Vulnerability Type: Unrestricted archive extraction
Risk Level: MediumVulnerable Code
python def extract_zip_if_needed(zip_path, extract_dir): """Extract ZIP file if needed""" if zipfile.is_zipfile(zip_path): print(f"Extracting ZIP file: {zip_path}") with zipfile.ZipFile(zip_path, 'r') as zip_ref: zip_ref.extractall(extract_dir) return extract_dir return os.path.dirname(zip_path)Technical Analysis
The script extracts every entry from a user-provided ZIP archive through
ZipFile.extractall()without first enforcing limits on:- The number of archive entries
- Total uncompressed size
- Maximum size of an individual entry
- Compression ratio
- Directory nesting depth
- Encrypted or otherwise abnormal archive entries
Resume archives are an expected untrusted input to this skill. A maliciously constructed ZIP bomb can have a small compressed size while expanding into a very large amount of data. Extraction occurs before the script filters for PDF files, so unsupported files can also consume storage even though they are never processed.
The subsequent PDF traversal and parsing can further amplify resource use if the archive contains many files or oversized PDF documents.
Attack Path
- An attacker creates a ZIP archive containing highly compressible data, many nested entries, or an excessive number of files.
- The attacker submits the archive as a resume package.
- The skill invokes
step1_extract_resumes.pywith the attacker-controlled archive. zipfile.is_zipfile()validates only that the input has a recognizable ZIP structure.zip_ref.extractall(extract_dir)expands every entry without resource limits.- The archive consumes available disk space, I/O capacity, memory, or processing time.
- Resume processing or other work ...[truncated 730 chars]
- Remediation
View remediation
Remediation Suggestions
Replace unrestricted
extractall()usage with validation followed by controlled, per-entry extraction.- Inspect all entries with
ZipFile.infolist()before extracting anything. - Enforce conservative limits on:
- Total entry count
- Total declared uncompressed size
- Individual entry size
- Compression ratio
- Path depth
- Reject encrypted entries, unsupported file types, symbolic links, and suspicious metadata.
- Permit only expected resume extensions such as
.pdf,.doc, and.docx. - Resolve each destination path and verify that it remains under the intended extraction directory.
- Stream each accepted entry with a byte limit instead of using
extractall(). - Apply filesystem quotas, execution timeouts, and process resource limits.
- Ensure temporary data is removed through a
try/finallyblock even when extraction or parsing fails.
Example validation pattern:
python MAX_FILES = 200 MAX_FILE_SIZE = 20 * 1024 * 1024 MAX_TOTAL_SIZE = 500 * 1024 * 1024 ALLOWED_EXTENSIONS = {".pdf", ".doc", ".docx"} entries = zip_ref.infolist() if len(entries) > MAX_FILES: raise ValueError("Archive contains too many files") total_size = 0 for entry in entries: total_size += entry.file_size if entry.file_size > MAX_FILE_SIZE: raise ValueError("Archive entry exceeds the size limit") if total_size > MAX_TOTAL_SIZE: raise ValueError("Archive exceeds the total size limit")These checks should be combined with canonical path validation and bounded streaming extraction.
- Inspect all entries with
