T09 · Insecure Skill Coding Practices
- Location
scripts/image_crawler.py:190- Finding
Unbounded and Insufficiently Validated Remote Image Downloads
- Content
View full analysis
- Remediation
View remediation
MAX_IMAGE_BYTES: return False except ValueError: return False written = 0 with open(temp_path, "wb") as f: for chunk in resp.iter_content(8192): if not chunk: continue written += len(chunk) if written > MAX_IMAGE_BYTES: raise ValueError("Image exceeds maximum permitted size") f.write(chunk) ``` 2. **Require explicit image MIME types** - Apply the same validation to Baidu and Bing. - Allow only necessary values such as `image/jpeg`, `image/png`, and `image/webp`. - Do not treat `application/octet-stream` as sufficient proof that the response is an image. 3. **Validate decoded image content** - Use a maintained image-processing library to parse and verify the completed file. - Check the actual decoded format, dimensions, and pixel count. - Configure a maximum pixel count to prevent decompression-bomb attacks. - Do not rely on URL extensions or server-controlled HTTP headers. 4. **Use temporary files and atomic publication** - Write each response to a securely created temporary file inside the output directory. - Delete it on every rejection or exception. - Atomically rename it to the final generated filename only after size, MIME, and image-structure checks succeed. 5. **Apply task-level resource controls** - Limit total downloaded bytes, candidate URLs, redirects, and execution time for each crawler run. - Consider filesystem quotas or process ...[truncated 315 chars]
