T09 · Insecure Skill Coding Practices
- Location
scripts/get_data.py:154- Finding
Unbounded Remote Response and Attachment Writes
- Content
View full analysis
Vulnerability Details
File Location:
scripts/get_data.py, lines 154–178 and 199
Vulnerability Type: Unbounded processing and storage of remotely supplied data
Risk Level: MediumVulnerable Code
python def _decode_attachment_base64(data: Dict[str, Any], output_dir: Path) -> List[Dict[str, str]]: attachments: List[Dict[str, str]] = [] output_dir.mkdir(parents=True, exist_ok=True) file_map = [ ("wordBase64", "docx", "DOCX"), ("pdfBase64", "pdf", "PDF"), ] safe_article_id = _safe_article_id(data.get("articleId")) for key, ext, ftype in file_map: b64_value = data.get(key) b64_str = b64_value.strip() if isinstance(b64_value, str) else "" if not b64_str: continue try: raw = base64.b64decode(b64_str) except Exception: continue file_name = "{0}_{1}.{2}".format(safe_article_id, ftype.lower(), ext) file_path = output_dir / file_name file_path.write_bytes(raw) attachments.append({"type": ftype, "path": str(file_path)}) return attachmentspython with urllib_request.urlopen(req, timeout=TIMEOUT_SECONDS) as resp: raw_body = resp.read().decode("utf-8", errors="replace")Technical Analysis
The implementation reads the entire HTTP response into memory without enforcing a maximum response size. It subsequently accepts arbitrarily large Base64 attachment fields, decodes each field entirely in memory, and writes the resulting bytes without a decoded-size limit or available-space check.
The declared attachment type is also trusted without validating file signatures or structure. Consequently, arbitrary remote bytes may be stored under
.docxor.pdfextensions. Although the upstream service uses HTTPS, a compromised or malicious API endpoint could still provide hostile content.Attack Path
- An attacker compromises the configured API service or otherwise causes it to return ...[truncated 1160 chars]
- Remediation
View remediation
Remediation Suggestions
- Stream the HTTP response and reject it when a strict maximum byte count is exceeded.
- Check
Content-Lengthwhen present, while still enforcing the limit during streaming because the header may be absent or inaccurate. - Enforce encoded and decoded attachment-size limits before allocating or writing data.
- Decode Base64 incrementally rather than creating a complete second in-memory copy.
- Check available disk capacity before writing attachments.
- Validate PDF magic bytes and structure before using a
.pdfextension. - Validate that DOCX output is a well-formed ZIP container with the expected Office document entries.
- Use strict Base64 validation, such as
base64.b64decode(value, validate=True). - Write to a temporary file with restrictive permissions and atomically rename it only after successful validation.
- Surface invalid or oversized attachments as explicit errors instead of silently accepting or discarding them.
