T09 · Insecure Skill Coding Practices
- Location
mx_data.py:306- Finding
Unsanitized API Data Allows Excel Formula Injection
- Content
View full analysis
Vulnerability Details
File Location:
mx_data.py, lines 136-140 and 306-308
Vulnerability Type: Excel formula injection caused by exporting untrusted data without neutralization
Risk Level: MediumVulnerable Code
API-supplied values are converted to strings without checking for spreadsheet formula prefixes:
python raw_values = table.get(key, []) value = raw_values[row_idx] if row_idx < len(raw_values) else "" row[label] = flatten_value(value) rows.append(row)The resulting values are subsequently written directly to an XLSX workbook:
python df = pd.DataFrame(table["rows"], columns=table["fieldnames"]) df.to_excel(writer, sheet_name=table["sheet_name"], index=False)Technical Analysis
Values returned by the remote financial-data API are treated as trusted spreadsheet content. The
flatten_value()function converts values to strings but does not neutralize strings beginning with spreadsheet formula markers, particularly=.When pandas exports the values through the
openpyxlengine, formula-like strings can be stored as active spreadsheet formulas instead of inert text. If an upstream response is compromised, manipulated, or otherwise contains attacker-controlled values, the generated workbook may therefore include formulas that are evaluated when opened in spreadsheet software.Formula behavior depends on the spreadsheet application and its security configuration. Potential payloads can reference external resources, manipulate displayed data, trigger link-resolution behavior, or abuse application-specific formula functionality. This issue does not establish that the current API is malicious; it creates an exploitable trust-boundary weakness if hostile data reaches the API response.
The same neutralization policy should be applied to all remotely derived workbook content, including cell values and column labels.
Attack Path
- An attacker gains influence ove ...[truncated 1713 chars]
- Remediation
View remediation
Remediation Suggestions
-
Sanitize every remotely sourced string before placing it in a spreadsheet cell. At minimum, detect values beginning with formula markers and prefix them with an apostrophe so the spreadsheet treats them as text.
python def safe_excel_value(value: Any) -> str: text = flatten_value(value) if text.startswith(("=", "+", "-", "@")): return "'" + text return text -
Apply the sanitizer when constructing rows:
python row[label] = safe_excel_value(value) -
Apply equivalent protection to remotely derived headers and labels, not only ordinary data cells.
-
Where supported, explicitly set exported cells to the string data type rather than relying solely on prefix escaping.
-
Keep the raw JSON response separate from the workbook and clearly treat it as untrusted external data.
-
Add regression tests covering values such as
=1+1,=HYPERLINK(...),+1,-1, and@SUM(...). Inspect the resulting workbook to verify that these cells are stored and displayed as literal text rather than formulas. -
Document that generated workbooks contain externally sourced financial data and should not be opened with external-link updates or active-content features enabled unless the source is trusted.
-
