T09 · Insecure Skill Coding Practices
- Location
scripts/build_roster.py:42- Finding
CSV and XLSX Formula Injection Through Untrusted Business Records
- Content
View full analysis
Vulnerability Details
File Location:
scripts/build_roster.py, lines 42–49, with dangerous output operations at lines 62–66 and 87–89
Vulnerability Type: Spreadsheet formula injection
Risk Level: MediumComplete Code Snippet
python def normalize(row, idx): out = [idx] for k in KEYS: v = row.get(k, "") if v is None or str(v).strip() == "": v = FILL out.append(str(v).strip()) return outThe normalized values are written directly to CSV:
python def write_csv(path, data): with open(path, "w", newline="", encoding="utf-8-sig") as f: w = csv.writer(f) w.writerow(HEADERS) w.writerows(data)They are also written directly to XLSX cells:
python for r in data: ws.append(r)Technical Analysis
The skill collects business information from public maps, directories, recruitment platforms, websites, and similar third-party sources. Fields such as business name, address, contact, business scope, and source are therefore potentially controlled by an external content provider.
normalize()converts these fields to strings but does not neutralize spreadsheet formula prefixes. Values beginning with=,+,-, or@are passed unchanged to both output formats. When the XLSX workbook is created, a value beginning with=may be stored as a formula. CSV files can likewise be interpreted as containing formulas when opened in spreadsheet applications.This crosses the trust boundary between untrusted third-party directory content and a spreadsheet interpreter running in the recipient's desktop context.
Attack Path
- An attacker publishes or modifies a business-directory field consumed by the skill, such as a business name or contact field, so that it begins with a spreadsheet formula marker.
- The skill retrieves that external record and places it into the input JSON used by
build_roster.py. normalize()preserves the formula-prefixe ...[truncated 1127 chars]
- Remediation
View remediation
Remediation Suggestions
- Introduce a centralized cell-sanitization function for every externally sourced textual value.
- If a value begins with
=,+,-, or@, prefix it with a single quote or otherwise encode it as literal text according to the target format. - For XLSX output, explicitly force untrusted cells to use the string data type rather than relying only on value prefixing.
- Apply the same policy consistently to CSV and XLSX generation.
- Consider handling leading tabs, carriage returns, line feeds, and whitespace before formula markers, because spreadsheet applications may normalize them.
- Add regression tests covering formula-prefixed values in every externally populated field and verify the resulting CSV and XLSX files contain literal text rather than formulas.
- Preserve the original value separately only if required for provenance, and never place an unsanitized version in an interpreted spreadsheet cell.
A suitable defensive pattern is:
python def sanitize_spreadsheet_cell(value): text = str(value).strip() if text and text[0] in ("=", "+", "-", "@"): return "'" + text return textFor XLSX output, additionally set externally sourced cells explicitly as strings.
