T09 · Insecure Skill Coding Practices
- Location
scripts/report_builder.py:272- Finding
Stored HTML and JavaScript Injection in Generated Analysis Reports
- Content
View full analysis
list: """Summary statistics for categorical columns.""" if not cat_cols: return [] results = [] for col in cat_cols: series = df[col] vc = series.value_counts() results.append({ 'column': col, 'count': int(series.count()), 'unique': int(series.nunique()), 'missing': int(series.isna().sum()), 'top_value': str(vc.index[0]) if len(vc) > 0 else None, 'top_count': int(vc.iloc[0]) if len(vc) > 0 else 0, 'top_pct': round(vc.iloc[0] / series.count() * 100, 1) if len(vc) > 0 and series.count() > 0 else 0, }) return results ``` The resulting column names and categorical values are inserted directly into the HTML report: ```python rows = '\n'.join(f''' {s['column']} {s['count']:,} {s['unique']} {s['missing']} {s['top_value']} {s['top_count']:,} {s['top_pct']}% ''' for s in summary) ``` Other report fields, including `dataset_name`, numeric column names, correlation feature names, and audit messages, are also interpolated into HTML without context-aware escaping. ### Technical Analysis The report builder uses Python formatted strings to construct HTML. Values originating from an analyzed dataset are treated as trusted markup instead of untrusted text. No HTML escaping or sanitization is applied before these values are placed inside element content or the re ...[truncated 2600 chars]- Remediation
View remediation
str: return escape(str(value), quote=True) ``` Use this helper for dataset names, column names, categorical values, audit messages, chart labels, and correlation feature names. 2. Prefer a templating engine with automatic escaping enabled rather than constructing the entire document with formatted strings. For example: ```python from jinja2 import Environment, select_autoescape env = Environment( autoescape=select_autoescape( enabled_extensions=('html', 'xml'), default_for_string=True ) ) template = env.from_string(template_source) report_html = template.render(report_data) ``` Do not mark dataset-derived content as safe HTML. 3. Keep markup generated by the application separate from untrusted report data. Escape values at the final rendering boundary even if earlier processing stages appear to sanitize them. 4. Add a restrictive Content Security Policy to the generated report, for example: ```html ``` Eliminating inline scripts and event-handler execution substantially limits the impact of any missed escaping location. 5. Validate chart embedding separately. Restrict embedded files to expected image formats such as PNG, JPEG, SVG only after appropriate sanitization, and GIF. Derive MIME types from an allowlist rather than directly trusting arbitrary file extensions. 6. Add regression tests using malicious dataset names, column names, categorical values, and audit messages. Tests should verify that strings such as `
