T09 · Insecure Skill Coding Practices
- Location
scripts/dlt_lottery.py:340- Finding
Insufficient Validation of Untrusted Lottery Data
- Content
View full analysis
Vulnerability Details
File Location:
scripts/dlt_lottery.py:340-356, 388-397
Vulnerability Type: Weak validation and overly broad parsing of untrusted third-party HTML
Risk Level: MediumVulnerable Code
python # 号码 - 彩经网通常有表格格式 # 尝试匹配表格中的号码 table_matches = re.findall(r'<td[^>]*>\s*(\d{2})\s*</td>', html) if len(table_matches) >= 7: result['front'] = table_matches[:5] result['back'] = table_matches[5:7] else: # 回退到通用匹配 - 找连续的 7 个两位数 combo_match = re.search( r'(\d{2})\s+(\d{2})\s+(\d{2})\s+(\d{2})\s+' r'(\d{2})\s+(\d{2})\s+(\d{2})', html ) if combo_match: result['front'] = list(combo_match.groups()[:5]) result['back'] = list(combo_match.groups()[5:7]) else: numbers = re.findall(r'\b(\d{2})\b', html) if len(numbers) >= 7: result['front'] = sorted( set(numbers[:5]), key=lambda x: int(x) ) result['back'] = numbers[5:7] return resultpython html = fetch_lottery_page(url, source['timeout']) if not html.startswith("ERROR:"): draw = parse_draw(html, source['name']) # 验证数据完整性 if draw['issue'] or ( len(draw['front']) >= 5 and len(draw['back']) >= 2 ): return (True, draw)Technical Analysis
The application downloads HTML from several external sources, including third-party lottery websites, and treats it as untrusted input. Some parser fallbacks search an entire page for generic two-digit values or table cells and then interpret the first seven matches as lottery numbers. These matches are not necessarily located in the lottery-result section and may instead represent dates, navigation content, advertisements, prices, or attacker-controlled page content.
The acceptance condition does not implement the validation strategy documented in
references/data_sources.md. It accepts a parsed response if either an issue number exists or enough number str ...[truncated 2162 chars]- Remediation
View remediation
Remediation Suggestions
-
Implement a central validation function and call it before returning success:
- Require a correctly formatted issue identifier.
- Require exactly five unique front-area numbers in the range 1–35.
- Require exactly two unique back-area numbers in the range 1–12.
- Reject non-numeric, duplicate, missing, and out-of-range values.
-
When a specific issue is requested, require the parsed issue to equal the requested issue. Do not report a list page's latest draw as the requested historical draw.
-
Replace page-wide numeric fallbacks with source-specific parsing scoped to a single draw-result container. Reject the response if the expected container or schema cannot be identified reliably.
-
Change the success condition so an issue number alone is never sufficient. Require a complete, validated draw record:
python def validate_draw(draw, requested_issue=None): issue = draw.get('issue', '') front = draw.get('front', []) back = draw.get('back', []) if not re.fullmatch(r'\d{7,8}', issue): return False if requested_issue and issue != requested_issue: return False if len(front) != 5 or len(set(front)) != 5: return False if len(back) != 2 or len(set(back)) != 2: return False try: front_values = [int(value) for value in front] back_values = [int(value) for value in back] except ValueError: return False return ( all(1 <= value <= 35 for value in front_values) and all(1 <= value <= 12 for value in back_values) )-
For third-party results, obtain matching data from a second independent source before presenting the result as verified. If agreement cannot be established, clearly mark the result as unverified or fail safely.
-
Add unit tests covering advertisements before result tables, duplicate values, out-of-range values, issue-only responses, mismatched requested issues, truncated pages, and c ...[truncated 19 chars]
-
