T09 · Insecure Skill Coding Practices
- Location
scripts/news_crawler.js:98- Finding
Unvalidated Remote URLs Permit Arbitrary Outbound Requests
- Content
View full analysis
a.getAttribute('href') || ''); const headers = { 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'Accept-Encoding': 'gzip, deflate', 'Accept-Language': 'zh-CN,zh;q=0.9,en;q=0.8', 'Cache-Control': 'no-cache', 'Cookie': 'cna=DLYSGBDthG4CAbRVCNxSxGT6', 'Host': 'tv.cctv.com', 'Pragma': 'no-cache', 'Proxy-Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.159 Safari/537.36' }; const data = await Promise.all(pageUrls.map(async pageUrl => { try { const pageResponse = await fetch(pageUrl, { headers }); ``` Equivalent unvalidated URL extraction and fetching occurs in all three crawler implementations: ```js const pageUrls = rawList.slice(1).map(item => item.match(/(http.*)/)?.[0].split('\'')[0] || ''); ``` ```js const pageUrls = soup.querySelectorAll('#contentELMT1368521805488378 li a') .slice(1) .map(a => a.getAttribute('href') || ''); ``` ### Technical Analysis The crawler initially contacts a fixed CCTV URL, but subsequently treats links extracted from the returned HTML as trusted request destinations. It does not validate: - The destination hostname - The URL scheme or port - Redirect destinations - Whether DNS resolution produces a loopback, private, link-local, or metadata-service address - Whether the destination remains within the expected CCTV domain Remote HTML is an external trust boundary. If the source page is compromised, modified in transit at an upstream system, or contai ...[truncated 1809 chars]- Remediation
View remediation
