T09 · Insecure Skill Coding Practices
- Location
scripts/fetch_transcript.py:61- Finding
Path Traversal Through Insufficient YouTube Video ID Validation
- Content
View full analysis
Vulnerability Details
File Location:
scripts/fetch_transcript.py:61-64,scripts/fetch_transcript.py:97-119
Vulnerability Type: Path traversal and arbitrary JSON file access
Risk Level: HighVulnerable Code
python # Handle youtu.be short links if "youtu.be" in url: path = urlparse(url).path return path.lstrip("/").split("?")[0]The resulting value is subsequently used directly as part of a filesystem path:
python def get_cache_path(video_id): """Get the cache file path for a video ID.""" return CACHE_DIR / f"{video_id}.json" def load_from_cache(video_id): """Load transcript from cache. Returns None if not cached.""" cache_path = get_cache_path(video_id) if cache_path.exists(): try: with open(cache_path, "r", encoding="utf-8") as f: return json.load(f) except (json.JSONDecodeError, IOError): return None return None def write_cache_data(video_id, cache_data): """Write cache data to disk.""" ensure_cache_dir() cache_path = get_cache_path(video_id) with open(cache_path, "w", encoding="utf-8") as f: json.dump(cache_data, f, ensure_ascii=False, indent=2)Technical Analysis
The short-URL branch checks whether the untrusted input contains the substring
youtu.be, rather than verifying that the parsed hostname is exactly an approved YouTube hostname. It then returns the entire URL path without enforcing the expected 11-character video ID format.Consequently, path components such as
../can become part ofvideo_id. The cache path is formed using:python CACHE_DIR / f"{video_id}.json"pathlibdoes not automatically confine this path toCACHE_DIR. Traversal components can therefore resolve to a JSON file outside the intended cache directory.The vulnerable value reaches both read and write operations. An external J ...[truncated 2017 chars]
- Remediation
View remediation
Remediation Suggestions
-
Parse the URL before making any domain decision and compare the normalized hostname against an explicit allowlist:
python ALLOWED_HOSTS = { "youtu.be", "youtube.com", "www.youtube.com", "m.youtube.com", } -
Extract only the expected path segment from short URLs.
-
Validate every extracted ID, regardless of URL format, against the canonical format:
python VIDEO_ID_RE = re.compile(r"^[A-Za-z0-9_-]{11}$") def validate_video_id(value): return value if VIDEO_ID_RE.fullmatch(value) else None -
Add defense-in-depth confinement before all cache operations:
python def get_cache_path(video_id): if not VIDEO_ID_RE.fullmatch(video_id): raise ValueError("Invalid YouTube video ID") cache_root = CACHE_DIR.resolve() cache_path = (cache_root / f"{video_id}.json").resolve() if cache_path.parent != cache_root: raise ValueError("Cache path escapes cache directory") return cache_path -
Use restrictive cache permissions where supported, such as a user-only cache directory and files created with mode
0600. -
Add tests covering traversal strings, misleading hostnames, encoded separators, extra path segments, empty IDs, and IDs longer or shorter than 11 characters.
-
