T01 · Skill Instruction Hijacking
- Location
scripts/aibase.py:131- Finding
Untrusted remote article HTML is forwarded directly to the AI Agent
- Content
View full analysis
str: try: resp = session.get(url, headers=HEADERS, timeout=15) resp.raise_for_status() soup = BeautifulSoup(resp.text, "html.parser") content_div = soup.find("div", class_="article-content") or soup.find("article") return content_div.decode_contents() if content_div else "" except Exception as e: print(f"Failed to retrieve article content from {url}: {e}", file=sys.stderr) return "" ``` The retrieved HTML is subsequently included in the generated result: ```python items = _extract_items(data, data_type)[:limit] results = [] for item in items: content = "" if with_content: article_url = f"https://news.aibase.cn/{data_type}/{item.get('oid', '')}" content = _fetch_article_content(session, article_url) results.append(_build_item(item, data_type, content if with_content else None)) return results ``` The content is added to the output object without sanitization: ```python if content is not None: result["content"] = content return result ``` ### Technical Analysis The Skill retrieves complete article content from a remote website and returns the HTML directly in JSON intended for consumption by an AI Agent. Although BeautifulSoup identifies an article container, `decode_contents()` preserves the remotely controlled markup and text inside that container. There is no trust-boundary annotation, instruction filtering, plain-text conversion, content-length restriction, or isolation between the remote article and the Agent's instruction context. An attacker who controls or compromises an AIBase article can place instruc ...[truncated 1620 chars]- Remediation
View remediation
