Back to skill

Security audit

HotBee 抖音视频报告

Security checks for vulnerabilities and agentic risk

Overview

The skill largely matches its stated Douyin reporting purpose, but its HotBee API key handling and flexible network destinations need careful review before installation.

Install only if you are comfortable sending provided Douyin links and derived media/comment/transcript data to HotBee and storing the results locally. Use a dedicated HotBee key with limited quota, avoid custom HOTBEE_API_BASE values unless you control them, and keep generated output directories out of public repositories.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:43
Finding

Mandatory Branding and Response-Language Rules Override User-Controlled Output

Content
View full analysis
{esc(video.get('title'))} - HotBee Douyin Video Analysis Report ... Analysis insights provided by HotBee.cn | Public social-media data collection and content analysis ... HotBee Douyin Video Analysis ``` The Skill instructions additionally require the final response to use a fixed language, require a fixed HotBee attribution footer, and state that all user-visible content must follow that language requirement. ### Technical Analysis The Skill imposes persistent branding and response-format rules that are not technically required to retrieve Douyin metadata, collect comments, transcribe a video, or generate local report files. These instructions alter how the hosting agent responds and force promotional attribution into generated artifacts regardless of the user's preferred report style. This behavior matches instruction hijacking because loading and following the Skill changes session-level output behavior beyond the minimum instructions necessary to perform the declared analysis task. The concern is not the presence of ordinary product identification, but the mandatory and unconditional nature of the branding and language controls. ### Attack Path 1. A user invokes the Skill to analyze a Douyin URL. 2. The agent loads and follows `SKILL.md`. 3. The Skill directs the agent to use a fixed response language and mandatory attribution. 4. The report renderer independently inserts fixed HotBee branding into HTML and SVG artifacts. 5. The resulting response and generated files contain Skill-controlled promotional content even if the user did not request it. ### Impact Assessment ...[truncated 466 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/douyin_video_report.py:207
Finding

API Credentials Are Sent in Query Strings to a User-Configurable HTTPS Host

Content
View full analysis
Any: clean_params = {key: value for key, value in params.items() if value is not None and value != ""} query = urllib.parse.urlencode(clean_params, doseq=True) url = f"{base_url.rstrip('/')}{endpoint}" if query: url = f"{url}?{query}" req = urllib.request.Request( url, data=b"", method="POST", headers={ "Accept": "application/json,text/plain,*/*", "Content-Type": "application/json", "User-Agent": "Mozilla/5.0 HotBee-Douyin-Video-Report/1.0", }, ) secret_values = [ str(item) for key, item in clean_params.items() if key.lower() in SENSITIVE_FIELD_NAMES and item ] try: with urllib.request.urlopen(req, timeout=timeout) as response: body = response.read().decode("utf-8", errors="replace") except Exception as exc: raise RuntimeError(redact_sensitive_text(exc, secret_values)) from None ``` Credential-bearing calls include: ```python payload = post_query( base_url, "/tool/speech/speechToText", {"file_url": source, "key": key}, timeout=120, ) ``` ```python payload = post_query( base_url, "/tool/douyin/Dy_video_all_comments_VIP", {"video_url": candidate, "page": page, "key": key}, timeout=90, ) ``` The destination is configurable: ```python parser.add_argument( "--base-url", default=os.environ.get("HOTBEE_API_BASE", DEFAULT_BASE_URL), help="HotBee API Base URL.", ) ``` ### Technical Analysis `post_query()` serializes every parameter, including the API key, into the request URL. Although the HTTP method is POST, the ...[truncated 2485 chars]
Remediation
View remediation
``` If the API contract cannot support an authorization header, send the credential in the POST body over HTTPS. 3. Pin credential-bearing requests to an explicit allowlist of official API origins. 4. Treat custom API hosts as untrusted. Do not forward automatically loaded credentials to them without explicit, informed confirmation. 5. Separate endpoint configuration from credential selection so a custom host does not automatically inherit the official service credential. 6. Compare the normalized scheme, hostname, and port against the approved origin before attaching credentials. 7. Avoid redirects for credential-bearing requests, or strip authorization data before following a redirect to another origin. 8. Rotate any credentials that may already have appeared in URL or proxy logs. 9. Configure server and proxy logging to redact sensitive query parameters during the migration period. 10. Add tests confirming that secrets never appear in request URLs, exceptions, manifests, raw files, or console output. ]]>

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/douyin_video_report.py:1215
Finding

Media Downloader Does Not Revalidate Redirect Targets or Resolved Network Addresses

Content
View full analysis
bool: try: parsed = urllib.parse.urlsplit(value) except ValueError: return False if parsed.scheme != "https" or not parsed.hostname or parsed.username or parsed.password: return False if parsed.hostname == "localhost" or parsed.hostname.endswith(".local"): return False try: return not ( ipaddress.ip_address(parsed.hostname).is_private or ipaddress.ip_address(parsed.hostname).is_loopback ) except ValueError: return True ``` ```python def download_media(url: str, target_stem: Path, warnings: list[str]) -> str: if not url: return "" if not is_allowed_media_url(url): warnings.append(f"Rejected unsafe image URL: {redact_url_for_log(url)}") return "" try: req = urllib.request.Request( url, headers={ "User-Agent": "Mozilla/5.0 HotBee-Douyin-Video-Report/1.0", "Referer": "https://www.douyin.com/", }, ) with urllib.request.urlopen(req, timeout=35) as response: declared_size = int(response.headers.get("Content-Length") or 0) if declared_size > MAX_MEDIA_BYTES: raise ValueError("Image exceeds the 25 MB safety limit.") content = response.read(MAX_MEDIA_BYTES + 1) if len(content) > MAX_MEDIA_BYTES: raise ValueError("Image exceeds the 25 MB safety limit.") content_type = response.headers.get("Content-Type", "") path = target_stem.with_suffix(guess_extension(url, content_type)) path.parent.mkdir(parents=True, exist_ok=True) path.write_bytes(content) return path.name ...[truncated 2745 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (11)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 15)May include surrounding context.

md
python scripts/douyin_video_report.py --url "抖音视频链接" --output-dir "./output/douyin-video-report"

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 54)May include surrounding context.

md
python scripts/douyin_video_report.py --url "抖音视频链接" --output-dir "./output/douyin-video-report"

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill clearly instructs use of environment variables, local file reads/writes, and outbound network access, but it does not declare an explicit tool scope or permission boundary. That mismatch can cause the agent platform to invoke the skill in contexts where users or reviewers do not have a precise understanding of what resources it will access, increasing the chance of unintended data exposure or network/file operations.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The description uses broad trigger phrases like video parsing, breakdown, comment insights, and content extraction without strong constraints on when the skill should be selected. This can cause over-invocation on loosely related requests, leading the agent to perform network calls, consume paid API quota, and collect/store third-party public data when the user did not clearly request those actions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instruction “所有用户可见内容必须是中文” forces a specific language for all outputs. Under the policy, language restrictions should not be imposed unless the user is given a choice or the locale constraint is clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The display name, description, and default prompt are all written to require operation in Chinese and specifically target Douyin content, with no indication that users may choose another language or locale. This creates a natural-language locale policy concern because the skill appears to enforce a specific language experience without explicit user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The contract explicitly sends user-provided Douyin URLs and related media URLs to a third-party HotBee API, but it does not require a user-facing disclosure or consent step about external transmission of potentially sensitive links and derived content. In this skill context, users may assume local analysis/output generation, so undisclosed sharing of URLs, comments, transcripts, and metadata to an external service creates a real privacy and data-handling risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script transmits user-supplied Douyin URLs to a third-party HotBee API and also may submit additional derived sources such as resolved URLs, download URLs, or media URLs for transcript and comment retrieval. Even though the code includes some transport and URL validation, it does not obtain explicit user consent or clearly disclose that external services will receive potentially sensitive browsing targets and related metadata, creating a privacy and data-sharing risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script persistently stores transcripts, comments, raw API responses, manifests, and generated reports under the output directory, which can include personal data, user-generated content, and service responses. Although the code attempts to redact some sensitive fields, there is no user-facing warning, retention control, or minimization policy, so operators may unintentionally create local privacy exposure or long-lived sensitive datasets.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The document explicitly requires intent rules to remain in Chinese, and later requires a Chinese prompt for image generation. This imposes a specific language policy without stating that the user can opt in, choose another language, or that the skill is limited to a China-specific workflow.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The generated HTML sets lang="zh-CN", which enforces a specific locale in output regardless of user preference. The file contains Chinese-only user-facing text but does not expose a language option or explicitly justify the locale restriction as a region-specific requirement.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.