Back to skill

Security audit

guaikei-xhs-data-tool

Security checks across malware telemetry and agentic risk

Overview

This skill is a disclosed Xiaohongshu public-data tool, but it needs Review because it can bulk collect social content, auto-save results locally, and may be invoked too broadly.

Install only if you are comfortable sending Xiaohongshu keywords or URLs to guaikei.com, using a provider API token, and retaining fetched public social-media results in local logs. Use it for explicit Xiaohongshu tasks, avoid personal or sensitive investigations, and review or delete the logs directory after use.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (25)

Description-Behavior Mismatch

Medium
Confidence
82% confidence
Finding
The package metadata markets the skill as a broad Xiaohongshu analytics, competitor monitoring, KOL screening, and precision marketing tool, while the stated skill capability is much narrower: keyword-based search over public notes. This mismatch can cause over-invocation, user deception, or downstream policy bypass by making operators and agents believe the tool is authorized for broader profiling and marketing use cases than it actually is.

Description-Behavior Mismatch

Medium
Confidence
90% confidence
Finding
The README advertises broader capabilities such as competitor monitoring, trend prediction, KOL screening, and comment analysis than the manifest says the skill should do. This scope mismatch is dangerous because users and orchestrators may authorize a seemingly narrow keyword-search skill while it actually encourages broader social-media intelligence collection, increasing privacy, compliance, and misuse risk.

Description-Behavior Mismatch

Medium
Confidence
95% confidence
Finding
The documented commands include note detail retrieval, profile/post monitoring, and comment analysis, which exceed the manifest's stated keyword-search-only scope. This creates a deceptive capability boundary: a consumer may install or trust the skill for simple search, but the documented operations support much deeper collection and monitoring of public user content.

Intent-Code Divergence

Medium
Confidence
95% confidence
Finding
The changelog claims the skill can fetch Xiaohongshu comment information, while the published skill metadata describes a narrower capability limited to searching public notes by keyword. This mismatch can cause downstream agents or reviewers to trust undocumented features and invoke data-collection behavior outside the declared scope, increasing the risk of unauthorized processing and policy bypass.

Intent-Code Divergence

High
Confidence
97% confidence
Finding
The changelog states support for blogger work monitoring/scraping, which is materially broader and more privacy-sensitive than simple keyword search over public notes. This kind of capability drift is dangerous because it suggests hidden scraping or persistent tracking functionality that operators and users may not expect, potentially enabling surveillance-style collection beyond the approved use case.

Intent-Code Divergence

Medium
Confidence
93% confidence
Finding
The version history says the skill was broadened from a search tool into a more general Xiaohongshu tool, conflicting with the current manifest's narrower description. In security review, inconsistent scope declarations are risky because they can conceal latent capabilities, weaken least-privilege assumptions, and make policy enforcement depend on incomplete metadata.

Description-Behavior Mismatch

High
Confidence
95% confidence
Finding
The file implements creation of Xiaohongshu comment-scraping tasks, which materially exceeds the declared skill purpose of keyword-based search over public notes. Capability drift like this is dangerous because it introduces undeclared data collection functionality that could be used to gather user-generated comments at scale without clear user expectation or policy justification.

Context-Inappropriate Capability

Medium
Confidence
88% confidence
Finding
The result-fetching path continues the undeclared comment-harvesting workflow, reinforcing that the skill can collect more than high-level content discovery data. Even if comments are public, collecting them through an undocumented feature expands data processing scope and can create privacy, compliance, and trust issues.

Description-Behavior Mismatch

Medium
Confidence
93% confidence
Finding
The file adds functionality for retrieving full note details and comments from a direct URL, which materially exceeds the declared skill behavior of keyword-based public note search returning listing metadata. This scope expansion is dangerous because downstream agents, reviewers, or users may grant the skill broader access and use than intended, enabling collection of richer user-generated content and profile-linked data without that capability being transparently disclosed in the manifest.

Description-Behavior Mismatch

Medium
Confidence
90% confidence
Finding
The implementation performs blogger-profile and published-post retrieval by URL, which materially differs from the declared skill purpose of keyword-based public note discovery. This kind of capability drift is dangerous because it expands data access and user expectations beyond the reviewed manifest, enabling collection of account-specific content and metadata under a misleading description.

Intent-Code Divergence

Medium
Confidence
86% confidence
Finding
The file comments explicitly describe blogger-detail and published-notes functionality that contradicts the stated skill behavior. Misleading internal documentation increases the risk that reviewers, operators, and downstream agents misunderstand the actual data collection scope, allowing unauthorized or unexpected scraping behavior to persist unnoticed.

Context-Inappropriate Capability

Medium
Confidence
93% confidence
Finding
This file provides arbitrary local file-writing capability via taskWrite(), which is not clearly necessary for a skill whose stated purpose is searching public Xiaohongshu content and returning results. Although the filename is partially sanitized, the function still enables persistence of attacker-controlled content on disk, increasing the attack surface for data leakage, log poisoning, storage abuse, or unintended local state changes if exposed to untrusted inputs.

Description-Behavior Mismatch

Medium
Confidence
92% confidence
Finding
The CLI performs note-detail retrieval and returns full task results, which exceeds the manifest description that limits the skill to keyword-based discovery of public Xiaohongshu content. This scope mismatch is dangerous because it enables collection of richer per-note data than users and reviewers would reasonably expect, weakening consent and policy controls around what the skill is supposed to do.

Context-Inappropriate Capability

Medium
Confidence
95% confidence
Finding
The --limit flag allows collecting up to 10,000 comments for a single note, which is disproportionate to the stated use case of understanding keyword/topic performance on social media. Even if the content is public, bulk comment harvesting materially increases privacy and compliance risk by enabling large-scale aggregation of user-generated content beyond the minimally necessary dataset.

Description-Behavior Mismatch

High
Confidence
96% confidence
Finding
The implemented interface requires a blogger profile URL and is designed to fetch posts from a specific account, which materially differs from the declared skill purpose of keyword/topic discovery over public notes. This mismatch increases the likelihood of unauthorized profile-focused scraping and can mislead reviewers or users about what data the tool actually collects.

Context-Inappropriate Capability

Medium
Confidence
91% confidence
Finding
The code validates profile URLs, creates a post collection task, and retrieves a blogger's posts, enabling account-centric scraping unrelated to the stated use case of topic trend discovery. In context, that makes the skill more dangerous because the declared social-topic analysis purpose does not justify targeted harvesting of an individual's/public account history.

Context-Inappropriate Capability

Medium
Confidence
95% confidence
Finding
The CLI writes a JSON file containing the user's keyword and the full search results to local storage after completing a read-only search. For a social-search skill, this creates unnecessary data retention and expands exposure to local disclosure, especially if queries contain sensitive business topics, personal interests, or investigatory terms.

Vague Triggers

Medium
Confidence
86% confidence
Finding
The trigger text says the skill should be used even when the user does not explicitly mention Xiaohongshu, as long as the intent is to understand a term's performance on social media. That can cause over-broad invocation on generic social-media analysis requests, unintentionally sending user queries or URLs to a third-party service when the user did not clearly ask for Xiaohongshu-specific processing.

Vague Triggers

Medium
Confidence
76% confidence
Finding
The description is overly broad and does not state when the tool should or should not be invoked, despite the skill definition saying it is intended for keyword/topic discovery on public Xiaohongshu content and not for SEO or ad targeting. Ambiguous metadata increases the chance an agent will route unrelated marketing, profiling, or growth-hacking tasks to this skill, expanding data access and use beyond intended boundaries.

Missing User Warnings

Medium
Confidence
86% confidence
Finding
The README states that all task results are automatically saved to the logs directory, but it does not clearly disclose that fetched social-media data will be persisted to disk. Silent local retention increases the risk of unintended storage, later exfiltration, over-collection, and mishandling of scraped content, especially on shared systems or developer machines.

Missing User Warnings

Medium
Confidence
91% confidence
Finding
The GET request places the API token in request parameters, which increases the chance of credential exposure through logs, proxies, browser/network tooling, or intermediary systems that record URLs. Tokens in URLs are widely considered unsafe because they are more likely to persist in observability pipelines and infrastructure metadata than header-based authentication.

Missing User Warnings

Low
Confidence
84% confidence
Finding
The CLI writes fetched comment results to a local JSON file automatically, without an explicit consent prompt or clear warning at the write site. Because comment content may include user-generated text and metadata, silent persistence can create unintended local data retention, leakage to other local users/processes, or accidental inclusion in backups and logs.

Missing User Warnings

Medium
Confidence
91% confidence
Finding
The tool writes fetched detail results to a local JSON file automatically, but the user is not clearly warned that retrieved content will be persisted on disk. Silent persistence increases the chance of unintended retention, later disclosure, or mishandling of scraped content and metadata, especially on shared systems or in automated pipelines.

Missing User Warnings

Medium
Confidence
88% confidence
Finding
The CLI persists fetched results to a local JSON file automatically, but does not clearly warn the user that data will be written to disk. This can create unintended retention of scraped content, links, author information, and interaction data, increasing privacy, compliance, and workstation data exposure risks.

Missing User Warnings

Medium
Confidence
96% confidence
Finding
This file silently writes search output to a local JSON log without warning the user in normal execution flow. Undisclosed persistence is risky because users may reasonably expect a search utility to return results transiently, not create artifacts containing their query history and retrieved content.

VirusTotal

VirusTotal findings are pending for this skill version.

View on VirusTotal

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
SKILL.md:16