Back to skill

Security audit

网页内容提取工具

Security checks for vulnerabilities and agentic risk

Overview

The skill appears to be a remote webpage extractor, but it needs review because it persists an API key in plaintext and can send sensitive URL or cookie-path data to an external service with weak endpoint controls.

Install only if you trust XiaoBenYang with the URLs you submit and with the API key. Avoid authenticated, internal, tokenized, or private URLs, do not provide browser cookie files, and treat any saved .env file as a secret that should not be committed, synced, or shared.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/call_api.py:49
Finding

API Credential Can Be Redirected to an Attacker-Controlled Endpoint

Content
View full analysis

Vulnerability Details

File Location: scripts/config.py:14-21; scripts/call_api.py:49-67
Vulnerability Type: Configurable credential destination without endpoint validation
Risk Level: High

Vulnerable Code

python
# scripts/config.py:14-21
model_config = SettingsConfigDict(
    env_prefix="XBY_GAOKAO_",
    env_file=".env",
    env_file_encoding="utf-8",
    extra="ignore",
)

# API configuration
base_url: str = "https://mcp.xiaobenyang.com"
python
# scripts/call_api.py:49-67
url = f"{settings.base_url}/api"
mcp_id = mcp_id or settings.mcp_id

api_key = get_api_key()
if not api_key:
    raise UpstreamError("API密钥未设置,请先调用 set_api_key()")

headers = {
    "XBY-APIKEY": api_key,
    "func": tool_name,
    "mcpid": mcp_id,
    "Content-Type": "application/json",
}

t0 = time.time()
try:
    resp = self._session.post(
        url=url,
        headers=headers,
        data=json.dumps(params),
        timeout=settings.timeout_seconds,
    )

Technical Analysis

SettingsConfigDict allows settings to be overridden through variables using the XBY_GAOKAO_ prefix. Consequently, XBY_GAOKAO_BASE_URL can replace the intended API origin. The client then attaches the user's XBY-APIKEY credential to a request sent to that configurable origin.

The implementation does not enforce HTTPS, verify that the destination hostname is mcp.xiaobenyang.com, or apply an explicit endpoint allowlist. It also does not explicitly disable or validate redirects. An actor who controls the launch environment or .env configuration can therefore redirect the credential and extraction parameters to an endpoint under their control.

Attack Path

  1. An attacker gains control over the Skill's environment variables or writable .env configuration.
  2. The attacker sets XBY_GAOKAO_BASE_URL to an attacker-controlled HTTP or HTTPS server.
  3. The user configures a ...[truncated 861 chars]
Remediation
View remediation

Remediation Suggestions

  • Make the production API endpoint immutable unless endpoint customization is explicitly required.
  • If customization is necessary, parse the URL and enforce:
    • The https scheme.
    • An exact hostname allowlist, such as mcp.xiaobenyang.com.
    • The expected port and path.
    • Rejection of embedded credentials and malformed hostnames.
  • Disable automatic redirects for authenticated requests, or validate every redirect destination before resending credentials.
  • Separate endpoint selection from secret-bearing request construction so credentials are attached only after origin validation.
  • Treat .env and process-environment configuration as untrusted input.
  • Add tests proving that HTTP endpoints, lookalike domains, user-information URL tricks, and cross-origin redirects are rejected.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/config.py:44
Finding

API Key Is Persisted in a Plaintext File Without Explicit Permission Hardening

Content
View full analysis

Vulnerability Details

File Location: scripts/config.py:44-63; persistence is required by SKILL.md:16-18
Vulnerability Type: Insecure plaintext credential storage
Risk Level: Medium

Vulnerable Code

python
# scripts/config.py:44-63
def save_api_key_to_env(api_key: str) -> bool:
    """将API key保存到.env文件"""
    try:
        env_path = Path(".env")
        lines = []
        if env_path.exists():
            lines = env_path.read_text(encoding="utf-8").splitlines()
        found = False
        new_lines = []
        for line in lines:
            if line.startswith("XBY_APIKEY="):
                new_lines.append(f"XBY_APIKEY={api_key}")
                found = True
            else:
                new_lines.append(line)
        if not found:
            new_lines.append(f"XBY_APIKEY={api_key}")
        env_path.write_text("\n".join(new_lines) + "\n", encoding="utf-8")
        os.environ["XBY_APIKEY"] = api_key
        return True
    except Exception as e:
        print(f"保存API key失败: {e}")
        return False

Technical Analysis

The API key is written directly into .env as plaintext. The code relies on default filesystem creation permissions and does not explicitly enforce owner-only access, inspect an existing file's ownership or permissions, prevent symbolic-link traversal, or use atomic secure-file creation.

The Skill documentation requires saving the supplied key rather than offering session-only use. Persistent plaintext storage is not strictly necessary to make an authenticated API request and expands the period and number of contexts in which the secret can be exposed.

If the current directory is shared, backed up, synchronized, or tracked by version control, the key may be available to unrelated users or systems. If an attacker can prepare the working directory, the non-atomic path-based write may also target an unintended location through a symbo ...[truncated 1237 chars]

Remediation
View remediation

Remediation Suggestions

  • Prefer session-only credential handling or an operating-system credential manager.
  • Obtain explicit user consent before persisting a secret.
  • If file storage is unavoidable:
    • Create the file atomically with owner-only mode 0600.
    • Verify file ownership and reject symbolic links.
    • Store it in a dedicated user configuration directory rather than an arbitrary current directory.
    • Preserve restrictive permissions when updating an existing file.
    • Add .env to version-control ignore rules and document that it must not be committed.
  • Avoid retaining the key in more locations than necessary.
  • Provide key rotation and deletion instructions in case the workspace or file is exposed.

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/tools.py:7
Finding

Sensitive URL Data and Local Cookie-File Paths Are Disclosed to a Third-Party API

Content
View full analysis

Vulnerability Details

File Location: scripts/tools.py:7-30; scripts/call_api.py:61-67
Vulnerability Type: Excessive transmission of user and local-system metadata
Risk Level: Medium

Vulnerable Code

python
# scripts/tools.py:7-30
def read_website(
    url: str,
    pages: Optional[float] = 1.0,
    cookiesFile: Optional[str] = None
) -> Dict[str, Any]:
    """
    Fast, token-efficient web content extraction - ideal for reading documentation, analyzing content, and gathering information from websites. Converts to clean Markdown while preserving links and structure.
    
    Args:
        url: HTTP/HTTPS URL to fetch and convert to markdown
        pages: Maximum number of pages to crawl (default: 1)
        cookiesFile: Path to Netscape cookie file for authenticated pages
    
    Returns:
        
    """
    arguments = {
        "url": url,
        "pages": pages,
        "cookiesFile": cookiesFile
    }
    
    return call_api("1777316659753987", "read_website", arguments)
python
# scripts/call_api.py:61-67
try:
    resp = self._session.post(
        url=url,
        headers=headers,
        data=json.dumps(params),
        timeout=settings.timeout_seconds,
    )

Technical Analysis

The Skill serializes and transmits the complete requested URL and the cookiesFile value to mcp.xiaobenyang.com. A complete URL can include bearer tokens, signed query parameters, session identifiers, internal hostnames, personal identifiers, or private resource names.

The cookiesFile parameter is documented as a local Netscape cookie-file path, but only the path string is sent to the remote API. A remote service cannot ordinarily access a file on the user's local filesystem merely from its path. Transmitting the path therefore reveals local directory structure or usernames without providing the documented authenticated-page functionality. This ...[truncated 1556 chars]

Remediation
View remediation

Remediation Suggestions

  • Clearly disclose that extraction requests are processed by a third-party API before sending user-provided data.
  • Request explicit confirmation before transmitting potentially sensitive or authenticated URLs.
  • Reject URLs containing embedded username/password credentials.
  • Detect and warn about sensitive query parameters such as token, key, signature, auth, and session identifiers.
  • Remove cookiesFile unless the remote protocol has a documented, secure, and functional mechanism for authenticated extraction.
  • Never transmit a local filesystem path when the remote service cannot use it.
  • If authenticated extraction is required, design a narrowly scoped mechanism that:
    • Requires explicit user consent.
    • Sends only cookies necessary for the selected origin.
    • Redacts unrelated cookies and metadata.
    • Uses encrypted transport and documented retention controls.
    • Avoids exposing the local source-file path.
  • Document the remote service's data retention and privacy expectations.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (30)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill claims to extract webpage content, but the project structure and instructions indicate reading and writing API keys in .env and managing configuration for a different external service. Persisting credentials for an unrelated service under a misleading description creates a serious trust and secret-handling risk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill claims to extract webpage content, but the project structure and instructions indicate reading and writing API keys in .env and managing configuration for a different external service. Persisting credentials for an unrelated service under a misleading description creates a serious trust and secret-handling risk.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config.py (reported line 13)May include surrounding context.

python
model_config = SettingsConfigDict(
        env_prefix="XBY_GAOKAO_",
        env_file=".env",
        env_file_encoding="utf-8",
        extra="ignore",
    )

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config.py (reported line 32)May include surrounding context.

python
model_config = SettingsConfigDict(
        env_prefix="XBY_GAOKAO_",
        env_file=".env",
        env_file_encoding="utf-8",
        extra="ignore",
    )

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

The module defines external service connection settings and an API key for a skill described as a web-to-Markdown extractor, which is a clear scope mismatch. Hidden outbound service configuration increases the risk that data or credentials are sent to an unrelated backend without the user's informed consent.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
89% confidence
Finding

The post-init hook manually reads .env to extract XBY_APIKEY, bypassing normal settings abstractions and broadening direct file access to credentials. In the context of a web content extraction skill, this unnecessary secret-reading behavior is suspicious and increases the chance of mishandled credentials.

Content

Scanner excerpt · scripts/config.py (reported line 31)May include surrounding context.

python
default_year: int = 2025

    def model_post_init(self, __context):
        # 强制从 .env 文件读取 XBY_APIKEY
        env_path = Path(".env")
        if env_path.exists():
            content = env_path.read_text(encoding="utf-8")

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config.py (reported line 39)May include surrounding context.

python
if line.startswith("XBY_APIKEY="):
                    self.api_key = line.split("=", 1)[1].strip()
                    break
        # 如果环境变量有值,覆盖 .env 的值
        env_val = os.getenv("XBY_APIKEY", "")
        if env_val:
            self.api_key = env_val

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code persists an API key into a local .env file, creating a durable credential store on disk for a skill that does not appear to require long-lived secrets for its stated purpose. This increases exposure to accidental disclosure through source control, backups, local compromise, or later reuse by unrelated processes.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

This function explicitly saves an API key into a .env file, creating persistent local storage of sensitive credentials. In a skill whose advertised purpose is HTML-to-Markdown extraction, this capability is disproportionate and materially increases the risk of credential leakage or unauthorized reuse.

Content

Scanner excerpt · scripts/config.py (reported line 48)May include surrounding context.

python
def save_api_key_to_env(api_key: str) -> bool:
    """将API key保存到.env文件"""
    try:
        env_path = Path(".env")
        lines = []
        if env_path.exists():
            lines = env_path.read_text(encoding="utf-8").splitlines()

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill declares no explicit tool scope or permissions despite requiring environment access, file reads/writes, and network use. This weakens least-privilege controls and makes it harder for users or platforms to understand that the skill can persist secrets and transmit data externally.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The skill metadata and instructions are entirely presented in Chinese, and there is no statement that the skill supports user language preference or that Chinese is a required locale for a region-specific compliance reason. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The instructions require collecting a user API key and then directly displaying raw tool output. If the API response echoes credentials, tokens, account metadata, request headers, or other sensitive fields, the model is explicitly told to expose them back to the user or chat transcript, creating a secret-disclosure path.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The body of the skill references gaokao/school-query examples and return conventions unrelated to webpage extraction. Such internal inconsistency is a supply-chain red flag because it suggests copy-paste reuse, hidden functionality, or poor review, increasing the chance that the agent routes user input to unintended services.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill accepts a URL and optional cookies file for authenticated pages, but does not warn users that these may be transmitted to an external API/service. Sending authenticated browsing context off-platform can expose session data, private content, or internal resources if users do not understand the trust boundary.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The project structure explicitly names the package as a gaokao skill while the document markets it as a web extractor. This naming contradiction undermines transparency and suggests the skill may be repurposed code with residual behaviors or hidden dependencies not obvious to users.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This file implements a generic upstream MCP API client rather than a narrowly scoped webpage-to-Markdown extractor. In the context of a skill advertised as a web content extraction tool, that capability mismatch is dangerous because it enables hidden or expanded remote actions through a generic tool invocation surface, reducing transparency and increasing the chance of unauthorized data handling or unintended behavior.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code allows callers to supply arbitrary tool_name, mcp_id, and params, then forwards them to a remote service with an authenticated API key. This creates a generic authenticated remote procedure call primitive that can be repurposed far beyond webpage extraction, potentially invoking sensitive upstream tools or exfiltrating user-supplied data under the cover of a benign skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This request sends arbitrary parameters and the API key to an external service, but the file contains no control that limits what data may be transmitted or ensures that users understand the disclosure. In an agent skill, such silent outbound transmission increases privacy and supply-chain risk because prompts, URLs, page contents, or other sensitive inputs could be forwarded to a third party without meaningful minimization or visibility.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The class docstring refers to an unrelated '高考' skill context, which is inconsistent with the declared web content extraction functionality. Such provenance mismatch is a supply-chain red flag because it suggests code reuse from another project, increasing the chance of hidden, unnecessary, or unsafe behavior remaining in the package.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The function silently writes the supplied API key to .env and updates process environment state without any user-facing warning, consent flow, or indication of storage location. This can lead users or calling agents to unintentionally persist secrets on disk in insecure contexts.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The tool exposes authenticated web retrieval by accepting a raw cookie file path and forwarding it to an external API without any visible validation, consent prompt, scope restriction, or warning about transmitting session credentials. In an agent context, this can lead to accidental exfiltration or misuse of authentication material to access private web content on behalf of the user.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

The dependency specification for requests is unpinned and allows installation of any future version above 2.31.0. This weakens supply-chain control and makes builds non-reproducible, increasing the chance of pulling in a vulnerable or breaking release without review.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.31.0
pydantic>=2.7.0
pydantic-settings>=2.2.0
python-dotenv>=1.0.1

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

requests has multiple known advisories, and because the manifest does not pin an exact version, it is impossible to verify whether deployed environments will avoid affected releases. For a web content extraction tool that likely makes outbound HTTP requests, this uncertainty is more relevant because the package is central to the skill's functionality.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
92% confidence
Finding

The pydantic requirement uses a lower-bound-only specifier, so dependency resolution may select different versions over time. This creates supply-chain uncertainty and may expose the skill to known or future vulnerabilities in upstream releases.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests>=2.31.0
pydantic>=2.7.0
pydantic-settings>=2.2.0
python-dotenv>=1.0.1

Unverifiable Dependency: pydantic has 4 known advisory(ies) (CVE-2021-29510 (Use of "infinity" as an input to datetime and date fields causes infinite loop i); CVE-2024-3772 (Pydantic regular expression denial of service); CVE-2021-29510 (Pydantic is a data validation and settings management using Python type hinting.) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
84% confidence
Finding

pydantic has known advisories, but the current requirement does not constrain installation to a verified safe release. This leaves the environment vulnerable to resolver-selected versions that may include denial-of-service or parsing-related flaws.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.