Back to skill

Security audit

网页内容抓取服务

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly implements a web-fetch proxy, but it has enough credential-storage and documentation-scope issues that users should review it before installing.

Install only if you trust XiaoBenYang with the URLs and parameters you fetch, and avoid using a sensitive or broadly privileged API key. Expect the key to be stored in a local plaintext `.env` file unless the skill is changed to use safer secret storage.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/config.py:45
Finding

API Key Persisted in a Plaintext File Without Enforced Access Controls

Content
View full analysis

Vulnerability Details

File Location: scripts/config.py, lines 45-67
Vulnerability Type: Plaintext sensitive-data storage
Risk Level: Medium

Vulnerable Code

python
def save_api_key_to_env(api_key: str) -> bool:
    """将API key保存到.env文件"""
    try:
        env_path = Path(".env")
        lines = []
        if env_path.exists():
            lines = env_path.read_text(encoding="utf-8").splitlines()
        found = False
        new_lines = []
        for line in lines:
            if line.startswith("XBY_APIKEY="):
                new_lines.append(f"XBY_APIKEY={api_key}")
                found = True
            else:
                new_lines.append(line)
        if not found:
            new_lines.append(f"XBY_APIKEY={api_key}")
        env_path.write_text("\n".join(new_lines) + "\n", encoding="utf-8")
        os.environ["XBY_APIKEY"] = api_key
        return True
    except Exception as e:
        print(f"保存API key失败: {e}")
        return False

Technical Analysis

The function persistently writes the user-supplied API key to a plaintext .env file. It does not explicitly enforce owner-only permissions, use a protected credential store, or validate that the destination is a regular file at a trusted location. The relative Path(".env") destination also depends on the process working directory.

Consequently, the file's effective permissions depend on the process umask, existing file permissions, and working-directory protections. In a shared or improperly configured environment, another local user or process may be able to read the credential. Persisting the same secret in os.environ additionally makes it available to code running inside the process and potentially to subsequently launched child processes.

The API key is legitimately transmitted in the XBY-APIKEY header to the configured HTTPS service in scripts/call_api.py; that network transmission is ne ...[truncated 1512 chars]

Remediation
View remediation

Remediation Suggestions

  1. Prefer session-only credential handling or an operating-system credential store instead of writing the key to the project directory.
  2. If file persistence is necessary, use a fixed, user-specific configuration directory rather than a working-directory-relative path.
  3. Create the credential file atomically with owner-only permissions such as 0600, and verify or correct permissions when updating an existing file.
  4. Refuse symbolic links and verify that the destination is a regular file owned by the expected user before reading or writing it.
  5. Avoid propagating the secret through environment variables unless required, particularly when the process may launch child processes.
  6. Ensure .env is excluded from source control, build artifacts, logs, backups, and shared workspace exports.
  7. Never include the API key in exceptions, diagnostic output, or returned tool data, and document credential rotation procedures for suspected exposure.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (29)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill claims to fetch webpages and convert HTML to Markdown, but the file also documents reading environment variables, writing .env, and managing API credentials without showing the claimed transformation behavior. This creates a deceptive trust boundary where operators may enable file and secret access for a seemingly simple content utility that in practice performs broader configuration and credential operations.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill claims to fetch webpages and convert HTML to Markdown, but the file also documents reading environment variables, writing .env, and managing API credentials without showing the claimed transformation behavior. This creates a deceptive trust boundary where operators may enable file and secret access for a seemingly simple content utility that in practice performs broader configuration and credential operations.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes a web-content fetching service, but the body documents an unrelated API-keyed gaokao skill, including different project naming and examples unrelated to webpage conversion. Such internal inconsistency is a strong indicator of copy-paste drift or deceptive packaging, and it increases the chance that users will expose credentials or approve actions under false assumptions about the skill's purpose.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file claims the agent will only select the webpage-fetching tool, but adjacent instructions and examples reference unrelated API operations, creating contradictory operational guidance. This ambiguity is dangerous because agents or users may be induced to follow hidden or broader behaviors not justified by the stated skill purpose, especially around credential collection and external service usage.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

This file implements a generic upstream MCP API client rather than a narrowly scoped webpage-fetching/HTML-to-Markdown service. Because it accepts arbitrary tool names and parameters and forwards them upstream, the actual capability surface is much broader than the declared skill purpose, creating a dangerous mismatch between stated functionality and effective behavior.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The call_tool method exposes a generic remote invocation primitive: callers can choose any tool_name and pass arbitrary params, which are then sent to the upstream service with authentication. In the context of a skill advertised as webpage content scraping, this expands into undeclared remote action execution and could be abused to access unrelated privileged MCP functions or cause unintended side effects.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config.py (reported line 13)May include surrounding context.

python
model_config = SettingsConfigDict(
        env_prefix="XBY_GAOKAO_",
        env_file=".env",
        env_file_encoding="utf-8",
        extra="ignore",
    )

Credential Access

High
Category
Privilege Escalation
Confidence
90% confidence
Finding

The model_post_init method manually reads .env and extracts XBY_APIKEY outside the standard settings flow, adding bespoke secret-parsing behavior. In the context of a web-content-fetching skill, this extra credential access is unnecessary and increases the chance of mishandling secrets or expanding scope beyond user expectations.

Content

Scanner excerpt · scripts/config.py (reported line 31)May include surrounding context.

python
default_year: int = 2025

    def model_post_init(self, __context):
        # 强制从 .env 文件读取 XBY_APIKEY
        env_path = Path(".env")
        if env_path.exists():
            content = env_path.read_text(encoding="utf-8")

Credential Access

High
Category
Privilege Escalation
Confidence
90% confidence
Finding

This line initiates direct access to the local .env file for credential retrieval, contributing to custom secret-handling code. While reading a secret is not inherently malicious, in this skill it is more dangerous because the declared purpose is unrelated to local credential management and the code later persists and mutates those credentials.

Content

Scanner excerpt · scripts/config.py (reported line 32)May include surrounding context.

python
def model_post_init(self, __context):
        # 强制从 .env 文件读取 XBY_APIKEY
        env_path = Path(".env")
        if env_path.exists():
            content = env_path.read_text(encoding="utf-8")
            for line in content.splitlines():

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config.py (reported line 39)May include surrounding context.

python
if line.startswith("XBY_APIKEY="):
                    self.api_key = line.split("=", 1)[1].strip()
                    break
        # 如果环境变量有值,覆盖 .env 的值
        env_val = os.getenv("XBY_APIKEY", "")
        if env_val:
            self.api_key = env_val

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The code persistently stores an API credential in a local .env file and exposes helper functions to update that file at runtime. For a skill described as a webpage scraping/HTML-to-Markdown service, local credential persistence is unnecessary to the stated purpose and increases the chance of credential leakage through source packaging, backups, logs, or unintended file access.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The functions both write credentials to disk and mutate the current process environment, creating secret-handling side effects unrelated to a simple content-fetching service. This broadens the attack surface because other code in the same process may read the injected secret, and local files may retain the credential after execution.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

This function explicitly saves an API key into a plaintext .env file, creating durable local secret storage. In a skill whose advertised purpose is webpage content extraction, such credential persistence is out of scope and materially raises the risk of secret disclosure through filesystem access, accidental commits, archives, or other tooling.

Content

Scanner excerpt · scripts/config.py (reported line 48)May include surrounding context.

python
def save_api_key_to_env(api_key: str) -> bool:
    """将API key保存到.env文件"""
    try:
        env_path = Path(".env")
        lines = []
        if env_path.exists():
            lines = env_path.read_text(encoding="utf-8").splitlines()

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill exposes or implies capabilities to access environment variables, read/write local files, and perform network requests, but it does not declare any explicit tool scope or permission boundaries. This weakens least-privilege controls and makes it harder for reviewers or platforms to constrain what the skill may access, especially since the document also instructs credential storage and external API use.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill instructs the agent to ask the user for an API key and then persist it via local configuration without disclosing storage location, retention, access controls, or masking requirements. This is dangerous because it normalizes secret collection and durable storage in a user-facing workflow, increasing the risk of credential leakage through files, logs, backups, or other tools.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This code performs a network POST to an upstream service and transmits the caller-provided params payload, which may contain user or system data. While the module logs success and error states, it does not provide any user-facing warning, confirmation, or explanatory notice about data being sent externally.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The function persists an API key to a local .env file without any user-facing notice, confirmation, or indication of retention. Silent persistence of credentials is dangerous because users may believe the key is transient while it remains on disk and becomes accessible to other local processes, users, or future packaging steps.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
73% confidence
Finding

The skill name, description, and operating instructions are entirely in Chinese, and no opt-in or alternative language behavior is described. Under the stated policy, forcing a specific language without user choice can be a natural-language policy issue unless the locale limitation is explicitly justified.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
95% confidence
Finding

The dependency is specified with a minimum version only, which allows future installs to resolve to different releases over time. This weakens supply-chain reproducibility and can unintentionally introduce vulnerable or breaking versions, especially for a network-facing web scraping service that relies on HTTP behavior.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.31.0
pydantic>=2.7.0
pydantic-settings>=2.2.0
python-dotenv>=1.0.1

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

Requests has multiple published advisories, and because the manifest does not pin a version, it is impossible to determine whether an affected release may be installed. In a web content fetching service, this matters more because requests is directly exposed to attacker-controlled URLs and remote servers.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

Using an unpinned pydantic version makes builds non-reproducible and may allow a later vulnerable or incompatible release to be installed. Since this package is commonly used for parsing and validation, unexpected version drift can affect security-sensitive input handling.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests>=2.31.0
pydantic>=2.7.0
pydantic-settings>=2.2.0
python-dotenv>=1.0.1

Unverifiable Dependency: pydantic has 4 known advisory(ies) (CVE-2021-29510 (Use of "infinity" as an input to datetime and date fields causes infinite loop i); CVE-2024-3772 (Pydantic regular expression denial of service); CVE-2021-29510 (Pydantic is a data validation and settings management using Python type hinting.) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
84% confidence
Finding

Pydantic has known advisories, and the absence of version pinning means the deployed version cannot be verified as safe. Since validation libraries often process untrusted input, unresolved version ambiguity creates avoidable risk even if no direct exploit is shown in this file.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
92% confidence
Finding

An unpinned pydantic-settings dependency permits uncontrolled upgrades, making it unclear which code will run in deployment. For configuration-loading libraries, version drift can be risky because changes may affect secrets loading and file resolution behavior.

Content

Scanner excerpt · requirements.txt (reported line 3)May include surrounding context.

text
requests>=2.31.0
pydantic>=2.7.0
pydantic-settings>=2.2.0
python-dotenv>=1.0.1

Unverifiable Dependency: pydantic-settings has 1 known advisory(ies) (CVE-2026-58203 (pydantic-settings: NestedSecretsSettingsSource follows symlinks outside secrets_)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
83% confidence
Finding

Pydantic-settings has a known advisory, and without an exact version the project cannot prove it avoids the affected range. This is relevant because settings packages may read secrets or files from disk, so vulnerable resolution behavior could expose sensitive configuration.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
94% confidence
Finding

The python-dotenv requirement is not pinned, so deployments may install different versions with different security properties. Because dotenv libraries interact with local configuration and environment files, unexpected vulnerable versions can increase the chance of configuration tampering or unsafe file handling.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
requests>=2.31.0
pydantic>=2.7.0
pydantic-settings>=2.2.0
python-dotenv>=1.0.1

Static analysis

No suspicious patterns detected.