Back to skill

Security audit

OECD数据查询服务

Security checks for vulnerabilities and agentic risk

Overview

The skill looks like an OECD data helper, but it routes queries and an API key through a private XiaoBenYang service and stores that key in a local plaintext .env file.

Review this before installing. Use only a narrowly scoped, revocable XiaoBenYang API key, assume query contents and the key are visible to that service, and avoid installing where a plaintext .env credential could be indexed, backed up, committed, or shared. Prefer a version that directly uses official OECD endpoints or clearly documents the proxy, retention, and credential storage behavior.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

other

Warning
Location
scripts/call_api.py:49
Finding

OECD Queries and API Credentials Are Routed Through a Private Third-Party Service

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/config.py:43
Finding

API Key Is Persisted in a Plaintext Relative .env File Without Permission Hardening

Content
View full analysis
` to `.env` in the current working directory. 4. The file is created with permissions determined by the environment's umask or retains existing p ...[truncated 1259 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
requirements.txt:1
Finding

Dependencies Use Open-Ended Version Ranges Without a Lock File or Integrity Hashes

Content
View full analysis
=2.31.0 pydantic>=2.7.0 pydantic-settings>=2.2.0 python-dotenv>=1.0.1 ``` ### Technical Analysis Every dependency is specified using only a lower bound. A future installation may therefore resolve to versions that did not exist and were not reviewed when the Skill was audited. Transitive dependencies are also neither locked nor protected by integrity hashes. The package names shown are established packages, and no evidence of typosquatting, dependency confusion, a malicious package, or an unsafe package index was identified. The confirmed weakness is the absence of reproducible dependency resolution and artifact integrity verification. This configuration increases exposure to: - Future compromised package releases. - Unexpected breaking changes. - Vulnerable transitive dependency versions. - Differences between development and production installations. - Package-index or distribution-artifact compromise. ### Attack Path 1. The project is installed at a later date using `pip install -r requirements.txt`. 2. The resolver selects the newest versions satisfying the lower bounds. 3. A selected direct or transitive release contains malicious code or a newly introduced vulnerability. 4. Package installation hooks or imported runtime code execute in the installer or application context. 5. The compromised dependency gains the same privileges and data access as the Python process. This is a supply-chain exposure scenario rather than evidence that any currently named dependency is malicious. ### Impact Assessment If a resolved dependency is compromised, its code executes with the privileges of the user or service installing or running the Skill. It could access the plaintext API key, modify requests or responses, read files available to the process, or perform netwo ...[truncated 203 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (30)

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill is presented as a data-query interface but also instructs the agent to manage and persist API keys in a local .env file and expose configuration interfaces unrelated to OECD data retrieval. This broadens the behavior from simple data access to credential handling and local state mutation, increasing the chance of secret leakage, unintended persistence, or misuse of stored credentials.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is presented as a data-query interface but also instructs the agent to manage and persist API keys in a local .env file and expose configuration interfaces unrelated to OECD data retrieval. This broadens the behavior from simple data access to credential handling and local state mutation, increasing the chance of secret leakage, unintended persistence, or misuse of stored credentials.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The docstring identifies this as a '小笨羊高考Skill' configuration, which contradicts the stated OECD data query purpose. Such provenance inconsistencies are dangerous in security review because they indicate copied or repurposed code that may carry hidden integrations, credentials, or behaviors unrelated to the declared functionality.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The configuration is clearly mismatched with the advertised OECD data query skill: it references a different service ('小笨羊高考'), hard-coded remote endpoint metadata, and separate API-key handling. In a skill that should only support OECD data access, unrelated external-service configuration is a strong supply-chain red flag because it can route traffic or secrets to an unexpected service.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config.py (reported line 13)May include surrounding context.

python
model_config = SettingsConfigDict(
        env_prefix="XBY_GAOKAO_",
        env_file=".env",
        env_file_encoding="utf-8",
        extra="ignore",
    )

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config.py (reported line 32)May include surrounding context.

python
model_config = SettingsConfigDict(
        env_prefix="XBY_GAOKAO_",
        env_file=".env",
        env_file_encoding="utf-8",
        extra="ignore",
    )

Credential Access

High
Category
Privilege Escalation
Confidence
78% confidence
Finding

The post-init logic forcibly reads a specific API key from a local .env file outside the normal scoped settings prefix, which is more invasive than standard configuration loading. In the context of a mismatched skill, this increases concern that the code is harvesting or reusing credentials for an unrelated service.

Content

Scanner excerpt · scripts/config.py (reported line 31)May include surrounding context.

python
default_year: int = 2025

    def model_post_init(self, __context):
        # 强制从 .env 文件读取 XBY_APIKEY
        env_path = Path(".env")
        if env_path.exists():
            content = env_path.read_text(encoding="utf-8")

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/config.py (reported line 39)May include surrounding context.

python
if line.startswith("XBY_APIKEY="):
                    self.api_key = line.split("=", 1)[1].strip()
                    break
        # 如果环境变量有值,覆盖 .env 的值
        env_val = os.getenv("XBY_APIKEY", "")
        if env_val:
            self.api_key = env_val

Credential Access

High
Category
Privilege Escalation
Confidence
90% confidence
Finding

The function explicitly saves API keys into a plaintext .env file, creating durable local secret storage. If the working directory is accessible to other users, included in backups, or accidentally committed, the credential can be exposed and reused against the remote service.

Content

Scanner excerpt · scripts/config.py (reported line 48)May include surrounding context.

python
def save_api_key_to_env(api_key: str) -> bool:
    """将API key保存到.env文件"""
    try:
        env_path = Path(".env")
        lines = []
        if env_path.exists():
            lines = env_path.read_text(encoding="utf-8").splitlines()

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill declares no explicit tool scope or permission boundaries even though it appears capable of reading environment variables, reading/writing local files, and making network calls. In an agent setting, this weakens least-privilege controls and makes it easier for a benign data-query skill to be repurposed for credential access, local persistence, or unintended external communication.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The skill’s operational instructions and requirements are presented in Chinese and include mandatory language such as '必须遵守' and '强制要求', but there is no indication that users may choose another language or that Chinese is required for a justified regional reason. This can constitute a language/locale policy issue because it effectively forces one language without opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly instructs the agent to ask for a user's API key and persist it via a local configuration mechanism, but provides no warning about sensitive credential handling, retention, or storage location. In an agent workflow, this can normalize unsafe secret collection and cause users to disclose credentials without informed consent or adequate protection.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation includes example calls and project-structure references to gaokao/school-query functionality that do not match the declared OECD dataset purpose. Such cross-domain inconsistencies are a supply-chain red flag because they suggest the skill may be repackaged from unrelated code, increasing the risk of hidden behaviors, wrong endpoints, or accidental invocation of unintended capabilities.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The implementation does not directly query OECD SDMX endpoints as the skill metadata suggests, but instead forwards requests to a generic third-party 'XiaoBenYang MCP API' using configurable headers and an API key. This creates a supply-chain/trust-boundary risk: user queries and parameters may be sent to an unrelated upstream service, enabling undisclosed data exfiltration, result manipulation, or policy bypass if the upstream is malicious or compromised.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

This code reads a credential via get_api_key() and sends it as the XBY-APIKEY header in an outbound network request. While the code includes internal logging for request outcomes, there is no confirmation prompt, user-facing warning, or explanatory comment/docstring in this file disclosing that credentials and request data will be transmitted to an upstream service.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code persists an API key into a local .env file even though the skill is described as a read-only OECD data query service. Secret persistence is behavior beyond user expectation for such a skill and increases exposure through accidental inclusion in logs, backups, source control, or other local file disclosure paths.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill implements local credential persistence capability that is not necessary for basic OECD data retrieval. Even if not overtly malicious, adding secret-writing behavior expands the attack surface and creates a path for sensitive tokens to remain on disk beyond the session.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code silently writes an API key to .env without any user-facing warning, confirmation, or disclosure in the persistence path. This is risky because users may believe they are supplying a transient credential while the skill is actually creating a durable local secret copy.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
94% confidence
Finding

The dependency is specified with a lower bound only, which allows future installs to resolve to different versions over time. This weakens build reproducibility and makes it harder to verify whether a deployed version includes known security fixes or reintroduces vulnerable releases through resolver behavior.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.31.0
pydantic>=2.7.0
pydantic-settings>=2.2.0
python-dotenv>=1.0.1

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

requests has multiple historical advisories, but the manifest does not pin the installed version, so it is impossible to determine from this file alone whether a vulnerable release will be deployed. In a network-facing OECD data query service, ambiguity around the HTTP client version increases risk because request handling and credential behaviors can be security-sensitive.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
94% confidence
Finding

Using pydantic>=2.7.0 without an upper bound or exact pin permits non-deterministic dependency resolution across environments. That increases supply-chain risk and makes security posture unverifiable because different installations may pull different versions with different vulnerability status.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
requests>=2.31.0
pydantic>=2.7.0
pydantic-settings>=2.2.0
python-dotenv>=1.0.1

Unverifiable Dependency: pydantic has 4 known advisory(ies) (CVE-2021-29510 (Use of "infinity" as an input to datetime and date fields causes infinite loop i); CVE-2024-3772 (Pydantic regular expression denial of service); CVE-2021-29510 (Pydantic is a data validation and settings management using Python type hinting.) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

pydantic has known advisories, and without an exact resolved version the deployment's exposure cannot be verified. Since this service likely validates API inputs and configuration through pydantic, uncertainty about parser and validation library security is a genuine though low-severity supply-chain concern.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

The pydantic-settings dependency is unpinned, so installs are not reproducible and may drift to versions with new bugs or advisories. For a service likely loading configuration and secrets, uncertainty around the exact package version is a meaningful supply-chain weakness.

Content

Scanner excerpt · requirements.txt (reported line 3)May include surrounding context.

text
requests>=2.31.0
pydantic>=2.7.0
pydantic-settings>=2.2.0
python-dotenv>=1.0.1

Unverifiable Dependency: pydantic-settings has 1 known advisory(ies) (CVE-2026-58203 (pydantic-settings: NestedSecretsSettingsSource follows symlinks outside secrets_)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
89% confidence
Finding

pydantic-settings has a known advisory and the absence of version pinning makes it impossible to verify whether the package version used in production is affected. This matters more in a service that likely loads local configuration and secrets, where package behavior can affect confidentiality or file-boundary assumptions.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.