Back to skill

Security audit

实验数据集成本估算(Experiment Dataset Cost Estimation)

Security checks for vulnerabilities and agentic risk

Overview

The skill appears aimed at legitimate dataset cost planning, but it needs review because its install and tool scope are under-declared and its dependencies run in a non-isolated, unpinned environment.

Review the dependency and tool-scope configuration before installing. Prefer running it in an isolated environment with pinned dependencies, require explicit authorization for PatSnap MCP and Office workflow calls, cap SAR polling, and treat compliance outputs as preliminary planning aids rather than legal approval.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
skill.manifest.json:9
Finding

Unbounded Dependencies Installed in a Non-Isolated Environment

Content
View full analysis

Vulnerability Details

File Location: skill.manifest.json, lines 9–16
Vulnerability Type: Supply-chain exposure caused by broad dependency constraints and lack of environment isolation
Risk Level: Medium

Vulnerable configuration:

json
"runtime": {
  "python_version": "3.12",
  "dependencies": [
    "openpyxl>=3.1.0",
    "pandas>=2.0.0",
    "python-docx>=1.1.0",
    "requests>=2.31.0"
  ],
  "isolated_env": false
}

Technical Analysis

Every dependency uses an open-ended minimum-version constraint. Consequently, any future release satisfying the constraint may be selected without having been reviewed with this Skill. The configuration also explicitly disables environment isolation.

Several declared packages are not used by the local runtime implementation, unnecessarily increasing the number of third-party components trusted during dependency installation and execution. Although no malicious dependency was identified in the reviewed package, this configuration increases exposure to future package compromise, malicious transitive dependencies, and incompatible releases.

Attack Path

  1. An upstream dependency or one of its transitive dependencies publishes a compromised future release that satisfies the open-ended version constraint.
  2. The Skill environment resolves and installs that release.
  3. Malicious installation or import-time behavior executes in the non-isolated Python environment.
  4. The dependency can access resources available to the host process, subject to the operating-system identity and sandbox restrictions under which installation or execution occurs.

Impact Assessment

Successful exploitation could execute code with the privileges of the process installing or running the Skill. In a shared environment, this may expose host-accessible files, environment variables, credentials, or other Python packages, and may modify the shared environment. ...[truncated 172 chars]

Remediation
View remediation

Remediation Suggestions

  1. Set isolated_env to true.
  2. Remove dependencies that the executable implementation does not use.
  3. Pin each required direct and transitive dependency to an audited version.
  4. Use a reproducible lockfile and require package hashes during installation.
  5. Retrieve packages only from an approved registry.
  6. Run dependency installation and Skill execution with a least-privileged account inside a restricted container or virtual environment.
  7. Add automated vulnerability and provenance scanning for dependency updates.

T09 · Insecure Skill Coding Practices

Note
Location
references/mcp_tool_routing.md:58
Finding

Unbounded Polling of External SAR Tasks

Content
View full analysis

Vulnerability Details

File Location: references/mcp_tool_routing.md, lines 58–66
Vulnerability Type: Unbounded retry loop for external MCP task status
Risk Level: Low

Vulnerable instructions:

python
# Step 3c: SAR structure-activity extraction
import time
for pn in selected_patents[:5]:
    task = ls_sar_submit(pn=pn, language="EN")
    while True:
        result = ls_sar_fetch(task_id=task["task_id"])
        if result["status"] in ["SUCCESS", "FAILED"]:
            break
        time.sleep(10)

Technical Analysis

The routing instructions tell the Agent to poll an external SAR task inside an unconditional while True loop. The loop has no maximum attempt count, absolute deadline, cancellation mechanism, exception handling, or treatment for unexpected terminal status values.

If the service continually returns a nonterminal or unrecognized status, the workflow can continue polling indefinitely. The ten-second delay reduces request frequency but does not bound total runtime or total requests.

Attack Path

  1. The Agent submits a SAR extraction task through ls_sar_submit.
  2. The service never returns SUCCESS or FAILED, whether because of a stalled task, malformed response, new status value, or service fault.
  3. The Agent invokes ls_sar_fetch every ten seconds without a retry limit.
  4. The workflow remains occupied and continues consuming Agent runtime and external service capacity until forcibly terminated.

Impact Assessment

Exploitation does not provide additional system privileges. Its primary impact is denial of service within the current workflow, repeated MCP/API consumption, delayed report delivery, and potentially avoidable service costs.

The affected scope is limited to executions that follow these documented MCP routing instructions; the local estimator script does not implement this polling loop.

Remediation
View remediation

Remediation Suggestions

  1. Set a maximum attempt count and an absolute elapsed-time deadline.
  2. Use capped exponential backoff with jitter.
  3. Recognize and handle all documented terminal, cancellation, expiration, and error states.
  4. Validate that task_id, status, and the response structure are present and correctly typed.
  5. Catch transient service exceptions separately from permanent failures.
  6. Stop polling when the user cancels the operation.
  7. Return a clear timeout result and preserve enough task metadata for safe manual retry.
  8. Apply service-side and client-side request budgets.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (22)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Yes, this is a mismatch. The declared description claims substantive budgeting and estimation functionality for experimental or training dataset construction, but the supplied code chunk contains only a trivial script that prints a readiness message. It does not implement the stated purpose, access relevant inputs, or perform any calculations related to resources, schedules, costs, procurement, or delivery.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
88% confidence
Finding

描述与代码部分匹配:代码确实在做数据集构建规模和成本/周期估算。但存在明显能力与范围偏差。首先,代码额外执行了数据源合规/许可风险评估,这是声明中未提及的独立能力。其次,代码的成本模型主要是按数据量、专利件数和靶点数推算人工工时,并生成相关清单与报告文件;它并未实现声明中强调的人力、样本、试剂、设备、算力、采购和交付成本的全面资源估算。其领域也更具体地偏向药物/靶点数据集,而非一般“实验和科研数据集”。因此属于描述与实际行为部分不一致,且实际能力范围与声明重点存在实质偏差。

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Line L147 states '本技能不直接执行数据下载', implying the skill does not directly fetch data. However, the documented workflow explicitly calls multiple search/fetch operations such as ls_target_fetch, ls_patent_structure_fetch, ls_sar_fetch, and literature searches, which are direct remote data retrieval actions, and also generates output files via an Office workflow.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The example output presents '智慧芽 MCP 检索结果摘要' as if it were an actual tool-produced result tied to a concrete session, even though this file is only a static sample. That can mislead users into believing the skill performed live external retrieval and validated the cited counts, which risks false trust in fabricated or stale data during budgeting, procurement, or research planning.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The statement '所有数据源均为 ✅ 绿色,无合规障碍' makes an absolute compliance claim without showing any decision logic, jurisdictional assumptions, license review, or source-by-source justification. In a costing and dataset-planning skill, this can cause users to underweight legal, licensing, privacy, or patent-use restrictions and proceed on an unsound compliance basis.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The title and all section headings in the template are written exclusively in Chinese, which establishes a fixed output language for generated reports. Under the policy rule, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is clearly justified, which is not present here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This markdown file presents all headings and operational guidance exclusively in Chinese, and there is no indication that users may choose another language or that the skill is intentionally restricted to a Chinese-speaking audience. Under the policy criteria, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The document presents all headings, instructions, and field descriptions only in Chinese. Under the policy rule, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly justified, which is not stated here.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documented routing significantly exceeds the stated purpose of dataset cost estimation by directing target normalization, drug intelligence retrieval, patent structure extraction, SAR processing, and ADMET screening. This scope mismatch can cause an agent to perform sensitive bio/pharma analysis work that users did not request, increasing the chance of unsafe capability exposure, policy bypass, and misuse under the cover of a benign budgeting skill.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

ADMET prediction is an active molecular triage capability, not a passive cost-estimation function. Embedding it inside a costing skill creates hidden high-value scientific functionality that could be invoked to prioritize compounds or support medicinal chemistry decisions, which is more dangerous in context because the skill description suggests a low-risk budgeting purpose.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script sets --output-language to zh by default and most descriptions, prompts, and status messages are hard-coded in Chinese. This can violate a language/locale policy when a specific language is imposed without explicit user choice, especially since the script is not documented as region-specific.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/test_estimator.py (reported line 107)May include surrounding context.

python
def test_cli_help():
    """cli: --help 输出"""
    result = subprocess.run(
        [sys.executable, _SCRIPT, "--help"],
        capture_output=True, text=True
    )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/test_estimator.py (reported line 119)May include surrounding context.

python
def test_cli_dual_target_academic():
    """cli: 双靶点学术场景端到端运行"""
    with tempfile.TemporaryDirectory() as tmpdir:
        result = subprocess.run(
            [sys.executable, _SCRIPT,
             "--targets", "PD-L1,4-1BB",
             "--task-type", "multi-target",

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/test_estimator.py (reported line 139)May include surrounding context.

python
def test_cli_single_target_commercial():
    """cli: 单靶点商业场景端到端运行"""
    with tempfile.TemporaryDirectory() as tmpdir:
        result = subprocess.run(
            [sys.executable, _SCRIPT,
             "--targets", "EGFR",
             "--task-type", "regression",

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The inputs table states that output_language defaults to zh while only optionally supporting en. This imposes a specific language by default rather than offering a neutral choice or explicit opt-in, which can violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The file presents the skill name and short description in Chinese while the default prompt is fixed in English. This creates a language/locale constraint without any explicit opt-in or documented language selection behavior, which matches the policy concern for forced language use.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file is entirely written in Chinese and does not mention any language preference, opt-in, or region-specific constraint. Under the policy for natural-language violations, forcing a specific language without user choice can be a locale-policy issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The entire document is written in Chinese, including the title and all operational guidance, with no indication that users may choose another language or that the document is intentionally limited to a Chinese-speaking audience. Under the policy, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

This markdown file contains user-facing instructional content only in Chinese, and there is no indication that users may choose another language or that the skill is intentionally limited to a Chinese-speaking context. Under the policy rule for language or locale constraints, forcing a single language without opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

This Python file contains natural-language instructions and test descriptions entirely in Chinese, including the top-level usage guidance. The policy allows locale-specific behavior only when the skill offers user choice or clearly documents a justified regional constraint, which is not present here.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.