Back to skill

Security audit

tushare-finance

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches Tushare financial-data use, but it also bundles credentialed web crawling, CAPTCHA OCR, broad local writes, and some broader-than-advertised datasets that deserve review before installation.

Install only if you are comfortable sending queries and a Tushare token to Tushare. Do not run scripts/crawl_docs.py or enable the documented CI sync unless you intentionally want a headless browser to use Tushare account/password secrets and crawl/update local reference files. Use an isolated environment, pin dependencies, and choose export paths carefully.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:26
Finding

Unpinned Third-Party Dependencies Permit Unreviewed Package Releases

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:26-31; requirements.txt:1-2
Vulnerability Type: Supply-chain exposure through mutable dependency resolution
Risk Level: Medium

Vulnerable Code

SKILL.md:26-31:

markdown
If an error occurs, install the dependencies:
```bash
pip install tushare pandas
text

`requirements.txt:1-2`:

```text
tushare>=1.2.60
pandas>=1.5.0

Technical Analysis

The Skill instructs the agent to install packages by name without exact versions or cryptographic integrity hashes. The requirements file also specifies only minimum versions using >=. Consequently, dependency resolution may select any later release available from the configured Python package index.

Python packages can execute code during installation, and their imported modules execute with the privileges of the Python process. A future compromised release of tushare, pandas, or one of their transitive dependencies could therefore execute code before the Skill performs its intended financial-data operations.

This is a supply-chain hardening issue rather than evidence that the currently named packages are malicious. The exploit requires compromise of a permitted package release, a transitive dependency, or the package index/resolution path.

Attack Path

  1. An attacker compromises a future release of a declared package, one of its transitive dependencies, or the package repository used by pip.
  2. The target environment does not already contain the required modules, causing the Skill's dependency check to fail.
  3. Following SKILL.md, the agent runs:
    bash
    pip install tushare pandas
    
    Alternatively, an operator installs from requirements.txt.
  4. Because no exact versions or hashes are enforced, pip resolves the attacker-controlled compatible release.
  5. Malicious installation logic or imported module initialization executes with the permissions of th ...[truncated 928 chars]
Remediation
View remediation

Remediation Suggestions

  1. Replace lower-bound constraints with exact, reviewed versions:
    text
    tushare==REVIEWED_VERSION
    pandas==REVIEWED_VERSION
    
  2. Generate a lock file that pins all transitive dependencies, not only direct dependencies.
  3. Require hashes during installation, for example with a hash-locked requirements file and:
    bash
    pip install --require-hashes -r requirements.lock
    
  4. Update SKILL.md so the agent installs only from the reviewed lock file rather than installing packages by unqualified name.
  5. Require explicit user approval before modifying the Python environment.
  6. Use an isolated virtual environment with minimal filesystem and network permissions.
  7. Configure an approved package index and disable unexpected fallback indexes to reduce dependency-confusion exposure.
  8. Add automated dependency review, vulnerability scanning, and controlled lock-file update procedures.
  9. Reconcile the inconsistent dependency declarations in requirements.txt and metadata.json so every installation path resolves the same reviewed versions.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (96)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The description overstates the implemented coverage. This code chunk is clearly a Tushare API wrapper, so its core domain matches financial data access, but the concrete behavior shown is much narrower than the declared purpose: it only exposes several stock/financial/index endpoints and lacks any demonstrated macroeconomic, fund, futures, bond, Hong Kong stock, or U.S. stock retrieval methods. Additionally, it includes a data export capability to CSV/JSON/Excel files, which is an undeclared capability. Therefore the declared description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description says the skill retrieves financial and macroeconomic market data from Tushare Pro APIs for end-user data queries. However, this code does not call Tushare data APIs or return market datasets. Its primary function is to crawl and maintain Tushare API documentation locally. It also performs undeclared capabilities such as browser automation, credential-based login, CAPTCHA OCR, filesystem synchronization, and CI/CD output generation. These are materially different from the declared purpose, so this is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill uses 'env' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The script launches Playwright and crawls external URLs on tushare.pro, which is a real network capability not covered by the declared permissions. In an agent context, undeclared outbound network access increases risk because the skill can fetch remote content, interact with login flows, and change behavior based on external state without transparent consent.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The script launches Playwright and crawls external URLs on tushare.pro, which is a real network capability not covered by the declared permissions. In an agent context, undeclared outbound network access increases risk because the skill can fetch remote content, interact with login flows, and change behavior based on external state without transparent consent.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README documents functionality beyond a normal finance-data skill: a credentialed workflow that logs into an external service, solves captchas, crawls large numbers of pages, and auto-generates pull requests. This expands the trust boundary from simple API consumption to automated account use and web automation, creating security and compliance risk if users enable it without understanding the consequences.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README instructs users to configure account credentials for automated login and scraping, but provides no warning about how those credentials are handled, what the automation does with the account, or possible privacy/account impact. That omission can lead users to expose sensitive secrets to CI workflows or enable automation that may trigger monitoring, suspension, or unintended data access.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Automated login, captcha solving, and mass crawling are not necessary for the declared user purpose of retrieving financial market data, so they represent excess capability. Unnecessary privileged automation increases the chance of credential misuse, account lockouts, terms-of-service violations, and abuse of CI secrets if this workflow is adopted casually.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This markdown file describes output fields including legal representative, general manager, board secretary, company email, and other identifying business-contact information, but provides no warning about handling or redistributing that data. Under the markdown-specific missing-warning rule, documentation that omits privacy-impact warnings for data-affecting behavior should be flagged.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documented film_record API exposes nationwide movie script filing data, which is unrelated to the skill’s declared finance, market, and macroeconomic scope. This scope expansion increases the chance that an agent may access or return non-financial government/cultural data unexpectedly, violating least-privilege and creating data-governance and trust-boundary issues.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document directs users to obtain data through off-platform delivery and to contact a specific WeChat account, but provides no warning about privacy, impersonation, secure transfer, or approved handling procedures. In an agent skill context, this can normalize bypassing platform controls and expose users to data leakage, social engineering, or unvetted data exchange outside audited channels.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

文档前文明确说明该页面描述的是“cb_factor_pro”可转债技术面因子接口,但示例代码在“获取交易日期为20250724当天所有可转债的因子数据”这一注释下实际调用的是“pro.hk_income”,这是港股利润表类接口,和注释及文档主题明显不一致。该示例会误导使用者调用与声明目的不同的接口。

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documented npr interface exposes a broad national policy and regulation text repository, which is outside the skill’s declared scope of finance-market and macroeconomic data. This scope expansion increases the chance the agent will retrieve and rely on unrestricted long-form policy text, enabling prompt-injection-style content propagation, unexpected data handling, or use of the skill for non-financial tasks not covered by the original trust boundary.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest describes the skill as serving Chinese financial market data such as stocks, funds, futures, bonds, financial statements, and macroeconomic indicators. This file documents a bo_weekly endpoint for weekly movie box office rankings and ticket metrics, which is entertainment-industry data rather than financial market or macroeconomic data as described.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This file contains natural-language descriptions that force a specific language for users, beginning with the module docstring and continuing throughout the interface text. Under the policy, language-specific behavior should either offer user opt-in/choice or be clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The client includes generic file export functionality that can write CSV, JSON, or Excel output to an arbitrary local path. While this is likely intended as a convenience feature, unrestricted write capability expands the skill beyond read-only data retrieval and could overwrite local files or persist sensitive data if an agent passes attacker-influenced paths.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Local file writing is not clearly necessary for a skill described primarily as fetching financial market data, so it increases the attack surface without strong contextual justification. In an agent setting, arbitrary path writes can be abused for unintended persistence, overwriting application files, or planting data in sensitive locations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code file contains natural-language documentation, CLI descriptions, and runtime messages exclusively in Chinese. That enforces a specific language/locale for users without offering an alternative or documenting a justified regional constraint.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script automates credentialed login and attempts CAPTCHA solving using OCR, going beyond simple documentation retrieval into access automation. In the context of a finance-data skill, this is more sensitive because it handles account credentials and bypass-like mechanisms for gated content, creating compliance, account-abuse, and secret-handling risks if invoked in shared or automated environments.

Content

No source excerpt is available for this finding.

Tainted flow: 'gh_output' from os.environ.get (line 816, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/crawl_docs.py (reported line 818)May include surrounding context.

python
# ---- GitHub Actions 输出 ----
        gh_output = os.environ.get("GITHUB_OUTPUT")
        if gh_output:
            with open(gh_output, "a", encoding="utf-8") as f:
                f.write(f"has_changes={'true' if has_changes else 'false'}\n")
                f.write(f"new_count={len(stats['new'])}\n")
                f.write(f"updated_count={len(stats['updated'])}\n")

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This markdown file includes examples that write data to data.csv and data.xlsx, which can affect user files or create artifacts on disk. The surrounding documentation provides no warning or note that these operations write local files or may overwrite existing files.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The natural-language content of the skill documentation is presented in a single forced language, with no indication that users can choose another language or that the skill is intentionally limited to a Chinese-only audience for compliance or regional reasons. This matches the policy category for language or locale constraints without user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill instructs users to configure a Tushare token and make API calls to an external third-party service, but it does not clearly warn that user queries and credentials are sent to Tushare. This can create privacy and consent issues, especially when requests may reveal proprietary watchlists, research interests, or account-linked API usage.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The user-facing metadata fields description and tags are entirely in Chinese, which can impose a language expectation on users without offering an alternative or explicit opt-in. Under the policy, locale or language constraints should be optional or clearly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The entire skill reference is presented only in Chinese, including headings and descriptions, with no indication that users may choose another language or that the skill is intentionally restricted to a Chinese-only audience. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.