Back to skill

Security audit

Theta Trading System

Security checks for vulnerabilities and agentic risk

Overview

This stock-picking skill needs review because it makes very strong performance claims that do not match the packaged code and it writes data and model files to fixed root workspace paths.

Install only if you treat it as experimental research tooling, not reliable trading advice. Run it in an isolated environment, avoid running as root, review or change the hard-coded storage paths before execution, and pin dependencies from a reviewed lockfile. Do not rely on the advertised 100% accuracy or position recommendations for real financial decisions.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:46
Finding

Unpinned Third-Party Dependencies Permit Supply-Chain Substitution

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/theta_simplified_scorer.py:111
Finding

SQL Injection Through Direct Interpolation of the Date Parameter

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/daily_data_update.py:12
Finding

Hard-Coded Root Workspace Paths Cause Cross-Project File Access and Overwrites

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (21)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

If the implementation uses RandomForest or other models instead of the declared Ridge model, then the documentation is materially false about core decision logic. In regulated or high-risk domains like finance, such inconsistency undermines auditability, reproducibility, and informed consent by users, increasing the chance of misuse and losses.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

If the implementation uses RandomForest or other models instead of the declared Ridge model, then the documentation is materially false about core decision logic. In regulated or high-risk domains like finance, such inconsistency undermines auditability, reproducibility, and informed consent by users, increasing the chance of misuse and losses.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the implementation uses RandomForest or other models instead of the declared Ridge model, then the documentation is materially false about core decision logic. In regulated or high-risk domains like finance, such inconsistency undermines auditability, reproducibility, and informed consent by users, increasing the chance of misuse and losses.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the implementation uses RandomForest or other models instead of the declared Ridge model, then the documentation is materially false about core decision logic. In regulated or high-risk domains like finance, such inconsistency undermines auditability, reproducibility, and informed consent by users, increasing the chance of misuse and losses.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

If the implementation uses RandomForest or other models instead of the declared Ridge model, then the documentation is materially false about core decision logic. In regulated or high-risk domains like finance, such inconsistency undermines auditability, reproducibility, and informed consent by users, increasing the chance of misuse and losses.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This markdown file contains end-user instructions exclusively in Chinese, including installation, publishing, and usage guidance. Under the policy rule for language/locale constraints, forcing a single language without user opt-in or justification is a natural-language policy issue.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README presents materially inconsistent claims about the system’s core model and feature set: elsewhere it advertises a v1.2.0 Ridge model with 100% accuracy and reduced features, while this section still describes GradientBoosting with 14 features. In a trading skill, such contradictions can mislead users about how predictions are generated and how much trust to place in reported performance, increasing the risk of unsafe financial decisions based on inaccurate documentation.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The primary skill description is written entirely in Chinese, and the rest of the document consistently presents the skill in that language without offering users a language or locale option. Under the policy provided, forcing a specific language without user opt-in is a natural-language policy concern unless the locale constraint is explicitly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The description at L04 is written entirely in Chinese and does not indicate any language option or user opt-in. Under the policy criteria, forcing a specific language or locale without choice is a natural-language policy violation unless it is clearly documented as region-specific and justified.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file’s implementation materially conflicts with the advertised capability: instead of a '100% accurate Ridge model' with hourly evolution and multi-model validation, it performs a simple heuristic ranking over a few SQLite columns. In a trading skill, this mismatch can mislead users into placing financial trust in output that lacks the claimed modeling rigor, creating a real integrity and consumer-deception risk even without code execution abuse.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code file contains user-facing natural language in Chinese in the module docstring and throughout the printed output, but there is no indication that the language is configurable or opt-in. That can violate language/locale policy requirements when skills must not force a specific language on users by default.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script reports a simple directional hit-rate as 'accuracy' while the skill metadata claims a '100% accurate Ridge model,' yet the code actually trains RandomForest regressors and computes non-classification metrics. In a trading skill, misleading performance claims can cause operators to place undue trust in model outputs, leading to financial harm through risky or automated decisions based on overstated efficacy.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest description is written entirely in Chinese and targets A-share trading context without offering any language or locale choice. Under the policy, forcing a specific language or locale without user opt-in can be a natural-language policy violation unless the regional constraint is clearly documented and justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill documentation begins entirely in Chinese and continues in that language throughout the file, with no indication that users may choose another language or that the skill is intended only for a Chinese-speaking or region-specific audience. Under the policy rule for language/locale constraints, this is a natural-language policy issue because the documentation effectively enforces a specific language without opt-in.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation prominently claims a "Ridge回归模型" with 100% accuracy, but the later configuration example specifies model_type: "GradientBoosting". This is an active contradiction in the skill's own documentation about what model the system actually uses, which can mislead users about the implemented approach.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script's user-facing comments and echoed instructions are entirely in Chinese, including status messages and publishing steps. For a general-purpose publish script, this imposes a specific language on users without opt-in or a documented justification that the skill is region- or language-restricted.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring is entirely in Chinese and presents the skill's purpose and behavior only in that language. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The displayed date can be wrong because the print path uses a hard-coded fallback ('2026-03-21') instead of the actual resolved scoring date when no date is supplied. In a trading context, mislabeling the evaluation date can cause users to act on stale or mismatched recommendations, undermining auditability and trust in the output.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The file header and console text describe an 'optimized Theta model,' but the implementation uses generic RandomForest regressors rather than a clearly defined Theta or Ridge-based method. In security terms this is an integrity and trust issue: mismatched descriptions can mislead users, reviewers, or downstream automation about what is actually being executed and validated.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code file contains natural-language instructions and status output that force a specific language/locale for users. Under the policy, language constraints should either offer user opt-in/choice or be clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.