Back to skill

Security audit

X2strategy

Security checks for vulnerabilities and agentic risk

Overview

This quant-research skill is coherent, but it should be reviewed because it stores API keys in project files, probes an unrelated secret path in tests, and runs generated trading code directly without a sandbox.

Install only if you are comfortable with external LLM calls, market-data network access, and local execution of generated Python. Use a throwaway workspace, a restricted API key with spending limits, avoid storing secrets in the project .env when possible, do not run real tests, and review generated strategy code before any backtest execution.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
references/spec2code.md:70
Finding

Generated Strategy Code Is Executed Without a Security Sandbox or Dangerous-Operation Validation

Content
View full analysis
/ uv run python strategy_1.py ``` If dependencies are absent, it creates an environment and installs packages before execution: ```bash cd library// uv venv && uv pip install backtrader yfinance akshare uv run python strategy_1.py ``` The validator performs syntax, structure, and Backtrader indicator checks, but it contains no policy against dangerous Python operations: ```python def validate_code(code: str) -> ValidationResult: errors: List[str] = [] warnings: List[str] = [] try: ast.parse(code) except SyntaxError as e: errors.append(f"SyntaxError at line {e.lineno}: {e.msg}") return ValidationResult(valid=False, errors=errors, warnings=warnings) tree = ast.parse(code) # Structural Backtrader checks omitted here. if _VALID_INDICATORS: invalid = _check_indicators(tree) errors.extend(invalid) return ValidationResult( valid=len(errors) == 0, errors=errors, warnings=warnings, ) ``` ### Tec ...[truncated 2473 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
tests/test_real_e2e.py:53
Finding

Real End-to-End Tests Read an API Key from an Unrelated Project’s Hard-Coded Secret Path

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/run_full_tests.sh:34
Finding

Full Test Runner Automatically Imports API Credentials from an Unrelated Absolute Path

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:80
Finding

Skill Instructs the Agent to Store API Keys in a Plaintext Project File Without Guaranteed Ignore or Permission Controls

Content
View full analysis
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
paper2spec/parser.py:196
Finding

Untrusted Document Content Is Interpolated into LLM Instructions Without Prompt-Injection Isolation

Content
View full analysis
100_000: ctx = full_text[:90_000] + "\n\n[...truncated...]\n\n" + full_text[-10_000:] else: ctx = full_text methodology, data_desc, signal_logic = await asyncio.gather( _extract_section_direct(ctx, METHODOLOGY_PROMPT, model=model), _extract_section_direct(ctx, DATA_DESCRIPTION_PROMPT, model=model), _extract_section_direct(ctx, SIGNAL_LOGIC_PROMPT, model=model), ) ``` ```python async def _extract_section_direct( ctx: str, prompt_template: str, *, model: Optional[str] = None ) -> str: prompt = prompt_template.format(context=ctx) return await achat(prompt, system=SYSTEM_PROMPT, model=model) ``` The system prompt in `paper2spec/prompts.py` does not establish an untrusted-data boundary: ```python SYSTEM_PROMPT = ( "You are an expert quantitative researcher and algorithmic trader. " "Extract precise, structured information from academic finance papers. " "Be rigorous — prefer exact values from the paper over guessed defaults." ) ``` The document content is embedded between ordinary prompt instructions: ```python METHODOLOGY_PROMPT = """Synthesize the trading strategy methodology from the provided text. Context from paper: {context} Instructions: 1. Describe the core trading idea ... 2. Explain the step-by-step process of the strategy ... ... """ ``` ### Technical Analysis PDF, Markdown, DOCX, and text inputs are attacker-controllable. Their contents are interpolated directly into the user message sent to the LLM. Neither the system prompt nor the prompt template explicitly tells the model: - That the document is ...[truncated 2058 chars]
Remediation
View remediation
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (110)

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The supplied code chunk is narrowly focused on parsing documents into a PaperContent object and using LLM prompts to extract a few semantic sections. This partially aligns with the declared 'paper2spec' idea, but the declaration describes a much larger end-to-end system with two core capabilities, including spec2code, validation, backtesting, and search. None of those downstream capabilities appear in this code. The code supports multiple document formats as declared, but its actual primary purpose is document text extraction plus section summarization, not full research-to-executable-results orchestration. Therefore the description materially overstates what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The supplied script matches part of the declared 'paper2spec' capability: it parses documents and extracts strategy specs, including multilayer extraction and output generation. However, the declared description emphasizes two core capabilities, including 'spec2code' that generates validated Backtrader code, runs backtests, and compares against paper metrics, plus search-oriented triggers. None of those behaviors appear in this code. The script's actual primary purpose is a one-shot document-to-content/spec analyzer and exporter, not the full end-to-end research-to-executable-results pipeline. Therefore the description overstates the implemented behavior for this code chunk in a materially important way.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The supplied code chunk is a narrow developer tooling script for schema generation, not an implementation of the advertised end-to-end research-to-strategy pipeline. It reads local Python model definitions and writes JSON schema files. While these schemas may support the larger system, this chunk does not itself analyze finance papers, extract strategies, generate executable trading code, run backtests, or search for papers. Therefore the code's actual behavior is materially different from the declared primary purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description presents a comprehensive research-to-implementation pipeline with two major capabilities: paper2spec and spec2code, covering multiple document types, code generation, backtesting, diagnosis, and search. The supplied code chunk only implements a parsing front end for a PDF-to-structured-JSON extraction step. That is related to one sub-part of the declared paper analysis flow, but it does not implement the broader claimed functionality and supports only PDF input. Because the code’s actual behavior is a narrow extractor rather than the advertised end-to-end strategy-to-backtest system, this is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The code chunk does not implement the user-facing quant research capabilities described. Instead, it is an operational script for validating the project via pytest across parser, extractor, render, pipeline, library, and quality tests. While the test names suggest the broader system may support the declared functionality, this specific code's primary purpose is test orchestration and reporting, not extracting strategies from documents, generating executable Backtrader code, searching papers, or running user-requested backtests. Therefore the description does not accurately represent this supplied code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a comprehensive pipeline with two core capabilities: paper2spec and spec2code, including document parsing, strategy extraction, code generation, backtesting, and diagnostic comparison. The supplied code chunk only implements a search utility for finding quant finance papers and outputting search results. While paper search is mentioned in the declaration, this code covers only that small sub-capability and not the primary advertised behavior. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents an end-to-end research-to-implementation system with document ingestion, paper analysis, code generation, and backtesting. The supplied code chunk only validates an existing strategy Python file without executing it. While validation could be a supporting component of a larger spec2code pipeline, this chunk’s primary behavior is much narrower and does not implement the headline capabilities described. Therefore the description does not accurately represent what this specific code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The supplied code chunk is unrelated to the declared end-to-end quant research and Backtrader workflow. It only configures pytest behavior for test execution, specifically gating 'real' tests behind a --run-real flag. There is no implementation of paper parsing, LLM extraction, strategy specification, Backtrader code generation, backtest execution, paper search, or metric comparison. This is a materially different primary purpose, so the description does not accurately represent the code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The description promises an end-to-end system from research input to structured strategy spec to executable Backtrader code to backtest and diagnosis report. However, this supplied code chunk only contains tests for the earlier parts of that pipeline: search, PDF extraction, parsing, multilayer extraction, markdown rendering, and example-quality validation. It does support parts of the declared 'paper2spec' functionality, including document parsing, multi-strategy extraction, and paper search. But it does not implement or exercise the declared 'spec2code' capability: there is no Backtrader code generation, no strategy execution engine, no backtest run, and no comparison against paper metrics beyond checking extracted expected performance fields. Since the declared description emphasizes two core capabilities and one of them is absent from the actual code chunk, and the chunk's primary nature is testing rather than executable skill behavior, this is a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description promises an end-to-end quantitative research workflow: extracting strategy specs from papers and other inputs, generating executable Backtrader code, running backtests, and comparing results. The supplied code does none of that. It is strictly a test file for example JSON artifacts. It loads local example files, parses them into model objects, validates titles/methodology, checks JSON round-tripping, and ensures extracted strategies have expected counts, names, and indicators. While these tests are loosely related to the paper2spec domain, they do not implement the advertised core capabilities and instead serve a materially different purpose: fixture/schema validation. Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a broad quant-research automation skill with two major capabilities: paper2spec extraction and spec2code/backtesting. The supplied code chunk only covers tests for the parser portion of paper2spec. It checks extraction helpers, prompt/truncation behavior, semantic retrieval, and PDF parsing pathways; it does not implement or test strategy code generation, Backtrader execution, backtest comparison, diagnosis reporting, or search. While the parser-related portion is consistent with part of the declaration, the overall declared purpose materially overstates what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The description promises a full research-to-implementation pipeline with two core capabilities: paper2spec extraction and spec2code/backtesting/diagnosis. The supplied code only evidences the first half plus validation infrastructure: arXiv search, PDF text extraction, LLM parsing into PaperContent, extraction into StrategySpec, markdown rendering, file output, and regression/quality tests. There is no visible functionality for generating Backtrader strategy code, running Backtrader backtests, or comparing backtest outputs to paper-reported metrics, which are central parts of the declared purpose. In addition, the code is specifically a pytest integration test module rather than the main skill implementation, so its primary purpose is validation rather than user-facing execution. The undeclared resource usage is also notable: real internet access, API key loading from a local .env path, reading local sample PDFs, and writing temp files. While some of that may support the overall domain, the major functional gap around code generation/backtesting makes this a material description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a broad research-to-implementation pipeline with extraction, code generation, and backtesting. The supplied code does none of that directly; it is narrowly focused on unit tests for rendering model objects into Markdown. While rendering extracted content/specs could be a supporting detail within a larger system, this chunk’s actual purpose is testing presentation formatting rather than implementing the core advertised capabilities. Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents a full end-to-end quant research workflow skill with document analysis, strategy extraction, code generation, backtesting, and diagnosis. The supplied code, however, is only a test module for model serialization and dataclass integrity. It defines fixtures and assertions around defaults, to_dict/from_dict, and JSON round-trips for model classes. While these model names align with the broader declared domain, this chunk does not itself perform the advertised core capabilities. Therefore, the actual behavior of the supplied code chunk is materially narrower and different from the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents a comprehensive research-to-backtest system with two major capabilities: extracting strategy specs from documents and generating/running Backtrader implementations. The supplied code does not implement those user-facing capabilities. Instead, it is only a test module for validating generated Backtrader code structure and indicator names. While such validation could be a supporting subcomponent of the declared spec2code system, this chunk by itself has a much narrower purpose and does not evidence the declared ingestion, extraction, generation, search, or backtesting behavior. Therefore the description materially overstates what this code chunk actually does.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
93% confidence
Finding

Referencing .env examples and configuration is not dangerous by itself, but in this skill context it is part of a workflow that actively collects and writes credentials. Because the skill also implies environment access and file operations, encouraging secret placement in .env raises the risk of credential disclosure, reuse by unrelated tasks, and accidental exfiltration through tooling or repository operations.

Content

Scanner excerpt · SKILL.md (reported line 316)May include surrounding context.

md
Any [litellm-supported model](https://docs.litellm.ai/docs/providers) works.
The `--model` flag on any script overrides `PAPER2SPEC_MODEL`.
Full config + .env examples: [references/skill-internals.md](references/skill-internals.md)

---

Credential Access

High
Category
Privilege Escalation
Confidence
92% confidence
Finding

This reference again reinforces .env-based secret handling in the skill's supporting documentation. In a code-generating, file-writing, network-enabled agent skill, centralizing credentials in .env without tighter controls materially increases the blast radius if the workspace is exposed or if the agent later reads and transmits environment-derived secrets.

Content

Scanner excerpt · SKILL.md (reported line 326)May include surrounding context.

md
- [references/paper2spec.md](references/paper2spec.md) — Parser modes, multi-strategy detection, output schemas
- [references/spec2code.md](references/spec2code.md) — Code generation workflow, Backtrader patterns
- [references/skill-internals.md](references/skill-internals.md) — Script flags, output formats, .env examples, library management, project structure
- [references/backtrader_patterns.md](references/backtrader_patterns.md) — Strategy class, data loading, position sizing
- [references/indicator_cookbook.md](references/indicator_cookbook.md) — Built-in and custom indicators
- [references/data_sources.md](references/data_sources.md) — yfinance, akshare, FRED API

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · paper2spec/config.py (reported line 14)May include surrounding context.

python
PROJECT_ROOT = Path(__file__).resolve().parent.parent
ENV_PATH = PROJECT_ROOT / ".env"


def load_project_env() -> None:

Credential Access

High
Category
Privilege Escalation
Confidence
87% confidence
Finding

The file gives explicit guidance for storing provider API keys in a local .env file and includes commands for runtime verification. In an agent skill context, this is sensitive because agent-accessible internal references that normalize handling raw secrets can lead to credential exposure through logs, file reads, accidental output to users, or reuse across tasks if workspace boundaries are weak.

Content

Scanner excerpt · references/skill-internals.md (reported line 6)May include surrounding context.

Agent-only reference. Contents here are for the agent's internal use — never present raw commands or file structures to users unless explicitly asked.

.env Format by Provider

DeepSeek (recommended)

text

Credential Access

High
Category
Privilege Escalation
Confidence
98% confidence
Finding

The hard-coded reference to /home/whlu/ALAGENT/deepagents-quickstarts/aegra/.env causes this script to reach into another workspace for secrets. In a skill that can generate and run code and invoke networked LLM-backed tests, automatically importing those credentials expands blast radius and can enable unintended credential use or disclosure.

Content

Scanner excerpt · scripts/run_full_tests.sh (reported line 35)May include surrounding context.

sh
cd "$PROJECT_ROOT"

# -- Load API keys --
AEGRA_ENV="/home/whlu/ALAGENT/deepagents-quickstarts/aegra/.env"
if [[ -f "$AEGRA_ENV" ]]; then
    echo "[env] Loading API keys from aegra/.env"
    set -a

Credential Access

High
Category
Privilege Escalation
Confidence
97% confidence
Finding

At this point the script conditionally loads and exports API keys from the external .env file, making those secrets available to all child processes in the test run. That behavior increases the chance of exfiltration through test code, third-party libraries, accidental debug output, or outbound API calls during E2E tests.

Content

Scanner excerpt · scripts/run_full_tests.sh (reported line 37)May include surrounding context.

sh
# -- Load API keys --
AEGRA_ENV="/home/whlu/ALAGENT/deepagents-quickstarts/aegra/.env"
if [[ -f "$AEGRA_ENV" ]]; then
    echo "[env] Loading API keys from aegra/.env"
    set -a
    # Only export DEEPSEEK_API_KEY and OPENROUTER_API_KEY
    while IFS='=' read -r key value; do

Credential Access

High
Category
Privilege Escalation
Confidence
93% confidence
Finding

The surrounding logic confirms fallback behavior tied to the external .env loading mechanism, reinforcing that credential acquisition is built into the test harness rather than explicitly provided by the operator. While this line itself is not the read operation, it is part of a true credential-access pattern that normalizes pulling secrets from outside project scope.

Content

Scanner excerpt · scripts/run_full_tests.sh (reported line 50)May include surrounding context.

sh
done < <(grep -E '^(DEEPSEEK|OPENROUTER)' "$AEGRA_ENV" | sed 's/#.*//' | grep '=')
    set +a
else
    echo "[env] WARNING: aegra/.env not found, relying on existing env vars"
fi

# -- Activate venv --

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 147)May include surrounding context.

md
"""Shared configuration for spec2code.

Reuses the same .env / library path as paper2spec so both halves of
the pipeline share a single configuration surface.
"""

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 323)May include surrounding context.

md
"""Shared configuration for spec2code.

Reuses the same .env / library path as paper2spec so both halves of
the pipeline share a single configuration surface.
"""

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README_CN.md (reported line 147)May include surrounding context.

md
"""Shared configuration for spec2code.

Reuses the same .env / library path as paper2spec so both halves of
the pipeline share a single configuration surface.
"""

Static analysis

No suspicious patterns detected.