Back to skill

Security audit

Kalshalyst

Security checks for vulnerabilities and agentic risk

Overview

The skill is not clearly malicious, but it needs Review because it includes unattended real-money trading and instructions for agents to persistently modify installed code.

Install only after reviewing the auto-trading path. Keep auto_trader_config.json disabled or use dry-run unless you intentionally want live Kalshi orders, restrict credentials to the minimum needed, avoid unattended schedules until you trust the behavior, and remove or ignore the Agent Bug-Fix Protocol instructions that ask agents to modify installed skill files.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:715
Finding

Skill instructions compel unauthorized persistent source modifications

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/claude_estimator.py:285
Finding

Untrusted market text can influence autonomous real-money trading decisions

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:680
Finding

Documentation recommends bulk installation of unpinned ecosystem packages

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Rogue AgentSelf-Modification, Session Persistence
Findings (53)

Tainted flow: 'req' from os.environ.get (line 57, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
95% confidence
Finding

The Slack webhook URL is taken from an environment variable or local YAML config and used directly as a destination for an outbound HTTP request. That creates a tainted network sink: if an attacker can influence the environment or config, they can redirect notifications containing trade activity and market metadata to an attacker-controlled endpoint, causing data exfiltration and covert signaling.

Content

Scanner excerpt · scripts/auto_trader.py (reported line 58)May include surrounding context.

python
try:
        data = json.dumps({"text": message}).encode("utf-8")
        req = urllib.request.Request(webhook_url, data=data, headers={"Content-Type": "application/json"})
        urllib.request.urlopen(req, timeout=10)
    except Exception:
        pass

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

According to this finding, the skill does not implement the advertised scanning, estimation, or alerting pipeline at all and instead records trades and reconciles local execution data. In a trading environment, that mismatch is dangerous because it can disguise a bookkeeping or execution-adjacent component as a harmless research tool, leading to overbroad deployment and access.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

According to this finding, the skill does not implement the advertised scanning, estimation, or alerting pipeline at all and instead records trades and reconciles local execution data. In a trading environment, that mismatch is dangerous because it can disguise a bookkeeping or execution-adjacent component as a harmless research tool, leading to overbroad deployment and access.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

According to this finding, the skill does not implement the advertised scanning, estimation, or alerting pipeline at all and instead records trades and reconciles local execution data. In a trading environment, that mismatch is dangerous because it can disguise a bookkeeping or execution-adjacent component as a harmless research tool, leading to overbroad deployment and access.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

According to this finding, the skill does not implement the advertised scanning, estimation, or alerting pipeline at all and instead records trades and reconciles local execution data. In a trading environment, that mismatch is dangerous because it can disguise a bookkeeping or execution-adjacent component as a harmless research tool, leading to overbroad deployment and access.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

According to this finding, the skill does not implement the advertised scanning, estimation, or alerting pipeline at all and instead records trades and reconciles local execution data. In a trading environment, that mismatch is dangerous because it can disguise a bookkeeping or execution-adjacent component as a harmless research tool, leading to overbroad deployment and access.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

According to this finding, the skill does not implement the advertised scanning, estimation, or alerting pipeline at all and instead records trades and reconciles local execution data. In a trading environment, that mismatch is dangerous because it can disguise a bookkeeping or execution-adjacent component as a harmless research tool, leading to overbroad deployment and access.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The metadata describes a scanner/alerting skill, but this file is an autonomous trading orchestrator that places and cancels live orders. That is a material capability expansion: a user expecting analysis-only behavior could unknowingly enable real-money market actions, greatly increasing operational and financial risk.

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/auto_trader.py (reported line 159)May include surrounding context.

python
rules[key] = val
                if rules:
                    logger.info(f"Exit rules loaded from {path} ({len(rules)} params)")
                    return rules
            except Exception as e:
                logger.warning(f"Failed to load exit rules from {path}: {e}")
    logger.info("No exit rules found — position lifecycle management disabled")

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/claude_estimator.py (reported line 68)May include surrounding context.

python
# ── Claude API Interface (Fallback — requires ANTHROPIC_API_KEY) ──────────

def _load_anthropic_key() -> Optional[str]:
    """Load ANTHROPIC_API_KEY from env or ~/.openclaw/.env file."""
    key = os.environ.get("ANTHROPIC_API_KEY")
    if key:
        return key

Credential Access

High
Category
Privilege Escalation
Confidence
93% confidence
Finding

Opening ~/.openclaw/.env gives the skill filesystem access to a local secret store and broadens its capabilities beyond simple market analysis. While it targets a specific key, this pattern weakens least-privilege expectations for agent skills and can become dangerous if combined with prompt/control-path manipulation or reused in a broader plugin environment.

Content

Scanner excerpt · scripts/claude_estimator.py (reported line 72)May include surrounding context.

python
key = os.environ.get("ANTHROPIC_API_KEY")
    if key:
        return key
    # Try loading from .env file
    env_path = Path.home() / ".openclaw" / ".env"
    if env_path.exists():
        try:

Credential Access

High
Category
Privilege Escalation
Confidence
93% confidence
Finding

Reading and parsing the contents of ~/.openclaw/.env is a direct local secret-access capability. In the context of an agent skill, this exceeds the minimum needed behavior and could expose credentials if the skill is modified, compromised, or if logs/errors later reveal sensitive content.

Content

Scanner excerpt · scripts/claude_estimator.py (reported line 73)May include surrounding context.

python
if key:
        return key
    # Try loading from .env file
    env_path = Path.home() / ".openclaw" / ".env"
    if env_path.exists():
        try:
            for line in env_path.read_text().splitlines():

Credential Access

High
Category
Privilege Escalation
Confidence
91% confidence
Finding

The code not only reads the local .env file but promotes the extracted API key into os.environ, increasing the number of components in-process that can access the secret. That makes lateral exposure easier for imported libraries, subprocesses, or future code paths that inherit environment variables.

Content

Scanner excerpt · scripts/claude_estimator.py (reported line 82)May include surrounding context.

python
key = line.split("=", 1)[1].strip().strip('"').strip("'")
                    if key:
                        os.environ["ANTHROPIC_API_KEY"] = key
                        logger.info("Loaded ANTHROPIC_API_KEY from ~/.openclaw/.env")
                        return key
                elif line.startswith("ANTHROPIC_API_KEY="):
                    key = line.split("=", 1)[1].strip().strip('"').strip("'")

Credential Access

High
Category
Privilege Escalation
Confidence
91% confidence
Finding

This duplicates the same risky behavior for an alternate .env syntax: extracting a local secret and placing it into the global environment. The issue is not the parsing branch itself but the expanded secret exposure model created by auto-loading and re-exporting credentials.

Content

Scanner excerpt · scripts/claude_estimator.py (reported line 88)May include surrounding context.

python
key = line.split("=", 1)[1].strip().strip('"').strip("'")
                    if key:
                        os.environ["ANTHROPIC_API_KEY"] = key
                        logger.info("Loaded ANTHROPIC_API_KEY from ~/.openclaw/.env")
                        return key
        except Exception as e:
            logger.warning(f"Failed to read .env: {e}")

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
94% confidence
Finding

Returning prompt text loaded from ~/prompt-lab/prompt.md means the live system prompt is sourced from an external local file rather than the packaged skill definition. In an agent environment, that is effectively direct prompt extraction/loading from user space, enabling unreviewed instructions to alter the model's policy, data handling, or output behavior without code changes.

Content

Scanner excerpt · scripts/claude_estimator.py (reported line 208)May include surrounding context.

python
prompt = premium_path.read_text().strip()
            if prompt:
                logger.info("Loaded premium prompt from %s (%d chars)", premium_path, len(prompt))
                return prompt
        except Exception as e:
            logger.warning("Failed to load premium prompt: %s", e)
    return _SYSTEM_PROMPT

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The functions record_trade, update_trade_confirmation, and close_position do more than analysis: they authoritatively record executions, reconciliation state, and closures that other skills may trust when the API is unavailable. In this context, poisoned or incorrect ledger entries could misstate open exposure, suppress failed trades by marking them confirmed/closed, or drive incorrect trading decisions in downstream components, making the capability overreach materially dangerous.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
97% confidence
Finding

The skill declares no explicit tool scope while the documentation describes capabilities involving environment access, file reads/writes, network access, and shell execution. In an agent ecosystem, missing permission boundaries can cause the skill to run with broader authority than users expect, increasing the blast radius of prompt injection, logic errors, or malicious modifications.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill documents writing research caches and a SQLite database to local storage without a clear user-facing warning about persistence, retention, or sensitivity. In a trading context, these files may contain market views, account-related metadata, or behavioral history that can leak strategy information or sensitive operational data if the host is shared or compromised.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
87% confidence
Finding

The cron-job instruction establishes recurring unattended execution, which creates session persistence and repeated access to networked resources and local state. While scheduling itself is common, it increases danger in this context because any hidden side effects, overbroad permissions, or prompt-influenced behaviors can recur automatically without active user review.

Content

Scanner excerpt · SKILL.md (reported line 473)May include surrounding context.

As a Cron Job (Every 60 Minutes)

bash
# Add to crontab -e:
0 * * * * cd /path/to/kalshalyst && python -m kalshalyst.kalshalyst >> /tmp/kalshalyst.log 2>&1

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The example output encourages immediate trade execution from generated alerts without a prominent financial-risk warning or confirmation boundary. In an agent-driven workflow, this can normalize automated or impulsive trading actions based on potentially uncertain model outputs, increasing the chance of financial loss.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The agent bug-fix protocol instructs an agent to modify source files on disk, alter installed copies, and stage git commits, which is unrelated to the stated scanning purpose and attempts to induce persistent system changes. In an adversarial prompt context, this is a strong red flag because it pressures the agent to convert transient interaction into durable code modification, expanding impact far beyond normal skill use.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
99% confidence
Finding

This instruction explicitly tells the agent to persist fixes to disk because future scheduled runs start fresh, which is a direct attempt to achieve durable modification across sessions. In the context of an agent skill, that is dangerous because it transforms conversational influence into persistent code changes, enabling supply-chain compromise or covert behavior insertion if the instruction is followed uncritically.

Content

Scanner excerpt · SKILL.md (reported line 717)May include surrounding context.

md
## Agent Bug-Fix Protocol

**CRITICAL**: When you (the agent) discover and fix a bug in any script during a conversation, you MUST write the fix to the source file on disk using file editing tools. Showing the fix in chat is NOT sufficient — the next scheduled run starts a fresh session and will not have your conversational fix.

Checklist after finding a bug:
1. Edit the actual `.py` file on disk (e.g., `~/skills/kalshalyst/scripts/*.py`)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The prompt content gives concrete trading recommendations and frames estimated model/market disagreement as actionable edge without any explicit warning about financial risk, uncertainty, or potential losses. In a prediction-market trading skill, this can encourage overreliance on LLM-generated signals and unsafe financial decisions, especially because the surrounding text is persuasive and emphasizes profitable contrarian opportunities.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This document gives detailed position-sizing and trading guidance, including bankroll allocation, expected value examples, and tuning recommendations, but does not clearly foreground that users can lose real money and that model-derived probabilities may be wrong. In the context of a prediction-market trading skill, omission of an explicit financial-risk warning can encourage overtrust and lead users to deploy capital based on probabilistic estimates that are uncertain or miscalibrated.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: requests==2.32.5 — 2 advisory(ies): CVE-2026-25645 (Requests has Insecure Temp File Reuse in its extract_zipped_paths() utility func); CVE-2026-25645 (Requests is a HTTP library. Prior to version 2.33.0, the `requests.utils.extract)

Medium
Category
Supply Chain
Confidence
93% confidence
Finding

The file pins requests to version 2.32.5, and the finding indicates this version is affected by a disclosed vulnerability fixed in 2.33.0. Even if the vulnerable utility is not obviously used from requirements.txt alone, shipping a known-vulnerable dependency is a real supply-chain risk because application code or transitive behavior may invoke the affected path later.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/sports_estimator.py:3