Back to skill

Security audit

Fs Worldcup Knockout

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly coherent for play-money propSPACE trading, but it uses weak default credentials and stores account tokens locally in plaintext.

Install only if you are comfortable with a play-money trading bot that can create a propSPACE account, store a bearer token on disk, and place account-affecting trades when --live is used. Set a unique FS_PASSWORD, keep .auth out of archives and source control, use only the official HTTPS FS_BASE_URL unless you intentionally test against a trusted endpoint, and avoid the PAT-in-URL and unpinned npx setup commands where possible.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
main.py:57
Finding

Unrestricted API endpoint can receive account credentials and bearer tokens

Content
View full analysis
dict: resp = self._post("/auth/signup", { "username": username, "password": password, }) self._persist_token(resp) return resp["user"] def _login(self, username: str, password: str) -> dict: resp = self._post("/auth/login", { "username": username, "password": password, }) self._persist_token(resp) return resp["user"] def _post(self, path: str, body: dict) -> dict: data = json.dumps(body).encode() req = urllib.request.Request( self.base + path, data=data, headers={"Content-Type": "application/json"}, ) return self._send(req) def _authed_post(self, path: str, body: dict) -> dict: data = json.dumps(body).encode() headers = {"Content-Type": "application/json"} if self.token: headers["Authorization"] = f"Bearer {self.token}" req = urllib.request.Request(self.base + path, data=data, headers=headers) return self._send(r ...[truncated 2154 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
main.py:57
Finding

Public shared default password enables predictable account authentication

Content
View full analysis
| `FS_PASSWORD` | `simmer-wc-bot` | Account password; set a stronger value for production bots | ``` ```json { "name": "FS_PASSWORD", "required": false, "description": "Account password (min 6 chars). Default: 'simmer-wc-bot'. Set a strong value for production bots." } ``` ### Technical Analysis When `FS_PASSWORD` is not supplied, the Skill authenticates or creates accounts using the publicly documented shared password `simmer-wc-bot`. Because account creation is automatic, users may unknowingly establish persistent accounts with this known credential. The documentation also refers to the process as “passwordless,” but the implementation sends a password to the backend. This mismatch may cause users to underestimate the need to configure and protect a unique password. A shared default password is not secret and cannot provide meaningful account authentication. Its risk is especially high when usernames appear in logs, reports, test briefs, screenshots, or predictable automation naming conventions. ### Attack Path 1. A user runs the Skill with `FS_USERNAME` configured but without `FS_PASSWORD`. 2. The Skill logs in or creates the account using `simmer-wc-bot`. 3. An attacker learns or predicts the username from logs, command examples, shared output, or naming conventions. 4. The attacker submits the known username and public default password to the engine's login endpoint. 5. If authentication succeeds, the attacker receives an account token and performs any operation allo ...[truncated 623 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
fs_client.py:273
Finding

Long-lived bearer token is persisted without explicit restrictive permissions

Content
View full analysis
None: self.token = resp.get("access_token") user = resp.get("user") or {} self.user_id = user.get("user_id") if self._token_store and self.token: self._token_store.parent.mkdir(parents=True, exist_ok=True) self._token_store.write_text(json.dumps({ "access_token": self.token, "user_id": self.user_id, }, indent=2)) def _load_stored_token(self) -> None: if self._token_store and self._token_store.exists(): try: data = json.loads(self._token_store.read_text()) self.token = data.get("access_token") self.user_id = data.get("user_id") except Exception: pass ``` ### Technical Analysis The client stores the bearer token in `.auth/.json` inside the Skill directory. The code relies entirely on the process umask and does not explicitly enforce a private directory mode or file mode. The client documentation states that tokens have an approximately 60-day lifetime. A bearer token grants access to whoever possesses it, so local disclosure is sufficient for account impersonation. Storing it in the project tree also increases the chance that it will be copied into archives, support bundles, backups, container layers, or source-control commits. The username is used directly in the token filename. Although typical usernames may be benign, the code does not validate that the resulting path remains inside `.auth`, creating an add ...[truncated 993 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
HERMES-TEST.md:14
Finding

Test documentation embeds a GitHub personal access token in the clone URL

Content
View full analysis
git clone --depth=1 https://$GITHUB_TOKEN@github.com/SupaFund/simmer.git /root/projects/simmer cd /root/projects/simmer/skills/fs-worldcup-knockout ``` ``` ### Technical Analysis The documented command interpolates a GitHub personal access token directly into a command-line URL. Depending on the operating system, shell, Git configuration, diagnostics, and CI environment, credential-bearing arguments may be exposed through: - Process listings and process-monitoring tools. - Shell tracing or CI job logs. - Error messages and debugging output. - Audit logs. - Credential-bearing Git remote configuration or copied command history. The suggested token is read-only and repository-scoped, which limits impact, but it remains a reusable secret and should not be placed in process arguments or URLs. ### Attack Path 1. A tester follows the documented setup instructions. 2. The shell expands `$GITHUB_TOKEN` into the `git clone` command's URL. 3. A local observer, CI logger, tracing facility, or diagnostic collector records the expanded command. 4. The attacker extracts the PAT. 5. The attacker authenticates to GitHub and reads the scoped private repository until the token expires or is revoked. ### Impact Assessment The attacker can obtain the GitHub permissions granted to the PAT. Under the recommended scope, this includes read access to the `SupaFund/simmer` repository and may expose source code, repository history, configuration, ...[truncated 211 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
FUNCTIONSPACE_HANDOFF.md:20
Finding

Installation instructions execute an unpinned mutable npm package release

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (21)

Credential Access

High
Category
Privilege Escalation
Confidence
89% confidence
Finding

The test brief directs users to create and export a GitHub personal access token and to embed it directly in the git clone URL, which exposes sensitive credentials to shell history, process listings, logs, and copy/paste leakage. Even though the requested scope is read-only, the workflow still normalizes unsafe credential handling and could lead to token compromise or reuse in other environments.

Content

Scanner excerpt · HERMES-TEST.md (reported line 16)May include surrounding context.

A) GitHub PAT (recommended — works for all future tests too)

Create a fine-grained PAT at github.com → Settings → Developer settings → Personal access tokens → Fine-grained tokens. Scope: repository SupaFund/simmer, permission Contents: Read-only. Then:

bash

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

If the skill performs account signup/login, token persistence, and generic trading actions not clearly disclosed in its declared purpose, users and reviewers may authorize it under false assumptions. Hidden authentication state and broader-than-advertised market actions can lead to unanticipated account changes, credential handling, and unauthorized trading behavior.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

If the skill performs account signup/login, token persistence, and generic trading actions not clearly disclosed in its declared purpose, users and reviewers may authorize it under false assumptions. Hidden authentication state and broader-than-advertised market actions can lead to unanticipated account changes, credential handling, and unauthorized trading behavior.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The install command uses npx --yes clawhub@latest, which fetches and executes the latest package version at runtime rather than a reviewed, immutable version. If the upstream package is compromised or a breaking/malicious release is published, users following the handoff can execute attacker-controlled code during installation.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

The published-package verification step again invokes npx --yes clawhub@latest, causing verification itself to depend on an unpinned remote package. That weakens supply-chain integrity because the verification workflow can silently change over time or execute malicious code if the package source is compromised.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

Using npx --yes clawhub@latest publish . --version 0.1.3 runs a network-fetched, mutable package in a publishing path. If an attacker controls or compromises the upstream package, they could tamper with the publish process, exfiltrate credentials, or publish altered artifacts.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

The fresh-install verification command also relies on npx --yes clawhub@latest, preserving the same unpinned supply-chain execution risk. Because this skill handles credentials and can perform live play-money mutations, compromise of the installer/tooling could plausibly expose secrets or alter behavior beyond a harmless documentation issue.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document instructs the operator to run python3 main.py --live, which will place real trades and change account state, but it does not present a prominent safety warning, confirmation step, or explicit statement that this performs irreversible live actions. In the context of a trading skill, operational instructions that can spend balance or open positions without strong warning materially increase the risk of accidental execution and unintended financial exposure.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill advertises network, environment, and file capabilities through its documented behavior, but it does not declare any explicit tool scope or permission boundaries. In an agent ecosystem, this weakens least-privilege controls and makes it easier for the skill to access secrets, write files, or perform network actions beyond what a reviewer or orchestrator expects.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The instruction 'Correct it in _market_position() if needed' is an implementation directive tied to Python code, effectively prescribing a specific language/tooling path without offering user choice. This can violate language or locale policy where the skill should not force one implementation language unless explicitly justified or user-selected.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest documents a default password value ('simmer-wc-bot') for FS_PASSWORD. Even though this appears in configuration text, publishing a predictable default credential materially increases the chance that deployed bots or auto-created accounts will share the same password, enabling account takeover, impersonation, or abuse if the service is internet-accessible.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code writes a long-lived bearer token to disk in plaintext JSON without access controls, encryption, or any warning to the caller. If another local user, malware, backup system, or inadvertently shared workspace can read that file, they can reuse the token to act as the user against the trading API until expiry.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/enrich_from_web.py (reported line 119)May include surrounding context.

python
"freshness":      "pm",
        "extra_snippets": "true",
    })
    url = f"https://api.search.brave.com/res/v1/web/search?{params}"
    req = urllib.request.Request(url, headers={
        "Accept":               "application/json",
        "Accept-Encoding":      "gzip",

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/enrich_from_web.py (reported line 186)May include surrounding context.

python
"freshness":      "pm",
        "extra_snippets": "true",
    })
    url = f"https://api.search.brave.com/res/v1/web/search?{params}"
    req = urllib.request.Request(url, headers={
        "Accept":               "application/json",
        "Accept-Encoding":      "gzip",

Vague Triggers

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The manifest description says to use the skill when a user wants to trade World Cup knockout fantasy-score markets, but it does not define concrete trigger phrases, boundaries, or negative examples. In a manifest/markdown context, this can create ambiguity about when the skill should activate versus when general discussion of World Cup fantasy or betting should not invoke it.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The env var description says FS_MAX_COLLATERAL defaults to 50, while the configuration table says the default is 333. This is an active contradiction in the skill's own documentation about trading behavior and bankroll sizing.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The env var description says FS_RECIPE_SHAPE defaults to multimodal density, while the configuration table lists the default as multimodal. These may refer to the same mode, but as written they are inconsistent and can mislead users about accepted values.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The description says the username is "Auto-created on first run," but the manifest does not clarify what exact action constitutes a first run or under what conditions account creation occurs. In a manifest file, this can create ambiguity about when the skill performs account-affecting actions versus when it only reads configuration.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The manifest describes a skill for trading World Cup knockout-stage fantasy-score markets using an evidence-backed distribution strategy, but this file also implements full account lifecycle support including signup, login, and token persistence across runs. While authentication is related to trading, automatic account creation and credential-backed session management go beyond the narrower strategy-focused behavior promised in the manifest.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The skill's stated purpose is to trade specific propSPACE markets based on modeled beliefs, but the code can create new platform accounts when login fails. Creating accounts is a separate platform-management capability that is not explicitly disclosed or obviously required for a market-trading strategy skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

This code persists merged sentiment data back to player_data.json, changing local user/project data. Although the module docstring states that sentiment fields are written back and a --dry-run mode exists, the actual write path performs the modification immediately without any confirmation prompt or inline warning at the point of mutation.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution, suspicious.exposed_secret_literal

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_fs_worldcup_knockout.py:17

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
FUNCTIONSPACE_HANDOFF.md:27