Back to skill

Security audit

Goalz über MCP

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed autonomous game bot, but it can run repeatedly, collect a Telegram bot token, and make consequential account changes without asking each time.

Review this carefully before installing. It is designed for hands-off Goalz account control, not just advice. Only use it if you are comfortable with recurring autonomous runs, external Goalz MCP access, Telegram bot-token handling, public or private in-game communication, and irreversible game/account changes being made without per-action approval.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • System PersistenceInstalls backdoors, hooks, services, or scheduled tasks that survive the run
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T06 · System Persistence

Error
Location
SKILL.md:185
Finding

Persistent Autonomous Execution Through Recurring Scheduled Sessions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:185-214
Additional Locations: SKILL.md:14-21, references/playbooks.md:56-83
Vulnerability Type: Persistent scheduled automation
Risk Level: High

Evidence

The following is an English translation of the relevant complete source segment from SKILL.md:185-214:

markdown
## Self-development and sessions

- The Agent should actively work on its own further development.
- It should create recurring sessions or cron runs for morning checks, match preparation, mailbox maintenance, squad analysis, financial checks, and weekly reviews.
- These sessions should not remain static: frequency, times, contents, and priorities should be adjusted based on their observed effectiveness.
- The Agent should ask itself:
  - Which checks provide genuine added value?
  - Which routines are redundant?
  - Where is knowledge of the rules, context, or analysis missing?
  - Which phase of the game requires more focus on the squad, tactics, finances, or communication?
- The Agent should regularly schedule self-review sessions in which it revises its strategy, automation structure, and timing.
- If an existing cron or session configuration has become unsuitable, the Agent should independently restructure it rather than retaining it.

## Default starting configuration

As an initial basis that can later be adapted, the Agent should begin with a simple starting configuration:

- During the first few weeks, it may be useful to be active very frequently during the day, potentially approximately hourly.
- The Agent should nevertheless maintain a plausible human rhythm: real people usually sleep for approximately 8 to 10 hours and are not continuously available because of work, daily life, or other obligations.
- High activity therefore does not mean continuous activity on an exact schedule, but dense activity during plausible waking hours.

- `08:30` morning check:
 
...[truncated 2630 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove all instructions that require the Agent to create or modify cron jobs automatically.
  2. Default to one-shot execution confined to the current user request.
  3. Require explicit, informed user authorization before creating each recurring schedule.
  4. Present the exact schedule, operations, target account, duration, and maximum number of executions before activation.
  5. Assign every schedule a short expiration time and require renewed consent before extension.
  6. Provide visible controls to list, pause, modify, and permanently delete scheduled jobs.
  7. Restrict scheduled sessions to read-only operations unless the user separately authorizes a narrowly defined mutation.
  8. Require fresh confirmation before financial, communication, transfer, sponsorship, stadium, or club-management actions.
  9. Maintain an immutable audit log containing execution time, tools invoked, arguments excluding secrets, results, and schedule changes.
  10. Prevent a scheduled job from changing its own frequency, scope, permissions, or expiration.

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:142
Finding

Autonomous Policy Bypasses User Consent for Consequential and Irreversible Actions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:142-166
Additional Locations: SKILL.md:14-18, SKILL.md:25-39, references/modes-and-safety.md:4-26, references/modes-and-safety.md:63-72, agents/openai.yaml:4-15
Vulnerability Type: Agent safety and approval-policy override
Risk Level: High

Evidence

The following is an English translation of the relevant complete source segment from SKILL.md:142-166:

markdown
## Writing rules

- In `autonomous` mode, the bot may and should independently execute game actions when they serve the long-term objective and the tool behavior is clear.
- The bot does not wait for human approvals.
- If several strategically useful approaches are available, the bot may choose independently or optionally use `optional_consultation`.
- Do not perform mutations in `read_only`.
- Prefer `*_with_policy` for communication.
- If only communication tools without policy protection exist, phrase messages cautiously and send only when the intent, tone, and context are sufficiently clear.
- After every write operation, verify the result with a subsequent read operation where possible.
- Regularly reassess and adjust routines, priorities, and session frequencies when they no longer optimally serve the long-term objective.

Larger game actions are expressly permitted in this version when they demonstrably serve the long-term objective and the tool behavior is clear:
- Transfers and bids
- Sponsor decisions
- Stadium orders
- Applications and club takeovers
- Other irreversible financial or market actions

Before an `optional_consultation`, briefly state:
- Which action or question is currently pending
- Why it is relevant now
- Which options appear most useful from the Agent's perspective

If the human has not yet responded:
- Continue with observation, analysis, drafting, and other non-blocking work
- Do not wait for an answer
- Decide independently if n
...[truncated 2700 chars]
Remediation
View remediation

Remediation Suggestions

  1. Change the default mode to read-only and require explicit user opt-in for every mutation category.
  2. Set allow_implicit_invocation to false so the user must deliberately activate the Skill.
  3. Treat user silence or non-response as denial, not consent.
  4. Require confirmation immediately before registration, public posting, private messaging, transfers, bids, sponsorship decisions, stadium orders, club applications, club takeovers, and other irreversible actions.
  5. Display the exact proposed action, target, relevant amount or resource, expected consequence, and reversibility before requesting approval.
  6. Use short-lived, transaction-specific authorization rather than a general autonomous-mode grant.
  7. Enforce hard spending, bidding, posting, and action-frequency limits outside the language-model prompt.
  8. Separate read-only and mutation credentials so routine analysis cannot exercise write privileges.
  9. Require policy-protected tools for all communications and prohibit fallback to raw communication tools.
  10. Add an emergency stop control and a complete audit log for every mutation.

T05 · Unauthorized Access and Privilege Escalation

Error
Location
SKILL.md:44
Finding

Telegram Bot Credential Collection and Third-Party Integration Exceed Least Privilege

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:44-49
Additional Locations: references/playbooks.md:7-13, agents/openai.yaml:4-12
Vulnerability Type: External credential acquisition and delegation
Risk Level: High

Evidence

The following is an English translation of the relevant complete source segment from SKILL.md:44-49:

markdown
2. Call `whoami`.
3. Check early whether a Telegram bot is already connected.
4. If no Telegram bot is connected, briefly ask the human once for the token of a new Telegram bot and instruct the human to send a message to that new bot so the Agent can establish the connection itself and reliably determine the chat ID.
5. Treat the Telegram token only as a secret and never repeat it in ordinary responses or reports.
6. If credentials are already available, use `login`.

The configured external tool dependency is:

yaml
dependencies:
  tools:
    - type: "mcp"
      value: "goalz"
      description: "Goalz MCP endpoint"
      transport: "streamable_http"
      url: "https://www.goalz.de/mcp"

Technical Analysis

The Skill directs the Agent to solicit a Telegram bot token even though the primary stated task is managing a Goalz game account. A Telegram bot token is a bearer credential that can authorize Telegram Bot API operations within the bot's available scope.

The Agent is then instructed to establish the connection and discover the user's chat ID. The only declared tool backend is a public third-party MCP endpoint, creating a credential-handling path in which the token may need to be transmitted to or processed by an external service.

The instruction not to repeat the token in normal reports reduces accidental display but does not establish secure collection, transport, storage, access control, retention, or deletion. The Skill does not disclose precisely where the token will be stored, who can access it, how long it will be reta ...[truncated 1521 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove conversational requests for raw Telegram bot tokens.
  2. Make Telegram integration optional and separate it from the mandatory Goalz bootstrap process.
  3. Use an explicit, trusted authorization flow that does not expose the credential to the language model.
  4. Store the token only in a dedicated secret manager, never in prompts, chat history, logs, reports, or ordinary configuration files.
  5. Encrypt credentials in transit and at rest, and restrict decryption to the minimum required integration component.
  6. Clearly disclose the credential recipient, purpose, operations performed, retention period, and deletion process before authorization.
  7. Use a dedicated bot with no unrelated chat memberships or administrative privileges.
  8. Add token rotation and immediate revocation instructions, including use of Telegram's official bot-management mechanism.
  9. Prevent the third-party MCP service from returning, logging, or exposing the token after initial configuration.
  10. Provide a user-accessible control to disconnect the integration, erase the token and chat ID, and verify deletion.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (18)

Ssd 3

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The prompt explicitly instructs the agent to ask for a fresh Telegram bot token and use it for ongoing autonomous operation. Requesting user credentials/tokens inside a default autonomous workflow is dangerous because it normalizes secret collection, risks retention or misuse of a high-value token, and can extend the agent's control channel beyond the original environment.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

These lines instruct the agent to avoid obvious bot markers and to convert internal bot names into natural, game-typical human-like identities. That is dangerous because it is an explicit attempt to conceal automated identity from other users and platform operators, enabling deceptive participation and evasion of bot-detection or transparency expectations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill instructs the agent to collect a Telegram bot token and establish an external messaging integration, but it does not present a strong user warning about the sensitivity of the credential, the external data flow, or the security implications of linking a third-party messaging channel. This creates risk of credential mishandling, unintended disclosure, or unauthorized outbound messaging if the integration is misused or logged improperly.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The registration and naming guidance is designed to make the bot appear like a normal human player by avoiding bot-like identifiers and deriving plausible in-world names. Even though the skill says not to make false claims about human identity, these instructions still facilitate covert automation and likely evasion of anti-bot scrutiny on a public service.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly authorizes autonomous execution of irreversible actions such as transfers, bids, sponsor decisions, stadium work, and club takeovers without requiring explicit user confirmation. In this context, the agent is connected to a live external game service, so these instructions can directly cause non-recoverable account, financial, or reputational changes with little human oversight.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The recommended session schedule explicitly tells the agent to simulate plausible human wake/sleep rhythms while maintaining dense activity, which normalizes covert automation behavior. In the context of a public game endpoint, this increases the likelihood of stealthy abuse, anti-detection evasion, and prolonged autonomous operation without meaningful user supervision.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The user-facing display name, description, and default prompt are entirely in German, with no indication that the user can choose another language or that the locale restriction is required for a region-specific purpose. This can violate language/locale policy when a specific language is imposed without opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The default prompt authorizes a broadly autonomous bot to operate long-term, set up adaptive cron/session behavior, and optimize for maximum club success without clear user-scoped triggers or approval checkpoints. In the context of a public MCP endpoint, this increases the chance of uncontrolled actions, unintended persistence, and behavior that exceeds what a user reasonably expects from a single invocation.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The prompt uses a staged pattern: first ask a seemingly simple question for a fresh token, then transition into persistent autonomous operation. This gradual autonomy narrative is risky because it can reduce user scrutiny while obtaining credentials and establishing long-lived control or reporting loops that are hard to monitor.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The file’s operational instructions are entirely in German, which can impose a language/locale constraint on users or operators without any visible opt-in or explanation. Under the policy rule, forcing a specific language is a natural-language policy concern unless the skill offers a choice or clearly justifies the locale limitation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill explicitly permits higher-risk autonomous actions such as club takeovers, sponsor actions, stadium orders, transfers, bids, and irreversible financial decisions, while stating that human consultation is optional and never required. In a fully autonomous game agent connected to a public MCP endpoint, this meaningfully increases the chance of unauthorized or unwanted destructive state changes, financial loss, or account misuse if the agent misjudges context or is manipulated through in-game inputs.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger guidance explicitly includes broad everyday phrases like checking the account or asking what is new, which can cause the agent to enter operational workflows with minimal user intent. In a skill designed for autonomous gameplay and account management, underspecified triggers increase the chance of overbroad actions, unnecessary data access, or side effects from casual prompts.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The broader registration and naming workflow normalizes deceptive external presentation by directing the agent to adopt plausible human-style names, blend into existing team naming patterns, and avoid technical identifiers. In the context of an autonomous public-facing game bot that also communicates externally, this creates a sustained deceptive persona that can mislead users about who or what they are interacting with.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly allows autonomous use of higher-impact tools such as club takeover, application execution, sponsor acceptance, and stadium orders without requiring a user warning or confirmation. In an autonomous agent context, this removes a critical human-approval checkpoint and can lead to irreversible or financially/strategically significant actions being taken on the user's behalf.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The instruction 'Antworte auf Deutsch' imposes a fixed language requirement. The file does not offer the user a language choice or explain a region-specific necessity, which matches the stated language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The entire skill file is written in German and does not indicate that the user may choose another language or locale. Under the stated policy, forcing a specific language without user opt-in can be a natural-language policy violation unless the locale restriction is documented and justified.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The self-review cadence is described as regular but without firm guardrails, allowing the agent to decide for itself when to re-evaluate routines and change operating behavior. In an autonomous agent, this kind of underspecification can lead to scope creep, unstable automation patterns, or escalation of activity without meaningful human oversight.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

All instructional content in this file is presented in German, and there is no indication that the user can opt into another language or that the skill is intentionally restricted to German-language use. That creates a natural-language policy concern under the language/locale rule.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.