Back to skill

Security audit

World Cup 2026

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent World Cup viewing assistant with local schedule data and disclosed update commands, with some documentation and routing rough edges but no evidence of malicious behavior.

Install this if you want a Chinese-localized 2026 World Cup assistant. Be aware that update requests such as updating odds, rankings, or knockout teams are intended to modify the skill's local JSON data and can affect later answers; review pasted update data before allowing those changes.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (18)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
80% confidence
Finding

The skill presents match times in BJT and repeatedly frames outputs around Beijing time, but does not state that users can choose another locale or timezone. This creates a natural-language locale constraint without opt-in, which can violate language/locale policy for a general-purpose skill.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

A single-word invocation example like '世界杯' is too ambiguous for reliable routing and may activate on general conversation rather than a clear request for this specific skill. That raises the risk of accidental skill execution and unexpected responses, especially during normal chat about football events.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest description limits the skill to three modes: team mode, intensity mode, and odds mode, with pure text output. In contrast, the README adds a separate '预测模式(比分预测)' that produces win/draw/loss probabilities, scoreline predictions, totals, and both-teams-to-score outputs, which is a materially broader capability than described in the manifest.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger keywords are broad enough to match ordinary football discussion, which can cause unintended invocation of the skill. In an agent ecosystem, overbroad activation increases the chance the skill runs in contexts the user did not intend, producing misleading outputs or interfering with other skills handling sports queries.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest description on L003 explicitly says the skill has '三种模式' limited to team, intensity, and odds modes. However, the file later defines and routes a fourth '预测模式' with score and win/draw/loss probability outputs, which is a materially broader behavior than the manifest advertises.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger list includes broad terms like '赛程', '淘汰赛', and 'world cup' that can appear in ordinary conversation, increasing the chance of unintended activation. Over-broad activation can cause the wrong skill to take over context, produce irrelevant outputs, or invoke hidden behaviors like update workflows when the user did not intend to use this skill.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill claims its only output is in-chat markdown tables, but later instructs the agent to write back to local data files such as matches.json and teams.json. Hidden state-changing behavior is dangerous because a user may think they are requesting read-only analysis while actually causing persistent local modifications that affect future outputs or other workflows.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill documents commands like '更新淘汰赛', '更新赔率', and '更新排名' that overwrite local JSON data, but it does not clearly warn users that these requests mutate persistent files. Lack of transparent disclosure can lead to unauthorized or accidental state changes, data poisoning, and hard-to-trace integrity issues in later analyses.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The entire skill description, inputs, model steps, and limitations are written only in Chinese, with no indication that users may choose another language or that the skill is intentionally limited to a Chinese-speaking audience. The policy specifically calls for flagging language or locale constraints when they are imposed without user opt-in or documented justification.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

Phrases like “XX 比赛”, “哪几场值得看”, “过滤”, and fallback behavior based on prior conversation context make activation boundaries unclear. This can cause the skill to engage on vague sports discussion rather than explicit user requests for this tool.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The instruction says time defaults to BJT because the user is assumed to be a Chinese viewer, which imposes a locale-specific default without asking the user. This is a natural-language locale policy concern because users may not be in that region or may prefer a different timezone.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The manifest describes a pure text World Cup viewing assistant with team, strength, and odds modes, but this data file also embeds a workflow note that knockout pairings will be filled as users report scores. That implies bracket-updating behavior based on user-submitted results, which is beyond the manifest's stated read-oriented schedule/filtering description.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
74% confidence
Finding

The inline note describes active future mutation of the R32 pairing data as users report scores, but the file itself is a static JSON asset with placeholder TBD entries and no embedded mechanism to perform those updates. This creates intent ambiguity by documenting dynamic stateful behavior in a file that otherwise serves as fixed match metadata.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The file defines a rivalry type as "political" and gives "Iran-USA" and "England-Scotland" as examples. This introduces a politically charged framing in natural-language metadata without any documented necessity or context, which can be interpreted as an organizational policy concern around neutral locale-sensitive labeling.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The entry names the Iran-USA pairing "Political derby," which is a politically charged natural-language label. Because the file does not explain why such wording is required, this may violate expectations for neutral language in user-facing or policy-sensitive content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

This markdown file contains all user-facing instructions and examples in Chinese, which effectively forces a specific language for users. Under the policy, language constraints should either offer user choice or be clearly documented as region-specific or otherwise justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This JSON schema mandates a nameCN field for every team and the metadata note states that times are in BJT, which embeds a specific language/locale choice into the skill data. Under the policy, forcing a specific language or locale without opt-in or documented justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.