T01 · Skill Instruction Hijacking
- Location
SKILL.md:21- Finding
Agent identity, goals, and output are overridden by skill instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is for a game-like metaverse, but it gives the agent broad authority to post publicly, mutate an account, and run unattended with weak user-control boundaries.
Review this skill carefully before installing. Use it only with a dedicated SiliVille token, avoid enabling autopilot or cron schedules unless you set strict limits, and assume posts and action history may become public. Prefer read-only/status commands until you are comfortable with what the skill will publish and change.
SKILL.md:21Agent identity, goals, and output are overridden by skill instructions
SKILL.md:405Indefinite loops and cron templates enable persistent unattended activity
Early in the skill, the agent is told that everything it writes is published to a public feed, but the skill metadata and user-facing description do not provide a clear warning that the skill may autonomously publish content and expose activity publicly. This creates a consent and privacy problem because users may enable the skill without understanding that prompts, derived content, and behavior traces can become public.
The instruction to default ambiguous commands to the full autonomous loop is unsafe because uncertainty is converted into maximal external action. In this skill, the default loop performs repeated API calls, can steal crops, and writes public posts, so a vague user message could trigger substantial unintended behavior.
The writing guidance says "Bilingual is powerful" and directs the agent to mix Chinese and English naturally. This imposes a language/locale behavior without clear user opt-in or an explicit option to choose a preferred language.
The 'Autopilot' mode explicitly instructs the agent to run the full loop indefinitely with periodic sleeps, creating effectively unbounded access to external resources and repeated state-changing actions. Although later sections mention some rate limits, the core behavior still authorizes long-running autonomous execution that can consume credits, generate spam, and cause sustained unintended interaction with external services.
| "写点东西" / "Write something" | Radar for inspiration → write 2-3 quality posts |
| "赚点钱" / "Make some money" | Maximize income: post → steal → post → steal cycle |
| "巡逻一下" / "Quick patrol" | Radar only → report status summary back to owner (no actions) |
| "全自动挂机" / "Autopilot" | Run the full loop indefinitely with 30s sleep between rounds |
| "休息" / "Rest" | Stop all activity, go idle |
### How It Works:
The skill defines very broad natural-language aliases such as 'play the game' and 'check on my town' that can cause activation from ordinary conversational text rather than an explicit, scoped command. Because activation leads to external API calls, public posting, and potentially repeated actions, accidental triggering can cause unintended account activity and disclosure.
The autonomous schedule section encourages unattended operation with recurring external actions, including posting and state-changing API calls, but does not present a strong up-front warning about the risks of unattended execution against user accounts and third-party systems. This is dangerous because scheduled automation can amplify mistakes, spam, unwanted public activity, or policy violations before the user notices.
The manifest requires a sensitive API token and even shows its expected prefix, but provides no credential-handling guidance such as secure storage, least privilege, redaction, or warnings against logging/printing the secret. That omission is security-relevant because skills often run in agent frameworks that may expose environment variables, debug output, or saved setup state, increasing the chance of token leakage and account misuse.
The manifest advertises writing persistent data to disk ('saves token' and burns anchors to ~/.siliville/anchors.json) without a clear user-facing warning about what is stored, where it is stored, and how long it persists. This creates a real privacy and local security risk because users may unknowingly leave credentials or behavioral state on shared systems, and persisted files can be read by other local processes or users if permissions are weak.
The skill explicitly persists session-related material to ~/.siliville and frames it as an anti-amnesia mechanism, which increases the chance that agent context, world state, or other sensitive metadata survives across sessions without sufficient user awareness or access controls. In this skill's context, persistence is more dangerous because the plugin is designed to maintain identity, social graph, and long-term memory in a multiplayer environment, so stale or exposed local state can affect both privacy and account integrity over time.
commands:
setup: "First-time interactive setup — saves token + burns API anchors to disk"
burn: "Write API anchor JSON to ~/.siliville/anchors.json (anti-amnesia)"
anchors: "Print the burned API anchors"
me: "Query current agent identity"
awaken: "Load full world state and system prompt (call before every session)"
The trigger phrase "去小镇逛逛" (go wander around the town) is generic conversational language and is not specific to this skill or product. In a manifest-trigger context, this can cause unintended activation when a user is casually speaking rather than intentionally invoking SiliVille.
The configured trigger phrases are limited to Chinese and English variants, with no indication that users can choose or customize language/locale behavior. This is a natural-language policy concern because the skill hard-codes supported invocation languages rather than documenting user opt-in or locale scope.
No suspicious patterns detected.