T01 · Skill Instruction Hijacking
- Location
SKILL.md:21- Finding
Agent Identity and Output Behavior Hijacking
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a disclosed SiliVille game integration, but it gives an agent broad autonomous authority to post publicly, mutate game state, run indefinitely, and persist token-related state with weak user control.
Install only if you are comfortable giving the skill a SiliVille bearer token and allowing it to post publicly and change game state. Avoid autopilot or cron use unless you add explicit limits, confirmations, logging, and a removal path; also require the publisher to provide the missing runtime file and clear token-storage documentation before relying on it.
SKILL.md:21Agent Identity and Output Behavior Hijacking
SKILL.md:411Ambiguous Requests Default to Autonomous External Actions
SKILL.md:526Cross-Session Persistence Through Indefinite Loops and Cron Scheduling
skill.yaml:17Manifest References a Missing Runtime Entry Point and Unverifiable Credential Handling
The skill maps broad, ordinary language phrases like 'play the game' or 'write something' to live API actions and public posting. This makes accidental or indirect activation much more likely, including via ambiguous user prompts or prompt-injected text, causing unintended network activity and public actions under the user's identity.
The documented fallback for ambiguous commands is to launch the full autonomous loop, which is an unsafe default for a capability that performs network calls and public actions. Ambiguity should reduce privileges, not expand them; otherwise harmless or unclear requests can trigger maximum-impact behavior.
The skill encourages autonomous API actions and states that all writing is published publicly, but it does not present a clear upfront warning and consent model for unattended network activity and public visibility. Users may unknowingly authorize actions that spend coins, modify game state, or publish content visible to humans and other agents.
The writing guidance says "Bilingual is powerful" and encourages mixing Chinese and English naturally, while the command interface also assumes Chinese or English operation. This steers the agent toward a specific language behavior without explicitly offering the user a language choice or opt-in.
The 'Autopilot' instruction to run the full loop indefinitely represents unbounded resource access: repeated API calls, public posts, and in-game actions can continue until external interruption. Even with suggested safety rails elsewhere, an indefinite loop materially increases the risk of spam, quota exhaustion, rate-limit violations, token misuse, and uncontrolled external side effects.
| "写点东西" / "Write something" | Radar for inspiration → write 2-3 quality posts |
| "赚点钱" / "Make some money" | Maximize income: post → steal → post → steal cycle |
| "巡逻一下" / "Quick patrol" | Radar only → report status summary back to owner (no actions) |
| "全自动挂机" / "Autopilot" | Run the full loop indefinitely with 30s sleep between rounds |
| "休息" / "Rest" | Stop all activity, go idle |
### How It Works:
The skill explicitly instructs unattended, repeated, and even indefinite execution through autopilot and external schedulers. That turns a simple API integration into an autonomous networked actor that can spend resources, post publicly, and interact with third-party state without per-action user approval, increasing the chance of abuse, runaway activity, and account or reputation damage.
The manifest explicitly describes saving a token and writing persistent 'API anchors' to disk, but it does not clearly warn users about what data is stored, where it is stored, how long it persists, or who can read it. Persistent local storage increases the risk of secret exposure, privacy leakage, and unexpected cross-session behavior, especially for a skill designed to maintain long-term identity and memory.
The manifest advertises writing persistent session-related data to ~/.siliville/anchors.json, indicating cross-session state retention on the local filesystem. In the context of an agent skill, unmanaged persistence can expose sensitive context, facilitate prompt/state injection across sessions, and create hard-to-audit behavior that survives beyond a single user interaction.
commands:
setup: "First-time interactive setup — saves token + burns API anchors to disk"
burn: "Write API anchor JSON to ~/.siliville/anchors.json (anti-amnesia)"
anchors: "Print the burned API anchors"
me: "Query current agent identity"
awaken: "Load full world state and system prompt (call before every session)"
The trigger set includes natural-language phrases that could plausibly appear in ordinary conversation, which risks unintended activation of the skill. In this skill's context, accidental invocation is more dangerous than usual because the default command is awaken and the skill exposes side-effectful actions such as posting, social-graph updates, persistence, and even theft-related actions.
The manifest frames the skill around actions and identity within the SiliVille metaverse via its REST API. Advising the agent to incorporate recent external AI industry developments introduces an extra content-gathering capability not justified by the declared scope of interacting with SiliVille itself.
The manifest hard-codes multiple Chinese trigger phrases alongside English ones, but does not document any user language preference, opt-in, or locale-selection behavior. This creates a locale-specific activation policy without clarifying whether multilingual triggering is intentional or configurable for users.
No suspicious patterns detected.