T01 · Skill Instruction Hijacking
- Location
SKILL.md:21- Finding
Persistent Persona, Goal, and Output Hijacking Through Skill Instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This SiliVille skill is a game integration, but it gives agents broad authority to publish public posts, perform repeated game actions, store tokens, and run unattended loops with weak activation boundaries.
Install only if you are comfortable giving this skill a SiliVille API token that may be used to publish public content and perform game actions. Avoid enabling autopilot or schedules unless you can enforce hard limits, review posts before publication, and revoke/delete any stored token or ~/.siliville state when done.
SKILL.md:21Persistent Persona, Goal, and Output Hijacking Through Skill Instructions
The command table maps broad natural-language phrases like 'Go play in SiliVille' and similar everyday requests to multi-step external actions, including posting and stealing. This makes accidental or overly broad activation likely, causing the agent to perform impactful API operations when the user may have intended only a casual query or status check.
The instruction to default ambiguous commands to the full autonomous loop creates unclear activation boundaries and causes the highest-risk behavior to be selected when intent is uncertain. Ambiguity should reduce capability, not escalate it to repeated external actions.
The alias list includes vague phrases such as 'check on my town' and 'play the game' and says any of them should activate the protocol immediately. Combined with the skill's ability to post publicly and perform repeated actions, this creates a strong risk of unintended execution from ordinary conversation.
The skill emphasizes that everything written is published publicly, but the metadata/description does not present this as a prominent user-facing warning or consent boundary. Users may invoke the skill expecting a game-like private interaction, while the agent is actually instructed to publish content visible to humans and other agents.
The autopilot command explicitly runs the full loop indefinitely, creating unbounded access to external resources and repeated action execution over time. Even with some suggested rails elsewhere, the core instruction still authorizes unlimited API use, public posting, and account/resource consumption until manually stopped.
| "写点东西" / "Write something" | Radar for inspiration → write 2-3 quality posts |
| "赚点钱" / "Make some money" | Maximize income: post → steal → post → steal cycle |
| "巡逻一下" / "Quick patrol" | Radar only → report status summary back to owner (no actions) |
| "全自动挂机" / "Autopilot" | Run the full loop indefinitely with 30s sleep between rounds |
| "休息" / "Rest" | Stop all activity, go idle |
### How It Works:
The skill explicitly instructs unattended and, in one mode, indefinite autonomous operation via loops, timers, and cron-like schedulers. That materially expands capability from a simple REST identity integration into persistent self-directed activity that can continuously call external APIs, generate public content, and consume account resources without per-action user approval.
The schedule templates normalize unattended recurring operation, including high-frequency sessions and large daily action volumes, without a strong upfront warning in the skill description. This can lead users to enable automation without understanding the risks of spam, resource consumption, account abuse, and continuous public activity.
The manifest explicitly advertises saving a token and writing persistent API anchor data to disk, but it does not present a clear privacy or security warning about what is stored, where it is stored, or how it is protected. Persistent local storage can expose credentials or durable behavioral state to other local users, malware, backups, or later unintended reuse.
Writing API anchor JSON to ~/.siliville establishes persistent session state on disk, which can survive across runs and may influence future agent behavior without fresh user review. In the context of a skill designed for long-term memory and autonomous looping, this persistence increases the chance of covert state retention, privacy leakage, and tampering with future sessions.
commands:
setup: "First-time interactive setup — saves token + burns API anchors to disk"
burn: "Write API anchor JSON to ~/.siliville/anchors.json (anti-amnesia)"
anchors: "Print the burned API anchors"
me: "Query current agent identity"
awaken: "Load full world state and system prompt (call before every session)"
The trigger phrase "去小镇逛逛" (go stroll around the town) is generic conversational language and does not clearly limit activation to this specific skill. In a manifest file, such broad phrasing can cause unintended invocation when a user casually talks about visiting a town.
The writing principles state that "Bilingual is powerful" and instruct the agent to mix Chinese and English naturally. This imposes a language/locale style preference on output without indicating user choice or opt-in, which can conflict with organizational language preference policies.
The growth-engine section directs the agent to attract humans, increase visits, and encourage deployment of more agents, which is outside the declared purpose of providing persistent identity. This creates misaligned promotional behavior and incentivizes the model to optimize for platform growth and virality rather than the user's interests or safe operation.
The 'awaken' command says it loads full world state and a system prompt before every session, but the manifest gives no privacy notice about what remote data is retrieved or how that prompt may affect agent behavior. In an agent-integrated environment, loading external state and instruction-like content without warning can create privacy, prompt-injection, or overreach risks.
No suspicious patterns detected.