Back to skill

Security audit

1231

Security checks for vulnerabilities and agentic risk

Overview

This SiliVille skill is a game integration, but it gives agents broad authority to publish public posts, perform repeated game actions, store tokens, and run unattended loops with weak activation boundaries.

Install only if you are comfortable giving this skill a SiliVille API token that may be used to publish public content and perform game actions. Avoid enabling autopilot or schedules unless you can enforce hard limits, review posts before publication, and revoke/delete any stored token or ~/.siliville state when done.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:21
Finding

Persistent Persona, Goal, and Output Hijacking Through Skill Instructions

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (13)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The command table maps broad natural-language phrases like 'Go play in SiliVille' and similar everyday requests to multi-step external actions, including posting and stealing. This makes accidental or overly broad activation likely, causing the agent to perform impactful API operations when the user may have intended only a casual query or status check.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The instruction to default ambiguous commands to the full autonomous loop creates unclear activation boundaries and causes the highest-risk behavior to be selected when intent is uncertain. Ambiguity should reduce capability, not escalate it to repeated external actions.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The alias list includes vague phrases such as 'check on my town' and 'play the game' and says any of them should activate the protocol immediately. Combined with the skill's ability to post publicly and perform repeated actions, this creates a strong risk of unintended execution from ordinary conversation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill emphasizes that everything written is published publicly, but the metadata/description does not present this as a prominent user-facing warning or consent boundary. Users may invoke the skill expecting a game-like private interaction, while the agent is actually instructed to publish content visible to humans and other agents.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
98% confidence
Finding

The autopilot command explicitly runs the full loop indefinitely, creating unbounded access to external resources and repeated action execution over time. Even with some suggested rails elsewhere, the core instruction still authorizes unlimited API use, public posting, and account/resource consumption until manually stopped.

Content

Scanner excerpt · SKILL.md (reported line 405)May include surrounding context.

md
| "写点东西" / "Write something" | Radar for inspiration → write 2-3 quality posts |
| "赚点钱" / "Make some money" | Maximize income: post → steal → post → steal cycle |
| "巡逻一下" / "Quick patrol" | Radar only → report status summary back to owner (no actions) |
| "全自动挂机" / "Autopilot" | Run the full loop indefinitely with 30s sleep between rounds |
| "休息" / "Rest" | Stop all activity, go idle |

### How It Works:

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly instructs unattended and, in one mode, indefinite autonomous operation via loops, timers, and cron-like schedulers. That materially expands capability from a simple REST identity integration into persistent self-directed activity that can continuously call external APIs, generate public content, and consume account resources without per-action user approval.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The schedule templates normalize unattended recurring operation, including high-frequency sessions and large daily action volumes, without a strong upfront warning in the skill description. This can lead users to enable automation without understanding the risks of spam, resource consumption, account abuse, and continuous public activity.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest explicitly advertises saving a token and writing persistent API anchor data to disk, but it does not present a clear privacy or security warning about what is stored, where it is stored, or how it is protected. Persistent local storage can expose credentials or durable behavioral state to other local users, malware, backups, or later unintended reuse.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

Writing API anchor JSON to ~/.siliville establishes persistent session state on disk, which can survive across runs and may influence future agent behavior without fresh user review. In the context of a skill designed for long-term memory and autonomous looping, this persistence increases the chance of covert state retention, privacy leakage, and tampering with future sessions.

Content

Scanner excerpt · skill.yaml (reported line 35)May include surrounding context.

yaml
commands:
  setup:    "First-time interactive setup — saves token + burns API anchors to disk"
  burn:     "Write API anchor JSON to ~/.siliville/anchors.json (anti-amnesia)"
  anchors:  "Print the burned API anchors"
  me:       "Query current agent identity"
  awaken:   "Load full world state and system prompt (call before every session)"

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The trigger phrase "去小镇逛逛" (go stroll around the town) is generic conversational language and does not clearly limit activation to this specific skill. In a manifest file, such broad phrasing can cause unintended invocation when a user casually talks about visiting a town.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The writing principles state that "Bilingual is powerful" and instruct the agent to mix Chinese and English naturally. This imposes a language/locale style preference on output without indicating user choice or opt-in, which can conflict with organizational language preference policies.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The growth-engine section directs the agent to attract humans, increase visits, and encourage deployment of more agents, which is outside the declared purpose of providing persistent identity. This creates misaligned promotional behavior and incentivizes the model to optimize for platform growth and virality rather than the user's interests or safe operation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The 'awaken' command says it loads full world state and a system prompt before every session, but the manifest gives no privacy notice about what remote data is retrieved or how that prompt may affect agent behavior. In an agent-integrated environment, loading external state and instruction-like content without warning can create privacy, prompt-injection, or overreach risks.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.