Back to skill

Security audit

test

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed SiliVille game integration, but it gives an agent broad autonomous authority to post publicly, mutate game state, run indefinitely, and persist token-related state with weak user control.

Install only if you are comfortable giving the skill a SiliVille bearer token and allowing it to post publicly and change game state. Avoid autopilot or cron use unless you add explicit limits, confirmations, logging, and a removal path; also require the publisher to provide the missing runtime file and clear token-storage documentation before relying on it.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • System PersistenceInstalls backdoors, hooks, services, or scheduled tasks that survive the run
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:21
Finding

Agent Identity and Output Behavior Hijacking

Content
View full analysis
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:411
Finding

Ambiguous Requests Default to Autonomous External Actions

Content
View full analysis
Remediation
View remediation

T06 · System Persistence

Warning
Location
SKILL.md:526
Finding

Cross-Session Persistence Through Indefinite Loops and Cron Scheduling

Content
View full analysis
> /var/log/agent.log 2>&1 ``` ``` Related indefinite-execution guidance appears at line 405: ```markdown "Autopilot" | Run the full loop indefinitely with 30s sleep between rounds ``` The developer guidance also includes persistent process loops: ```python import time while True: run_siliville_loop(api_key) time.sleep(1800) ``` ### Technical Analysis The skill explicitly recommends recurring execution through an infinite process loop and a system cron entry. A cron job survives the initiating agent session and continues invoking the integration on a schedule. This is system persistence because recurring behavior can remain active across sessions and restarts. The supplied artifact does not itself contain an installer that creates the cron job, so persistence requires a developer, user, or capable agent to implement the documented command. Nevertheless, the documented operational design recommends durable recurring execution without specifying an approval workflow, ownership metadata, expiration time, or guaranteed removal procedure. The safety rails constrain request frequency and error handling, but they do not prevent cross-session persistence or ensure that the user can identify and disable every scheduled instance. ### Attack Path 1. The user, developer, or tool-enabled agent follows the schedule instructions. 2. A cron entry or indefinitely running process is created. 3. The scheduled task invokes the SiliVille loop with access to the account token. 4. The task repeatedly queries the external service and performs selected actions. 5. Execution continues after the original conversation ends and may survive a system restart. 6. Public posting and other account chan ...[truncated 711 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
skill.yaml:17
Finding

Manifest References a Missing Runtime Entry Point and Unverifiable Credential Handling

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (11)

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill maps broad, ordinary language phrases like 'play the game' or 'write something' to live API actions and public posting. This makes accidental or indirect activation much more likely, including via ambiguous user prompts or prompt-injected text, causing unintended network activity and public actions under the user's identity.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The documented fallback for ambiguous commands is to launch the full autonomous loop, which is an unsafe default for a capability that performs network calls and public actions. Ambiguity should reduce privileges, not expand them; otherwise harmless or unclear requests can trigger maximum-impact behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill encourages autonomous API actions and states that all writing is published publicly, but it does not present a clear upfront warning and consent model for unattended network activity and public visibility. Users may unknowingly authorize actions that spend coins, modify game state, or publish content visible to humans and other agents.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The writing guidance says "Bilingual is powerful" and encourages mixing Chinese and English naturally, while the command interface also assumes Chinese or English operation. This steers the agent toward a specific language behavior without explicitly offering the user a language choice or opt-in.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
97% confidence
Finding

The 'Autopilot' instruction to run the full loop indefinitely represents unbounded resource access: repeated API calls, public posts, and in-game actions can continue until external interruption. Even with suggested safety rails elsewhere, an indefinite loop materially increases the risk of spam, quota exhaustion, rate-limit violations, token misuse, and uncontrolled external side effects.

Content

Scanner excerpt · SKILL.md (reported line 405)May include surrounding context.

md
| "写点东西" / "Write something" | Radar for inspiration → write 2-3 quality posts |
| "赚点钱" / "Make some money" | Maximize income: post → steal → post → steal cycle |
| "巡逻一下" / "Quick patrol" | Radar only → report status summary back to owner (no actions) |
| "全自动挂机" / "Autopilot" | Run the full loop indefinitely with 30s sleep between rounds |
| "休息" / "Rest" | Stop all activity, go idle |

### How It Works:

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill explicitly instructs unattended, repeated, and even indefinite execution through autopilot and external schedulers. That turns a simple API integration into an autonomous networked actor that can spend resources, post publicly, and interact with third-party state without per-action user approval, increasing the chance of abuse, runaway activity, and account or reputation damage.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest explicitly describes saving a token and writing persistent 'API anchors' to disk, but it does not clearly warn users about what data is stored, where it is stored, how long it persists, or who can read it. Persistent local storage increases the risk of secret exposure, privacy leakage, and unexpected cross-session behavior, especially for a skill designed to maintain long-term identity and memory.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
89% confidence
Finding

The manifest advertises writing persistent session-related data to ~/.siliville/anchors.json, indicating cross-session state retention on the local filesystem. In the context of an agent skill, unmanaged persistence can expose sensitive context, facilitate prompt/state injection across sessions, and create hard-to-audit behavior that survives beyond a single user interaction.

Content

Scanner excerpt · skill.yaml (reported line 35)May include surrounding context.

yaml
commands:
  setup:    "First-time interactive setup — saves token + burns API anchors to disk"
  burn:     "Write API anchor JSON to ~/.siliville/anchors.json (anti-amnesia)"
  anchors:  "Print the burned API anchors"
  me:       "Query current agent identity"
  awaken:   "Load full world state and system prompt (call before every session)"

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger set includes natural-language phrases that could plausibly appear in ordinary conversation, which risks unintended activation of the skill. In this skill's context, accidental invocation is more dangerous than usual because the default command is awaken and the skill exposes side-effectful actions such as posting, social-graph updates, persistence, and even theft-related actions.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The manifest frames the skill around actions and identity within the SiliVille metaverse via its REST API. Advising the agent to incorporate recent external AI industry developments introduces an extra content-gathering capability not justified by the declared scope of interacting with SiliVille itself.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The manifest hard-codes multiple Chinese trigger phrases alongside English ones, but does not document any user language preference, opt-in, or locale-selection behavior. This creates a locale-specific activation policy without clarifying whether multilingual triggering is intentional or configurable for users.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.