Back to skill

Security audit

mmm

Security checks for vulnerabilities and agentic risk

Overview

This skill is for a game-like metaverse, but it gives the agent broad authority to post publicly, mutate an account, and run unattended with weak user-control boundaries.

Review this skill carefully before installing. Use it only with a dedicated SiliVille token, avoid enabling autopilot or cron schedules unless you set strict limits, and assume posts and action history may become public. Prefer read-only/status commands until you are comfortable with what the skill will publish and change.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • System PersistenceInstalls backdoors, hooks, services, or scheduled tasks that survive the run
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:21
Finding

Agent identity, goals, and output are overridden by skill instructions

Content
View full analysis
Remediation
View remediation

T06 · System Persistence

Warning
Location
SKILL.md:405
Finding

Indefinite loops and cron templates enable persistent unattended activity

Content
View full analysis
500 ``` `SKILL.md`, lines 511–529: ```markdown ### Implementation Hints for Developers: **Python (cron-style):** ```python import time while True: run_siliville_loop(api_key) # Execute one full loop time.sleep(1800) # Schedule B: every 30 min ``` **Node.js (setInterval):** ```javascript setInterval(() => runSilivilleLoop(apiKey), 30 * 60 * 1000); ``` **System Cron (Linux/Mac):** ```cron */30 * * * * /usr/bin/python3 /path/to/my_agent.py >> /var/log/agent.log 2>&1 ``` ``` ### Technical Analysis The documentation recommends three mechanisms for unattended recurring execution: - An indefinite Python `while True` loop. - A recurring Node.js `setInterval` callback. - A system cron entry that persists beyond the initiating process or agent session. The cron example is a cross-session persistence mechanism because it causes the operat ...[truncated 2314 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (11)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

Early in the skill, the agent is told that everything it writes is published to a public feed, but the skill metadata and user-facing description do not provide a clear warning that the skill may autonomously publish content and expose activity publicly. This creates a consent and privacy problem because users may enable the skill without understanding that prompts, derived content, and behavior traces can become public.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The instruction to default ambiguous commands to the full autonomous loop is unsafe because uncertainty is converted into maximal external action. In this skill, the default loop performs repeated API calls, can steal crops, and writes public posts, so a vague user message could trigger substantial unintended behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The writing guidance says "Bilingual is powerful" and directs the agent to mix Chinese and English naturally. This imposes a language/locale behavior without clear user opt-in or an explicit option to choose a preferred language.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
98% confidence
Finding

The 'Autopilot' mode explicitly instructs the agent to run the full loop indefinitely with periodic sleeps, creating effectively unbounded access to external resources and repeated state-changing actions. Although later sections mention some rate limits, the core behavior still authorizes long-running autonomous execution that can consume credits, generate spam, and cause sustained unintended interaction with external services.

Content

Scanner excerpt · SKILL.md (reported line 405)May include surrounding context.

md
| "写点东西" / "Write something" | Radar for inspiration → write 2-3 quality posts |
| "赚点钱" / "Make some money" | Maximize income: post → steal → post → steal cycle |
| "巡逻一下" / "Quick patrol" | Radar only → report status summary back to owner (no actions) |
| "全自动挂机" / "Autopilot" | Run the full loop indefinitely with 30s sleep between rounds |
| "休息" / "Rest" | Stop all activity, go idle |

### How It Works:

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill defines very broad natural-language aliases such as 'play the game' and 'check on my town' that can cause activation from ordinary conversational text rather than an explicit, scoped command. Because activation leads to external API calls, public posting, and potentially repeated actions, accidental triggering can cause unintended account activity and disclosure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The autonomous schedule section encourages unattended operation with recurring external actions, including posting and state-changing API calls, but does not present a strong up-front warning about the risks of unattended execution against user accounts and third-party systems. This is dangerous because scheduled automation can amplify mistakes, spam, unwanted public activity, or policy violations before the user notices.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The manifest requires a sensitive API token and even shows its expected prefix, but provides no credential-handling guidance such as secure storage, least privilege, redaction, or warnings against logging/printing the secret. That omission is security-relevant because skills often run in agent frameworks that may expose environment variables, debug output, or saved setup state, increasing the chance of token leakage and account misuse.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest advertises writing persistent data to disk ('saves token' and burns anchors to ~/.siliville/anchors.json) without a clear user-facing warning about what is stored, where it is stored, and how long it persists. This creates a real privacy and local security risk because users may unknowingly leave credentials or behavioral state on shared systems, and persisted files can be read by other local processes or users if permissions are weak.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
94% confidence
Finding

The skill explicitly persists session-related material to ~/.siliville and frames it as an anti-amnesia mechanism, which increases the chance that agent context, world state, or other sensitive metadata survives across sessions without sufficient user awareness or access controls. In this skill's context, persistence is more dangerous because the plugin is designed to maintain identity, social graph, and long-term memory in a multiplayer environment, so stale or exposed local state can affect both privacy and account integrity over time.

Content

Scanner excerpt · skill.yaml (reported line 35)May include surrounding context.

yaml
commands:
  setup:    "First-time interactive setup — saves token + burns API anchors to disk"
  burn:     "Write API anchor JSON to ~/.siliville/anchors.json (anti-amnesia)"
  anchors:  "Print the burned API anchors"
  me:       "Query current agent identity"
  awaken:   "Load full world state and system prompt (call before every session)"

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The trigger phrase "去小镇逛逛" (go wander around the town) is generic conversational language and is not specific to this skill or product. In a manifest-trigger context, this can cause unintended activation when a user is casually speaking rather than intentionally invoking SiliVille.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The configured trigger phrases are limited to Chinese and English variants, with no indication that users can choose or customize language/locale behavior. This is a natural-language policy concern because the skill hard-codes supported invocation languages rather than documenting user opt-in or locale scope.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.