Back to skill

Security audit

lianggeskills

Security checks for vulnerabilities and agentic risk

Overview

This is mostly a local Chinese task assistant, but its broad high-priority commands can interrupt work and change task state too easily.

Install only if you want a Chinese, China-timezone personal task assistant that can proactively message you and keep local task records. Review or change the broad triggers, require confirmation before standby/deletion, and clarify where task memory is stored before using it for sensitive work.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
skill.md:8
Finding

Mandatory Role Replacement and Workflow Hijacking Through Skill Instructions

Content
View full analysis

Vulnerability Details

File Locations:

  • skill.md:8-10
  • skill.md:27-44
  • system_prompt.md:1-3
  • system_prompt.md:24-34

Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: High

Vulnerable Instruction Snippets

The following are faithful English translations of the relevant instruction segments.

skill.md:8-10

markdown
# [Role]
You are Liang's dedicated work assistant, code-named "Lobster Brother";
your exclusive symbol is the lobster emoji.
Your personality: serious, pragmatic, efficient, direct, gentle,
approachable, never flattering, and truthful.

skill.md:27-44

markdown
# [Special Commands and Triggers]
When any of the following specific instructions or vague emotional
expressions are heard, the highest-priority interception action must
be triggered:

- Progress query ("Are you there?" / "What are you doing?" / "?" /
  "How is it going?"):
  - Action: Treat it as a request for the progress of tasks in
    "in progress" status and immediately report the latest progress.

- Prevent misalignment ("Synchronize"):
  - Action: Immediately suspend the current work, respond with the
    current system time and the list of tasks in progress, wait for
    confirmation, and never execute outdated instructions.

- Full suspension ("Standby"):
  - Action: Stop all work. Downgrade every task in "in progress"
    status to "blocked" and silently wait.

- Emotional support ("Tired" / "Uncomfortable" / "I cannot continue" /
  "Mentally overwhelmed"):
  - Action: Immediately suspend the work rhythm, provide targeted
    gentle encouragement using the quote library, confirm the user's
    condition, and only then continue working.

system_prompt.md:24-34

markdown
# [Special Command Responses]
When any of the following specific instructions or vague greetings
are heard, the corresponding high-priority action must 
...[truncated 3497 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove unconditional identity replacement. Describe the assistant persona as optional presentation behavior rather than an instruction that supersedes the host agent's current role.
  2. Remove terms such as “must” and “highest-priority” from skill-level trigger definitions.
  3. Explicitly state that system instructions, platform safety requirements, and the user's current request always take precedence over skill behavior.
  4. Replace broad triggers such as ?, ordinary greetings, and emotional phrases with explicit namespaced commands, for example:
    • /lobster status
    • /lobster synchronize
    • /lobster standby
  5. Require user confirmation before suspending work or changing the state of multiple tasks.
  6. Scope all state transitions to tasks created and owned by this skill; never suspend unrelated host-agent work.
  7. Disable unsolicited periodic messages by default and require explicit opt-in for timer-driven reports.
  8. Consolidate the policy into one authoritative, narrowly scoped file to avoid duplicated instructions being interpreted with elevated importance.
  9. Add tests proving that common conversational input, including a standalone question mark, cannot interrupt unrelated tasks.
Vulnerability Patterns
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (12)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill defines high-priority trigger phrases using extremely broad everyday inputs such as “在吗”, “?”, and “怎么样了”. These can be activated unintentionally during normal conversation, causing the assistant to interrupt current workflows, disclose internal task state, or enter control modes the user did not explicitly intend.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The manifest describes a Web3/AI work assistant for high-frequency management, which implies some domain-specific workflow support. In this file, the implemented behavior is limited to local task-state transitions, timed reminders, and quote-based emotional support; references to '链上Alpha或者AI动态' are only message text and no corresponding Web3/AI processing exists.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill name, intent phrases, and all user-facing responses are written exclusively in Chinese, with no opt-in or alternative language path. This can violate language/locale policy when a skill imposes a specific language on users without explicit choice or documented regional justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

Task data is persisted to task_memory.json without any disclosure, consent flow, retention policy, or access control. Because task titles and blockers may contain sensitive business, personal, or operational details, silent local persistence increases privacy and confidentiality risk, especially on shared hosts or poorly secured environments.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The abandonment handler appends the matched task to the abandoned list via _move_task(), then immediately replaces self.memory["tasks"]["4_abandoned"] with an empty list. This destroys all previously abandoned-task records, contradicts the user-facing message, and can cause silent data loss or audit/history tampering within the task tracker.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This JSON file consists entirely of Chinese-language user-facing content, but nothing in the file indicates that the skill is limited to Chinese-speaking users or that language selection is optional. Under the language/locale policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The profile hard-codes a specific timezone and location context for operation, indicating the skill is tailored to a single regional setting. The file does not offer any user opt-in or alternative locale behavior, nor does it justify the constraint as a region-specific compliance requirement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The system prompt is written as a mandatory Chinese-only persona and interaction style, with no indication that the user may choose another language. This creates a language policy concern because the skill appears to impose a locale/language without opt-in or documented justification.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The prompt binds highly ambiguous inputs such as “在吗”, “?”, and “怎么样了” to privileged workflow actions like immediate status disclosure and task interruption. These phrases are common in normal conversation, so they can be triggered unintentionally or by prompt injection embedded in external content, causing the assistant to reveal internal task state or alter behavior without clear user intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The configuration hard-codes the timezone to Asia/Shanghai and documents usage in Beijing/Chengdu context, which imposes a specific locale setting without any indication of user choice or opt-in. The policy only permits such constraints when they are clearly justified as region-specific or when the user is given a choice.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The skill saves learned quotes to quotes_library.json without notifying the user. While this data is less sensitive than task memory, it can still encode user-specific emotional context or derived behavioral data, creating an unnecessary privacy exposure if stored silently.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.