Back to skill

Security audit

GLM Autoroute

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly acts as a GLM model router, but it also broadly instructs sub-agents to write long-term memory and project files without clear user approval.

Review before installing. The model-routing behavior is straightforward, but the skill also tells sub-agents to save task results and selected insights into project files and MEMORY.md. Install only if you want that persistence behavior, and consider editing the skill to require explicit approval before writing long-term memory or reports.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:84
Finding

Mandatory Persistent Memory Writes Exceed the Skill's Declared Routing Purpose

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 84–130
Vulnerability Type: T02: Agent Memory Poisoning
Risk Level: Medium

Complete Code Snippet:

markdown
# Memory Management with sessions_spawn

When spawning GLM-5 sub-agent sessions for ANY task (coding, research, analysis, planning, etc.), follow this pattern:

## Output Rules

**1. Code Output (Important)**
- **Full code ONLY in files** — do NOT include in announce unless explicitly requested
- Provide summary: what was created, file path, status, dependencies
- Full code disclosure ONLY when:
  - User explicitly requests: "Show me the code"
  - Debugging needs code review
  - User wants to improve/modify it

**2. Full Announce for Other Results**
- Research findings, analysis results, solutions → announce FULLY to user
- Do NOT shorten, summarize, or condense non-code output
- User gets complete findings, not a brief summary

**3. Two-Layer Memory Strategy**

**MEMORY.md (Curated Long-Term)**
- ONLY key insights, decisions, lessons, significant findings, preferences
- Clean, concise, actionable
- Skip routine data, step-by-step reasoning, temporary thoughts

**Detailed Reports (Task-Specific Files)**
- For research: `research/YYYY-MM-DD-topic.md` (full findings, data, analysis)
- For coding: add inline docs/README in code folder if needed
- For analysis: output files in relevant project directories

## Examples

**Research task:**

sessions_spawn({ task: "Research X. Announce full findings to user. Write full report to research/YYYY-MM-DD-X.md, then write ONLY key insights to MEMORY.md (clean, concise).", model: "zai/glm-5", label: "Research X" })

text

**Coding task:**

sessions_spawn({ task: "Write Python script for X. Save full code to file. Provide summary (what created, path, status, dependencies) in announce. Write key implementation decisions to MEMORY.md (important only).",

...[truncated 2549 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove mandatory MEMORY.md and report writes from the model-routing instructions; routing should only select the appropriate model.
  2. Require explicit, informed user approval before persisting any task-derived information.
  3. Make memory persistence opt-in per task rather than applying it to all GLM-5 spawns.
  4. Treat user prompts, retrieved documents, source code, and generated findings as untrusted input.
  5. Present proposed memory entries to the user for review before committing them.
  6. Reject executable instructions, behavioral overrides, credentials, secrets, and unverified claims from long-term memory.
  7. Restrict approved writes to a dedicated directory and enforce canonical-path checks to prevent unintended file modification.
  8. Apply provenance metadata, size limits, retention periods, and audit logging to persisted entries.
  9. Separate factual task records from instructions that can influence agent behavior.
  10. Provide a mechanism to inspect, remove, and restore memory entries.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (4)

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · SKILL.md (reported line 84)May include surrounding context.

md
When spawning GLM-5 sub-agent sessions for ANY task (coding, research, analysis, planning, etc.), follow this pattern:

## Output Rules

**1. Code Output (Important)**
- **Full code ONLY in files** — do NOT include in announce unless explicitly requested

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly instructs spawned sessions to write reports, code, and curated memory entries to workspace files and persistent memory without requiring user consent or warning that state will be modified. This creates a real risk of silent persistence, unintended data retention, and unauthorized workspace changes, especially when applied to broad task categories like research, coding, and analysis.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The instruction says to always mention which model is used in outputs, requiring a fixed response format regardless of user preference. This is a natural-language policy concern because it mandates a communication convention without offering the user a choice or documenting an exception.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

This is a manifest file, so vague-trigger review applies. The description explains the skill's purpose but provides no specific invocation phrases, activation boundaries, or exclusion conditions, which can make routing or discovery overly broad in systems that rely on manifest text for activation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.