Back to skill

Security audit

glitchward-llm-shield

Security checks for vulnerabilities and agentic risk

Overview

This skill is a clearly disclosed prompt-scanning integration that sends selected prompt content to Glitchward's external API, with appropriate warnings about sensitive data.

Install only if you are comfortable sending the prompts you scan to Glitchward's service. Do not use it on secrets, credentials, personal data, regulated data, or proprietary conversations unless your organization has approved that data flow and retention model.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (7)

Instruction Override

High
Category
Prompt Injection
Confidence
90% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · README.md (reported line 92)May include surrounding context.

md
"severity": "critical",
      "pattern": "ignore_previous_rules",
      "matched_text": "...",
      "description": "Detected attempt to override system behavior"
    }
  ],
  "request_id": "abc-123"

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The README explicitly demonstrates sending prompt content to an external service, which is a real data egress behavior rather than a false positive. In this skill’s context, that transmission is the core feature, but it still creates confidentiality and compliance risk if users scan secrets, PII, or proprietary prompts, since prompt contents leave the local environment.

Content

Scanner excerpt · README.md (reported line 63)May include surrounding context.

bash
PROMPT_TEXT="What is the capital of France?"
echo "$PROMPT_TEXT" | jq -Rs '{prompt: .}' | \
  curl -s -X POST "https://glitchward.com/api/shield/validate" \
    -H "X-Shield-Token: $GLITCHWARD_SHIELD_TOKEN" \
    -H "Content-Type: application/json" \
    -d @- | jq .

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

This shell example transmits user prompt text and the API token to Glitchward's external validation endpoint. The use of jq reduces command-injection risk, but it does not change the core issue that potentially sensitive prompt content is exported to a third-party service, which can violate confidentiality expectations or policy.

Content

Scanner excerpt · SKILL.md (reported line 59)May include surrounding context.

bash
PROMPT_TEXT="user input goes here"
echo "$PROMPT_TEXT" | jq -Rs '{prompt: .}' | \
  curl -s -X POST "https://glitchward.com/api/shield/validate" \
    -H "X-Shield-Token: $GLITCHWARD_SHIELD_TOKEN" \
    -H "Content-Type: application/json" \
    -d @- | jq .

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

This messages-format example sends a full conversation payload to an external API, potentially including more context than a single prompt and thereby increasing confidentiality exposure. Because conversation histories often contain system prompts, tool outputs, or prior user data, the skill context makes this transmission more sensitive than a generic outbound HTTP request.

Content

Scanner excerpt · SKILL.md (reported line 70)May include surrounding context.

bash
PROMPT_TEXT="user input goes here"
echo "$PROMPT_TEXT" | jq -Rs '{messages: [{role: "user", content: .}]}' | \
  curl -s -X POST "https://glitchward.com/api/shield/validate" \
    -H "X-Shield-Token: $GLITCHWARD_SHIELD_TOKEN" \
    -H "Content-Type: application/json" \
    -d @- | jq .

External Transmission

Medium
Category
Data Exfiltration
Confidence
96% confidence
Finding

The batch-validation example exports multiple prompts in one request to a third-party service, multiplying the amount of potentially sensitive data disclosed in a single operation. Batch handling increases blast radius because one misuse can leak several user inputs or workflow artifacts at once.

Content

Scanner excerpt · SKILL.md (reported line 93)May include surrounding context.

md
PROMPT2="second user input"
jq -n --arg p1 "$PROMPT1" --arg p2 "$PROMPT2" \
  '{items: [{prompt: $p1}, {prompt: $p2}]}' | \
  curl -s -X POST "https://glitchward.com/api/shield/validate/batch" \
    -H "X-Shield-Token: $GLITCHWARD_SHIELD_TOKEN" \
    -H "Content-Type: application/json" \
    -d @- | jq .

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The Python example posts prompt content to an external endpoint, which is a genuine external transmission risk even though the HTTP client usage avoids shell injection. The surrounding skill context makes this more dangerous because the advertised use case is scanning arbitrary user and workflow content before LLM calls, which may include sensitive intermediate data unless strictly controlled.

Content

Scanner excerpt · SKILL.md (reported line 119)May include surrounding context.

md
token = os.environ.get("GLITCHWARD_SHIELD_TOKEN")
prompt = "user input goes here"  # Safe: passed as data, not shell syntax

response = requests.post(
    "https://glitchward.com/api/shield/validate",
    headers={
        "X-Shield-Token": token,

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

The Python example posts prompt content to an external endpoint, which is a genuine external transmission risk even though the HTTP client usage avoids shell injection. The surrounding skill context makes this more dangerous because the advertised use case is scanning arbitrary user and workflow content before LLM calls, which may include sensitive intermediate data unless strictly controlled.

Content

Scanner excerpt · SKILL.md (reported line 119)May include surrounding context.

md
token = os.environ.get("GLITCHWARD_SHIELD_TOKEN")
prompt = "user input goes here"  # Safe: passed as data, not shell syntax

response = requests.post(
    "https://glitchward.com/api/shield/validate",
    headers={
        "X-Shield-Token": token,

Static analysis

No suspicious patterns detected.