Back to skill

Security audit

Medical Triage

Security checks for vulnerabilities and agentic risk

Overview

This skill is a simple medical-message triage classifier, but it lacks guardrails for a high-stakes medical workflow.

Install only if you will use it as human-reviewed decision support. Do not let it autonomously triage, route, delay, or close patient messages, and add explicit emergency-care and clinician-review safeguards before operational use.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (1)

Missing User Warnings

High
Confidence
97% confidence
Finding
This is a true safety vulnerability because the skill performs medical urgency classification but nowhere warns that outputs are informational only and must not replace clinician judgment or emergency services. In a medical triage context, users or downstream systems may over-trust the model's categorization, causing delayed emergency care, unsafe automation, or misrouting of life-threatening messages.

Static analysis

No suspicious patterns detected.