Back to skill

Security audit

Prompt Design Tuning

Security checks for vulnerabilities and agentic risk

Overview

This is an instruction-only prompt-tuning workflow skill with disclosed model-evaluation behavior and human approval gates, but users should control data sharing and spending before execution.

Install only if you want a structured, agent-driven prompt-tuning workflow. Before using execution mode, confirm that the evaluation set can be shared with the selected model providers, redact secrets or sensitive data, use scoped test API keys, set budget and rate limits, and review generated scripts before running them.

Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (1)

Vague Triggers

Medium
Confidence
90% confidence
Finding
The trigger phrase "Help me automate this prompt tuning workflow" is broad enough to match ordinary prompt-help requests and can cause the skill to activate outside its intended niche. Because this skill is execution-oriented and explicitly guides batch generation, evaluation loops, script writing, and model comparisons, accidental invocation could lead an agent to take actions that are more powerful or costly than the user intended.

Static analysis

No suspicious patterns detected.