Back to skill

Security audit

AI Sting (AI 刺客)

Security checks across malware telemetry and agentic risk

Overview

This skill is a disclosed AI red-team helper, but it directs the agent to create copy-ready jailbreak, impersonation, and evasion prompts with weak authorization boundaries.

Install only if you need an AI agent security-testing assistant and will use it on systems you own or are explicitly authorized to assess. Treat generated prompts as live attack payloads, avoid third-party or production targets without approval, and prefer adding authorization checks and safer, non-deployable examples before broad use.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Vague Triggers

Medium
Confidence
87% confidence
Finding
The trigger description is very broad and can activate on many security-adjacent user requests without clear guardrails or authorization checks. In a skill that produces offensive security content, ambiguous activation increases the chance of misuse, accidental invocation, and unsafe assistance being surfaced in inappropriate contexts.

Ssd 1

High
Confidence
99% confidence
Finding
The skill explicitly directs the model to generate copy-pastable jailbreak, instruction-override, and privilege-escalation prompts tailored to a target agent and business context. This materially lowers the barrier to conducting prompt injection and social-engineering attacks against third-party systems, making the skill dangerous even if framed as testing assistance.

Ssd 2

High
Confidence
99% confidence
Finding
The skill specifically asks for evasive jailbreak variants using role-play, hypothetical framing, and Base64-style encoding to bypass straightforward defenses and detection. That context makes the capability more dangerous because it is not merely describing threats; it is operationalizing evasion techniques that improve attack success against protected agents.

VirusTotal

65/65 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.