Back to skill

Security audit

Photo Caption Tool Free

Security checks across malware telemetry and agentic risk

Overview

This captioning skill is mostly a normal text-generation aid, but it requests shell execution even though its own documentation says no command execution is needed.

Install only if you are comfortable with a captioning skill that currently asks for shell execution permission it does not appear to need. Prefer a version with exec removed and triggers narrowed to photo-caption requests; specify the desired output language in your prompt.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Intent-Code Divergence

Medium
Confidence
95% confidence
Finding
The skill describes itself as pure Markdown that does not require command execution, yet the manifest allows the exec tool. This creates an unnecessary capability expansion that could let an agent run shell commands despite users and reviewers being told no execution is needed, increasing the risk of misuse or prompt-induced command execution.

Vague Triggers

Medium
Confidence
82% confidence
Finding
The trigger condition says the skill should be used for broad tasks like marketing copy, writing content, title optimization, and content creation, which extend beyond photo-caption generation. Overbroad activation criteria can cause the agent to invoke this skill in unrelated contexts, leading to inappropriate scope expansion and increasing the chance that unsafe or unintended instructions are followed under this skill's looser framing.

Natural-Language Policy Violations

Medium
Confidence
78% confidence
Finding
The skill states that caption language is chosen by default based on location and user language, with domestic locations tending toward Chinese and overseas locations toward English, without requiring explicit user confirmation. This can infer or act on sensitive contextual attributes and may produce outputs in an unintended language, which is a privacy and user-control issue even in a low-risk captioning skill.

VirusTotal

VirusTotal findings are pending for this skill version.

View on VirusTotal

Static analysis

No suspicious patterns detected.