Back to skill

Security audit

Verify Claims

Security checks across malware telemetry and agentic risk

Overview

This fact-checking skill is mostly coherent, but it can automatically create delayed follow-up checks and retain the user's original claim/context without clear opt-in or retention limits.

Install only if you are comfortable with web-based fact-checking and search queries leaving the chat. For sensitive political, medical, legal, personal, or workplace claims, do not allow scheduled follow-ups unless the agent first asks permission and explains what will be stored, for how long, and how to cancel or delete it.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (6)

Context-Inappropriate Capability

Medium
Confidence
92% confidence
Finding
The skill instructs the agent to schedule a future task and notify the user later, which goes beyond the declared purpose of one-shot fact-checking. This creates unscoped delayed actions and persistence behavior that may surprise users, especially if the platform does not clearly surface background tasks, consent, or notification behavior.

Context-Inappropriate Capability

Medium
Confidence
88% confidence
Finding
The skill broadens from curated professional fact-checking into general web research and preliminary analysis, materially expanding its operational scope. That expansion can lead the agent to use lower-assurance sources than the user expects, increasing the risk of misinformation, unsafe tool use patterns, or policy bypass via a seemingly narrow fact-checking skill.

Intent-Code Divergence

Medium
Confidence
85% confidence
Finding
The documentation states that the skill focuses on professional fact-checking rather than original research, yet elsewhere directs fallback general web research and preliminary analysis. This internal contradiction weakens policy boundaries and makes it easier for an agent to exceed intended scope under ambiguous conditions.

Vague Triggers

Medium
Confidence
77% confidence
Finding
The manifest uses broad trigger phrases such as 'verify this' and 'is this true,' which overlap with ordinary conversation and can cause the skill to activate in situations beyond the user's intended request. Over-broad activation increases the chance of unnecessary tool use, unexpected browsing, and unintended application of the skill's expanded behaviors.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The skill describes automatically scheduling a follow-up and notifying the user later without requiring upfront notice and consent for persistence or delayed actions. This is dangerous because it can create background behavior the user did not knowingly authorize, with privacy and trust implications even if the task itself is low-risk.

Missing User Warnings

Medium
Confidence
95% confidence
Finding
The skill instructs storing the original query, claims, and context for later scheduled runs without a clear disclosure of retention, purpose, or duration. Retaining user-provided content for future processing without transparent consent is a privacy and data-governance risk, especially when the content may include sensitive topics or personally identifying details.

VirusTotal

66/66 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.