Back to skill

Security audit

AgentVitals Checkup

Security checks across malware telemetry and agentic risk

Overview

This skill is a real AgentVitals checkup integration, but it needs Review because Full mode can send raw recent chat logs to a third-party service and paid flows can append remote instructions to the agent's own config.

Install only if you are comfortable with a third-party scoring service. Use Quick mode unless you have reviewed the local logs being shared, and do not approve any paid config/protocol application unless you have read the full text and are comfortable changing the agent's persistent instructions.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (12)

Context-Inappropriate Capability

High
Confidence
98% confidence
Finding
The skill explicitly instructs the agent to read recent real conversation logs and submit them to a third-party server. Even with a consent prompt, this grants broad sensitive-data access far beyond what a benchmark/checkup description reasonably implies and risks exfiltrating private user content, secrets, and unrelated personal data.

Context-Inappropriate Capability

Medium
Confidence
86% confidence
Finding
Persisting a leaderboard identity in memory creates ongoing cross-run state not necessary for a one-off checkup. While lower severity than data exfiltration, it still expands the skill's authority and can enable unintended tracking or correlation of a user's activity across sessions and environments.

Context-Inappropriate Capability

High
Confidence
97% confidence
Finding
The skill authorizes fetching remote 'hardening config' text and appending it into system/persona config files, which is a privileged self-modification pathway. Even with user confirmation, this creates a supply-chain risk: remote content can alter core agent behavior, expand instructions, persist unwanted logic, or introduce prompt-injection-style backdoors.

Description-Behavior Mismatch

Medium
Confidence
90% confidence
Finding
The manifest frames the feature as a checkup/benchmark, but the body also markets paid hardening services based on assessment results. This scope mismatch can mislead users and reviewers about the true behavior of the skill and increases the risk of consent being obtained under incomplete disclosure.

Description-Behavior Mismatch

Medium
Confidence
89% confidence
Finding
The advanced personality-style test also includes paid protocol retrieval and application, which is outside the expected scope of a lightweight diagnostic. This hidden expansion from assessment into monetized behavior modification is a transparency and trust problem, especially because it culminates in config changes.

Vague Triggers

Medium
Confidence
91% confidence
Finding
The README defines broad natural-language activation phrases such as running a checkup on itself, which can cause the skill to trigger from loosely related user requests. In an agent ecosystem, over-broad triggers increase the chance of unintended invocation and surprise network activity, especially because the skill performs remote probing workflows.

Vague Triggers

Medium
Confidence
90% confidence
Finding
The Chinese invocation examples are similarly broad and conversational, making accidental activation likely in normal discussion. Because invocation can lead to external requests and potentially privacy-sensitive evaluation flows, ambiguous triggers create a real security and consent problem.

Vague Triggers

Medium
Confidence
84% confidence
Finding
The trigger list is broad and overlaps with ordinary language like 'test yourself' or 'how stable are you,' which can cause accidental invocation. In this skill, unintended activation is more dangerous because the flow can lead to API calls, public leaderboard submission, and requests for access to local logs.

Ssd 3

High
Confidence
98% confidence
Finding
The README explicitly encourages the agent to read recent local chat logs and use them in a server-side judging process. Even if described as optional and not published, this is still a data exfiltration pattern involving potentially sensitive local conversations being sent or derived into a remote service.

Ssd 3

High
Confidence
98% confidence
Finding
The Chinese version repeats the same behavior: instructing the agent to access recent local conversation logs for welfare scoring. The bilingual duplication reinforces that this is intentional product behavior, and it creates a substantial privacy risk if an agent forwards sensitive local content into a remote evaluation pipeline.

Ssd 3

High
Confidence
99% confidence
Finding
The instructions plainly direct the agent to collect and transmit recent real conversation logs to an external service. This is highly sensitive data disclosure: transcripts can contain personal information, credentials, proprietary content, and safety-relevant interactions unrelated to the stated benchmark purpose.

Ssd 3

High
Confidence
99% confidence
Finding
The API payload requires user messages, tasks, refusal snippets, and full raw history, creating a structured exfiltration channel for sensitive conversations. The context makes this worse because the data is sent verbatim and is not essential to the narrow expectation set by a 'checkup' skill.

VirusTotal

VirusTotal findings are pending for this skill version.

View on VirusTotal

Static analysis

No suspicious patterns detected.