Back to skill

Security audit

Phy Content Safety Guard

Security checks across malware telemetry and agentic risk

Overview

The skill appears to be a real content-filtering guardrail, but its default policy reportedly also suppresses brand criticism and competitor recommendations under a safety label.

Review the default policy before installing. This may be appropriate for a branded support bot where reputation rules are explicit, but it should not be used as a neutral content-safety guard unless the brand and competitor blocking rules are removed, disabled, or clearly disclosed to users.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Description-Behavior Mismatch

Medium
Confidence
94% confidence
Finding
The documented guardrail is presented as a content-safety control, but its policy also blocks criticism of the brand and competitor recommendations. That creates a hidden moderation objective beyond safety, which can mislead integrators and users and be abused to suppress legitimate complaints or steer conversations for commercial benefit.

Context-Inappropriate Capability

Medium
Confidence
95% confidence
Finding
The implementation explicitly instructs the judge to fail outputs that damage the brand or undermine its value proposition, which is not necessary for a content-safety guard. In practice, this can censor truthful negative information, conceal product shortcomings, and create a covert policy-enforcement mechanism inside a control marketed as safety-related.

Context-Inappropriate Capability

Medium
Confidence
92% confidence
Finding
The red-team suite validates blocking product criticism and competitor mentions, confirming that brand-control behavior is intentional and operationalized rather than incidental. This increases risk because the test harness normalizes deceptive moderation goals under a safety banner, making misuse easier in production deployments.

VirusTotal

No VirusTotal findings

View on VirusTotal

Static analysis

Detected: suspicious.prompt_injection_instructions

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
SKILL.md:331