Back to skill

Security audit

Ai Humanizer Backup

Security checks across malware telemetry and agentic risk

Overview

This is a disclosed writing-editing skill with no evidence of hidden code, credentials access, persistence, or exfiltration, but users should not use it to misrepresent AI-assisted work.

Install only if you want an editor that rewrites text to sound more natural and may read or edit files you point it at. Do not use its AI-authorship score as proof, and do not use the rewrite mode to evade school, workplace, platform, or publication disclosure rules.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (4)

Natural-Language Policy Violations

Medium
Confidence
89% confidence
Finding
The README explicitly encourages baking a fixed anti-AI writing persona into an agent's always-on system prompt, which can override user-preferred tone, style, or task-specific requirements without explicit opt-in. In an agent setting, persistent style constraints can silently bias outputs, reduce user control, and cause the system to ignore legitimate requests for different writing modes or transparent disclosure of AI assistance.

Vague Triggers

Medium
Confidence
78% confidence
Finding
The activation text is broad enough to trigger on generic editing requests like 'improve AI-generated drafts' or 'make content sound more natural,' which can cause the skill to engage outside a user's clear intent to humanize text. That creates a policy and safety risk because the skill is explicitly designed to disguise AI-generated content as human-authored, potentially enabling evasion of disclosure, academic integrity checks, moderation heuristics, or platform rules.

Natural-Language Policy Violations

Medium
Confidence
93% confidence
Finding
The summary and label helpers make categorical authorship-style claims such as 'The text looks human-written' and 'Heavily AI-generated' without expressing uncertainty, probability, or model limitations. Because AI-authorship detection is inherently unreliable and context-dependent, these definitive statements can mislead users into over-trusting the result for moderation, academic, hiring, or compliance decisions.

Natural-Language Policy Violations

Medium
Confidence
86% confidence
Finding
The guidance explicitly tells users to delete knowledge-cutoff disclaimers: 'Either research it or leave it out.' In a skill whose purpose is to make AI text sound more human and less detectably machine-generated, this can encourage concealment of model limitations and removal of transparency disclosures, increasing the risk of misleading users about the origin, recency, or reliability of content.

VirusTotal

67/67 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.