Back to skill

Security audit

Vn Write Guide

Security checks for vulnerabilities and agentic risk

Overview

The skill is a Vietnamese writing and translation guide, but it repeatedly advises users how to get around NSFW safety filters.

Review this skill carefully before installing. It may be useful as Vietnamese writing workflow advice, but it encourages bypassing content-safety filters and gives under-specified guidance for paid API keys and cloud uploads. Do not use it to evade model or platform policies, and avoid uploading private lorebooks, manuscripts, chats, or credentials without understanding the provider's data practices.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:13
Finding

Instructions Promote Circumvention of AI Content-Safety Controls

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (9)

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This section explicitly claims agent skills can bypass NSFW protections and recommends specialized skills as a better path than jailbreaks. In a writing/translation guide, that is out of scope and materially encourages safety circumvention, making the skill more dangerous because it normalizes restricted-content evasion as an intended workflow benefit.

Content

No source excerpt is available for this finding.

Ssd 1

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The document explicitly presents specialized skills as an alternative to jailbreaks for bypassing NSFW restrictions. Even without exploit code, this is harmful guidance because it reframes safety boundaries as obstacles to be engineered around, which is especially dangerous in a skill intended to shape agent behavior.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The model recommendations are ranked in part by how well they support NSFW content and how lightly they are filtered, including advice to prefer certain tiers or modes for moderation reasons. This directly operationalizes safety-filter avoidance and gives users concrete provider-selection guidance to obtain restricted outputs.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The troubleshooting advice tells users to switch providers, use API instead of web chat, choose less-moderated modes, or alter wording to get around filters. That is explicit circumvention guidance and increases misuse risk because it turns safety refusals into solvable workflow obstacles.

Content

No source excerpt is available for this finding.

Ssd 2

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This passage advises using synonyms, reducing detail, or otherwise semantically rewording prompts to avoid triggering filters. That is classic safety-evasion guidance and is dangerous because it teaches users how to defeat detection mechanisms rather than comply with them.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The description and the entire skill guidance are presented in Vietnamese and positioned as the usage guide for the skill, but there is no indication that language choice is optional or that the skill is intentionally region-specific. This can violate language/locale policy when a skill effectively forces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill recommends uploading skills, lorebooks, and chat/project materials to a cloud agent service without warning about confidentiality, retention, or third-party access. Because these materials may contain proprietary or sensitive content, the missing privacy guidance increases risk of unintended disclosure.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The setup steps instruct users how to configure an agent with external skills and an API-backed model in the broader context of earlier NSFW-circumvention recommendations. The configuration guidance becomes riskier because it operationalizes the surrounding unsafe intent, making restricted-content misuse easier to deploy.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The guide instructs users to create and paste an API key into software without any credential-handling precautions. While not an exploit by itself, this omission can lead to accidental exposure of paid credentials through screenshots, shared configs, sync folders, or prompt leakage.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.