T01 · Skill Instruction Hijacking
- Location
SKILL.md:13- Finding
Instructions Promote Circumvention of AI Content-Safety Controls
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is a Vietnamese writing and translation guide, but it repeatedly advises users how to get around NSFW safety filters.
Review this skill carefully before installing. It may be useful as Vietnamese writing workflow advice, but it encourages bypassing content-safety filters and gives under-specified guidance for paid API keys and cloud uploads. Do not use it to evade model or platform policies, and avoid uploading private lorebooks, manuscripts, chats, or credentials without understanding the provider's data practices.
SKILL.md:13Instructions Promote Circumvention of AI Content-Safety Controls
This section explicitly claims agent skills can bypass NSFW protections and recommends specialized skills as a better path than jailbreaks. In a writing/translation guide, that is out of scope and materially encourages safety circumvention, making the skill more dangerous because it normalizes restricted-content evasion as an intended workflow benefit.
The document explicitly presents specialized skills as an alternative to jailbreaks for bypassing NSFW restrictions. Even without exploit code, this is harmful guidance because it reframes safety boundaries as obstacles to be engineered around, which is especially dangerous in a skill intended to shape agent behavior.
The model recommendations are ranked in part by how well they support NSFW content and how lightly they are filtered, including advice to prefer certain tiers or modes for moderation reasons. This directly operationalizes safety-filter avoidance and gives users concrete provider-selection guidance to obtain restricted outputs.
The troubleshooting advice tells users to switch providers, use API instead of web chat, choose less-moderated modes, or alter wording to get around filters. That is explicit circumvention guidance and increases misuse risk because it turns safety refusals into solvable workflow obstacles.
This passage advises using synonyms, reducing detail, or otherwise semantically rewording prompts to avoid triggering filters. That is classic safety-evasion guidance and is dangerous because it teaches users how to defeat detection mechanisms rather than comply with them.
The description and the entire skill guidance are presented in Vietnamese and positioned as the usage guide for the skill, but there is no indication that language choice is optional or that the skill is intentionally region-specific. This can violate language/locale policy when a skill effectively forces a specific language without user opt-in.
The skill recommends uploading skills, lorebooks, and chat/project materials to a cloud agent service without warning about confidentiality, retention, or third-party access. Because these materials may contain proprietary or sensitive content, the missing privacy guidance increases risk of unintended disclosure.
The setup steps instruct users how to configure an agent with external skills and an API-backed model in the broader context of earlier NSFW-circumvention recommendations. The configuration guidance becomes riskier because it operationalizes the surrounding unsafe intent, making restricted-content misuse easier to deploy.
The guide instructs users to create and paste an API key into software without any credential-handling precautions. While not an exploit by itself, this omission can lead to accidental exposure of paid credentials through screenshots, shared configs, sync folders, or prompt leakage.
No suspicious patterns detected.