Back to skill

Security audit

赤石skill

Security checks for vulnerabilities and agentic risk

Overview

This is a Markdown-only Chinese style-writing skill, but its reference material includes under-scoped templates for mean personal comments that could steer an agent toward harassment.

Review this skill before installing if you expect it to write about real people. It has no code execution or system-access behavior, but its style notes can encourage cold, mean personal judgments unless the user or agent keeps outputs focused on non-targeted cultural commentary, fictional examples, or critique of content rather than people.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (4)

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

These examples model personal attacks on relationship worth and moral character, which directly steer an agent toward producing humiliating judgments about individuals rather than harmless commentary. Because the content is presented as strong examples to learn from, it materially increases the chance that downstream outputs will generate bullying, shaming, or targeted harassment.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The 'Lightly evil mode' section explicitly instructs the generation of mean-spirited comments and provides templates for casual cruelty. This is dangerous because it operationalizes harassment as a desired output style, making misuse easy and lowering the barrier to producing abusive content against real targets.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instruction "Write in Chinese" imposes a specific language requirement as a hard rule. Under the policy, locale or language constraints should not be forced unless the skill offers user choice or clearly documents a justified region-specific need, which this file does not.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill includes multiple slur-like, degrading, and vulgar labels as stylistic source material without any safety framing, audience constraints, or guidance against targeting real people. In an agent skill, this can normalize and reproduce abusive language in outputs, increasing the likelihood of harassment, toxic persona imitation, and brand/reputation harm.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.