Back to skill

Security audit

oss-persona

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed Markdown-only maintainer-response style guide, with some overbroad persona triggers that users should apply carefully before public posting.

Install only if you want help drafting OSS maintainer-style public responses. Before posting, make sure the output is clearly authorized, factually verified, and not presented as a real person's own words unless that person approved it.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (5)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill advertises a very broad trigger list, including generic phrases like ‘维护者回复’, ‘人设’, and ‘对外形象’, which can overlap with ordinary writing or support tasks. That creates overbroad activation risk: the agent may invoke this persona-imitation skill in contexts where the user only wanted drafting help, causing unintended identity mimicry or misleading communications.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill hard-codes language behavior for specific personas, including rules like always replying in English regardless of the user's language and restricting response style by persona. In practice, this can override user intent and produce deceptive or manipulative communications that appear authentically authored by a particular maintainer, especially when combined with the skill’s detailed voiceprint library.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The listed triggers include phrases like “lyingbug 风格”, “WeKnora 维护者口吻”, and especially “中文排障回复”, which are broad enough to match common requests for help or drafting style rather than a narrowly scoped invocation. The file does not provide concrete negative examples or clear boundaries for when these phrases should not activate the skill.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The instruction says “中文提问 → 中文回复” and elsewhere sets Chinese as the default for a Chinese repo when uncertain, which is a language-policy choice embedded in the skill behavior. This does not appear to offer the user an explicit language choice or opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The rule states that when unsure, a Chinese repository defaults to Chinese, which is a locale/language constraint applied automatically. Because the file does not pair this with user opt-in or a clear justified regional constraint, it fits the language-policy violation category.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.