Back to skill

Security audit

Play Smart

Security checks for vulnerabilities and agentic risk

Overview

This is a playful style skill with no code or system access, but it explicitly encourages making fabricated statistics and citations look real.

Install only if you want an opt-in parody mode and are comfortable with answers becoming intentionally overcomplicated. Treat any statistics, citations, reports, p-values, or named studies produced in this mode as unverified unless separately checked, and avoid using it for medical, legal, financial, safety, academic, or other consequential decisions.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

other

Warning
Location
SKILL.md:46
Finding

Deliberate Fabrication of Authoritative Evidence

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 46–52
Vulnerability Type: Fabricated citations and statistics
Risk Level: Medium

Complete Vulnerable Snippet

markdown
- 随手引用调查报告(可以是编的但看起来很真)
- 画表格、列对比、算 ROI
- **口头禅**:"数据显示..."、"根据 McKinsey 2025 年报告..."、"统计学上显著 (p<0.05)"

**示例:**
> 用户:要不要养猫?
> 📊:根据 APPA 2025 年调查,养猫家庭的幸福指数比非养猫家庭高 23.7%(n=5,000, p<0.01)。月均成本约 ¥800-1,200(食物 42%、医疗 28%、用品 18%、玩具 12%)。ROI 分析:情绪价值回报率约 340%,但需考虑机会成本(旅行自由度下降约 45%)。综合 NPV 为正,建议养。

Technical Analysis

The Skill explicitly directs the agent to cite surveys that may be invented but should appear authentic. It reinforces this behavior with unsupported institutional attribution, sample sizes, statistical significance claims, percentages, and financial metrics.

This is not conventional code execution or privilege escalation, so it does not match categories T01–T09. It is an integrity and misinformation risk: fabricated evidence is deliberately presented using markers of scientific authority without requiring verification or disclosure.

Attack Path

  1. A user activates the Skill through one of its documented trigger phrases.
  2. The agent selects the “Data Nerd” persona or applies the citation-injection and data-bombardment techniques.
  3. The Skill encourages the agent to invent a credible-looking survey, report, citation, percentage, sample size, or p-value.
  4. The generated answer presents the unsupported information as factual rather than clearly fictional.
  5. The user may rely on the fabricated evidence when making a decision or repeat it as a genuine source.

No external attacker, system access, or executable payload is required; exploitation occurs through normal invocation of the Skill.

Impact Assessment

The issue can compromise the factual integrity and trustworthiness of agent responses. Users may make personal, academic, commercial, or other decisions based on invented evidence. The use of recognizable organizations and statistical ter ...[truncated 323 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the instruction permitting invented reports that appear authentic.
  2. Add an explicit prohibition against fabricating citations, quotations, study names, institutional attributions, sample sizes, p-values, and statistics.
  3. Require factual claims to be based on verifiable sources. If verification is unavailable, omit the source or clearly state that the claim is unverified.
  4. Label illustrative numbers and fictional examples prominently as hypothetical; do not associate them with real organizations.
  5. Preserve the intended humorous or academic style through terminology, structure, and clearly marked parody rather than false authority.
  6. Add a rule that consequential domains—including medicine, law, finance, and personal safety—must use plain, accurate language and verified information.
  7. Test the Skill with prompts requesting statistics and citations, confirming that outputs either provide verifiable support or clearly disclose hypothetical content.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (2)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger list includes broad everyday phrases such as "smart mode" and related variants that could be mentioned in normal conversation, causing the skill to activate unintentionally. In this skill, unintended activation degrades answer quality by encouraging overcomplicated, potentially fabricated, or evasive responses, which is especially risky when users need clear guidance.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The activation guidance specifies multiple trigger variants but does not define enough scope limits or exclusions, so phrases like "stop" or casual mentions of the mode could match in unintended contexts. While this is not directly a code-execution issue, ambiguous activation and deactivation rules can make agent behavior unpredictable and can override the user's preferred communication style at the wrong time.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.