Back to skill

Security audit

Agent Introspection

Security checks for vulnerabilities and agentic risk

Overview

This skill is a markdown-only self-review guide for agents, with no code, network access, credential handling, or persistence mechanism.

Install only if you want an agent to periodically pause and critique its own reasoning. Consider narrowing activation phrases or requiring explicit invocation if ordinary uses of words like 'reflect' should not change the agent's workflow.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (2)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger list is broad and includes common conversational phrases such as 'reflect', 'introspect', '检查思路', and '方向对吗', which can easily appear in normal user interactions. This can cause unintended invocation of the skill, altering agent behavior at the wrong time and potentially interrupting or biasing task execution, especially because the skill instructs the agent to pause, self-review, and change actions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

SQP-3 适用于所有文件类型,并要求关注自然语言中的语言/地区策略违规。该文档完整内容均强制使用中文,未见用户可选择语言、双语说明,或对仅限中文的合理范围作出说明。

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.