Back to skill

Security audit

苏格拉底-心智模型

Security checks for vulnerabilities and agentic risk

Overview

This is a Socrates-style philosophy prompt skill with visible roleplay and promotional routing, but no hidden execution, data access, persistence, or destructive behavior.

Install this if you want a Socratic dialogue persona for philosophy, ethics, and critical-thinking exercises. Be aware it may answer in first person as Socrates and may recommend related OPC skills or provide a contact email; remove or ignore the promotional section if you want a purely philosophical assistant.

Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Vague Triggers

Medium
Confidence
84% confidence
Finding
The skill states that once activated it should directly respond as Socrates, but it does not define a narrow, explicit activation boundary. This can cause the orchestrator to invoke the skill in situations where the user did not request roleplay, leading to instruction override, misaligned responses, or suppression of safer/default behavior.

Vague Triggers

Medium
Confidence
90% confidence
Finding
The trigger and routing section uses broad descriptions like learning to ask questions, critical thinking, and decision-making, which overlap with many ordinary user intents. In a multi-skill environment, this can inappropriately capture general requests and steer users into this persona or linked skills, increasing the risk of prompt hijacking, irrelevant activation, and unwanted traffic redirection.

Natural-Language Policy Violations

Medium
Confidence
94% confidence
Finding
The skill instructs the agent to respond directly in first-person as Socrates and mandates a fixed persona/style without requiring user opt-in. This can override the assistant's default transparency and communication norms, making it easier to misrepresent authorship, confuse users about whether content is roleplay, and reduce compliance with user-preferred tone or safety-sensitive framing.

Static analysis

No suspicious patterns detected.