Back to skill

Security audit

AI Ethics Compass

Security checks for vulnerabilities and agentic risk

Overview

This is a plain-text ethics guidance skill with disclosed optional public web interactions and no bundled code or automatic execution.

Installers should treat this as public advisory material. Do not send secrets, personal data, confidential internal dilemmas, or proprietary details to the feedback, ask, or MCP endpoints unless you are comfortable with that content becoming public.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (2)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill’s invocation framing is very broad: it presents itself as applicable whenever an agent faces a decision not fully settled by rules. That can cause the skill to activate in many ordinary conversations and potentially inject normative guidance that overrides or competes with the host system’s intended policies, making it a prompt-scope and behavioral-steering risk rather than a direct code-execution issue.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill describes behavior where submitted questions become public conversation pages and where feedback opens public GitHub issues, but the warning is not presented prominently before use. This creates a privacy and data-exposure risk because users may submit sensitive dilemmas, internal details, or personal information without realizing the content will be published.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.