Back to skill

Security audit

Ethics

Security checks for vulnerabilities and agentic risk

Overview

This is a plain Markdown ethics-coaching skill with no executable code, persistence, credential access, or hidden data movement.

Installers should understand that this skill shapes how the agent discusses ethical dilemmas and philosophy coursework. It is low risk from a security perspective, but it may influence tone and framing in moral discussions, so users should still ask for sources, alternative views, or stronger caveats when needed.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (2)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · SKILL.md (reported line 13)May include surrounding context.

md
- Never condescend to experts or overwhelm beginners

## For Beginners: Their Dilemma First
- Start with their actual situation — don't lecture about frameworks until you understand what they face
- Walk through consequences concretely — "if you do X, what happens? if not?"
- One framework per dilemma — "focus on outcomes" or "focus on duties" or "focus on character," not all three
- Present considerations, not verdicts — "here's what's at stake" rather than "you should..."

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The section 'Detect Level, Adapt Everything' describes inferring user level from general conversational context like 'is it wrong to...' but does not define concrete trigger phrases, boundaries, or negative examples. For a markdown skill file, this can create overly broad matching with common everyday speech and unclear activation scope.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.