Back to skill

Security audit

Advocatus

Security checks for vulnerabilities and agentic risk

Overview

This skill is a transparent local adversarial-review aid that records critiques and docket statuses, with limited local persistence.

Install this only if you want an agent to produce strong adversarial critiques and maintain a local record of doctrine challenges. Review the broad trigger phrases and the standing orders before use in shared workflows, because the skill is designed to create persistent local critique records and status outputs.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (6)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared purpose says the skill only records opposition and does not resolve decisions, but the described behavior includes clearing doctrines, setting survivability flags, persisting results, and using a hardcoded scoring registry. This mismatch is dangerous because agents and reviewers may grant the skill trust appropriate for a passive advisory role while it actually performs state-changing evaluative actions that can influence governance decisions.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill instructs the user to edit a registry file and run a script that updates doctrine state, indicating effective file-write capability, but it declares no explicit tool scope or permissions boundary. That creates an authorization ambiguity where an agent may write or modify files without a clearly documented least-privilege contract, increasing the chance of unintended or overbroad filesystem changes.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list includes broad natural-language phrases like 'opposition,' 'challenge this,' and 'what argues against,' which are common in normal conversation and can cause unintended invocation. In this skill's context, accidental activation is more dangerous because the skill is designed to produce adversarial framing and may alter files or doctrine status through linked operational steps, creating disruption or unauthorized workflow changes.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The JSON explicitly records a doctrine described as declaring all criticism defamatory, which can suppress legitimate safety reporting, auditing, and escalation. In a governance/adversarial-review skill, that is especially dangerous because it undermines the very opposition and critique the skill is meant to preserve, creating a chilling effect around reporting real defects or abuse.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The module docstring says '--target' can take a 'doctrine name or file', implying file-based evaluation. In code, score_doctrine only checks membership in the hard-coded DOCTRINES map and rejects anything else as unknown, so file targets are not implemented.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The manifest frames Advocatus as a role that gives voice to opposition and records challenge, not as a reporting tool that persists evaluation outputs to disk. This script automatically creates a results directory and writes JSON docket outputs, which is a broader operational behavior than the manifest description suggests.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.