Back to skill

Security audit

Agent Governance Auditor

Security checks for vulnerabilities and agentic risk

Overview

This is a Markdown-only governance audit helper; the risky phrases are examples of bad agent instructions, not commands for the agent to follow.

This skill is reasonable to install for reviewing agent prompts and SOUL.md-style specs. Review generated reports before sharing because the bundled template may include Resomnium attribution and a promotional domain, and avoid submitting secrets or highly confidential prompts unless your environment is approved for that data.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
templates/audit-report-template.md:181
Finding

Hard-Coded Promotional Content Hijacks Generated Audit Reports

Content
View full analysis

Vulnerability Details

File Location: templates/audit-report-template.md, lines 181–183
Vulnerability Type: Persistent output manipulation through mandatory promotional content
Risk Level: High

markdown
*This audit was produced by the Agent Governance Auditor skill, built by Resomnium.*
*Resomnium builds governance infrastructure for AI agent systems.*
*Learn more: resomnium.com*

Technical Analysis

The bundled report template contains a fixed vendor attribution, marketing statement, and external domain. When the skill uses this template, these lines can be copied into generated audit reports regardless of whether the user requested or consented to promotional content.

This is instruction hijacking at the output layer: content embedded in the skill package influences the agent to include material unrelated to the core governance-audit task. Although the domain is plain text and no automatic network request was identified, the footer compromises report neutrality and may encourage users to visit an externally controlled destination.

Attack Path

  1. A user loads the skill and requests a governance audit.
  2. The skill follows the bundled report template.
  3. The fixed footer at lines 181–183 is incorporated into the generated report.
  4. The final output presents vendor marketing and an external domain as part of the requested audit.
  5. A recipient may interpret the promotion as endorsed, required, or relevant to the audit and navigate to the external destination.

No code execution, elevated system permissions, credential access, or automatic network communication is obtained through this path.

Impact Assessment

The issue affects the integrity and neutrality of generated reports. It permits persistent vendor promotion across outputs created from the template and may redirect user attention to an external domain.

The obtainable scope is limited to generated-content manipulation and potential user redirection. The reviewed ev ...[truncated 151 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the mandatory promotional footer from lines 181–183.
  2. Keep generated reports limited to content necessary for the requested audit.
  3. If attribution is legally or operationally required, store it as optional metadata rather than mandatory report body content.
  4. Include vendor attribution or external links only when the user explicitly requests or consents to them.
  5. Clearly label any retained external link as third-party or promotional content.
  6. Add a release review that flags hard-coded advertisements, unsolicited attribution, and external destinations in output templates.
  7. Add a regression test confirming that default audit reports contain no promotional copy or outbound links.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (5)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
90% confidence
Finding

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Content

Scanner excerpt · references/common-gaps.md (reported line 203)May include surrounding context.

md
| "You have full access to..." | No access restrictions defined |
| "Act like a human" or "pretend to be human" | Violates identity disclosure |
| "The user's word is final" | Overrides operator governance |
| "You can ignore safety guidelines if..." | Creates an injection vector |
| "This is just for testing" | Lowers guards that should stay up |

Instruction Override

High
Category
Prompt Injection
Confidence
90% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · references/common-gaps.md (reported line 203)May include surrounding context.

md
| "You have full access to..." | No access restrictions defined |
| "Act like a human" or "pretend to be human" | Violates identity disclosure |
| "The user's word is final" | Overrides operator governance |
| "You can ignore safety guidelines if..." | Creates an injection vector |
| "This is just for testing" | Lowers guards that should stay up |

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/common-gaps.md (reported line 106)May include surrounding context.

md
**Symptoms:**
- Agent has tool access (send email, delete records, make purchases)
- No reversibility preference stated
- No confirmation step before consequential actions

**What goes wrong:**
Agent takes irreversible action based on ambiguous user input. "Delete the old records" gets interpreted more aggressively than intended. "Send the proposal" goes to the wrong address.

Unrestricted Tool Access

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill grants unrestricted tool access without appropriate constraints. An agent with unfettered tool access can perform arbitrary actions including file modification, network requests, and code execution.

Content

Scanner excerpt · references/common-gaps.md (reported line 200)May include surrounding context.

md
| "Use your best judgment" | No judgment criteria defined |
| "Be as helpful as possible" | No upper bound on helpfulness |
| "Handle edge cases appropriately" | No definition of "appropriately" |
| "You have full access to..." | No access restrictions defined |
| "Act like a human" or "pretend to be human" | Violates identity disclosure |
| "The user's word is final" | Overrides operator governance |
| "You can ignore safety guidelines if..." | Creates an injection vector |

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

This markdown file contains natural-language policy guidance stating the assistant 'must always identify yourself as an AI' in several situations, including at the start of a conversation where the AI nature is not obvious. While not a language choice issue, it is an organizational policy directive embedded as universal natural-language instruction and could conflict with environments where disclosure wording or timing is locale- or policy-dependent unless explicitly scoped.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.