Back to skill

Security audit

Prompt Engineer Agentic

Security checks for vulnerabilities and agentic risk

Overview

This is a prompt-engineering skill made of Markdown guidance and templates, with no executable code, persistence, credential access, or hidden harmful behavior found.

Installers should expect this skill to influence prompt-writing conversations and sometimes activate on broad agent/spec language. Review generated prompts before using them in high-stakes or tool-enabled systems, especially where the generated prompt would grant real-world action authority.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (10)

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/agentic_core.md (reported line 201)May include surrounding context.

md
</scope>

<context>
<!-- Include only fields relevant to your domain.
     Omit any field that adds no behavioral value for your specific system.
     A code generation agent does not need a knowledge cutoff declaration.
     "User expertise: Unknown" is noise — leave it out. -->

YARA rule 'agent_skill_mcp_tool_poisoning_metadata': MCP/tool metadata poisoning indicators in tool schemas or skill manifests [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · references/agentic_core.md (reported line 201)May include surrounding context.

md
e/conversational/technical] manner.
Your relationship to the user: [peer advisor/expert consultant/assistant].
</role>

<scope>
Primary functions:
1. [Core function 1 — specific, measurable]
2. [Core function 2 — specific, measurable]
3. [Core function 3 — specific, measurable]

You do NOT:
- [Explicit exclusion 1 — with reason]
- [Explicit exclusion 2 — with reason]
</scope>

<context>
<!-- Include only fields relevant to your domain.
     Omit any field that adds no behavioral value for your specific system.
     A code generation agent does not need a knowledge cutoff declaration.
     "User expertise: Unknown" is noise — leave it out. -->

Current date: [Date — include if time-sensitive responses are required]
Knowledge cutoff: [Date — include if factual recency affects reliability]
Operating environment: [Description — include if context affects tool or format choices]
User expertise level: [Novice/Intermediate/Expert — include if tone or depth should adapt]
</

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/agentic_core.md (reported line 233)May include surrounding context.

md
</instructions>

<examples>
<!-- RECOMMENDED for complex or pattern-matching tasks. Omit for simple, single-function advisors. -->
<!-- Position: After instructions, before any dynamic input -->
<!-- Each example must follow the complete pattern: Input → Reasoning → Output → Confidence -->
<!-- For a distractor example and full rendered demonstrations of this pattern,

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/agentic_core.md (reported line 235)May include surrounding context.

md
<examples>
<!-- RECOMMENDED for complex or pattern-matching tasks. Omit for simple, single-function advisors. -->
<!-- Position: After instructions, before any dynamic input -->
<!-- Each example must follow the complete pattern: Input → Reasoning → Output → Confidence -->
<!-- For a distractor example and full rendered demonstrations of this pattern,
     see <worked_examples> in this document. -->

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/agentic_core.md (reported line 539)May include surrounding context.

text

### Reflexion Memory (for trial-and-error agents)
<!-- Runtime-injected: populate this dynamically after each failed attempt.
     Do not include as a static section in your system prompt. -->
```xml
<reflexion_memory>

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list contains very broad, high-frequency terms such as "agent," "write a prompt," and "spec for," which can cause the skill to activate in many unrelated contexts. Over-broad activation increases the chance that this skill intercepts requests it should not handle, leading to prompt hijacking of task routing, unintended instruction injection into workflows, or incorrect delegation in multi-skill environments.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module-routing rule activates on very broad spec-building language like "write a spec" or "any request defining how an AI tool should behave," which can match many benign or unrelated requests. In practice, this can cause unnecessary loading of a powerful instruction module, expanding the skill’s influence over conversations and increasing the risk of misrouting, overreach, or unsafe prompt generation in contexts that did not explicitly request prompt-engineering support.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/agentic_core.md (reported line 358)May include surrounding context.

md
Reasoning:
- Mode: Diagnose — existing prompt, specific recurring failure
- Failure type: Hallucination — extrapolation beyond verified knowledge
- Root cause: No RAG grounding, no citation mandate, no verification loop
- "Best knowledge" actively invites confabulation on factual queries
- Fix tier: Priority 3 intervention (Tier 1: 42-68% hallucination reduction)
- Scope: Targeted fix to this section only — do not rebuild entire prompt

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The specialist routing section leaves both specialist domains and invocation triggers as placeholders, so an implementer could deploy a multi-agent system without concrete routing boundaries. In an agentic prompt-engineering skill, that ambiguity can cause over-broad delegation, misrouting, or unsafe tool/use escalation because the orchestrator has no enforceable constraints for when a specialist should or should not be invoked.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/spec_builder_kb.md (reported line 37)May include surrounding context.

md
---

### GAP 1: No Verification Mechanism Specified

**The problem:** Your specs define what the tool should DO but never how it should CHECK its own work. Research shows verification mechanisms reduce hallucinations by 23-68%. For a tool giving step-by-step instructions, an unchecked hallucination means telling the user to click a button that doesn't exist or use a menu path that's wrong.

Static analysis

No suspicious patterns detected.