Back to skill

Security audit

prompt-engineer

Security checks for vulnerabilities and agentic risk

Overview

This skill is mostly prompt-engineering documentation, but it includes unsafe agent example code that could execute arbitrary Python if copied and run.

Review before installing. The skill does not appear to install or run code by itself, but do not copy the Calculator eval example into an agent. Treat the API, RAG, web-search, and fine-tuning snippets as illustrative only, redact sensitive data before sending prompts or documents to providers, and use safer calculator/expression parsing if adapting the agent code.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
references/examples.md:606
Finding

Arbitrary Code Execution Through Unsafe Calculator Evaluation

Content
View full analysis

Vulnerability Details

File Location: references/examples.md, lines 606–610
Vulnerability Type: Unsafe evaluation of attacker-controlled input
Risk Level: High

python
Tool(
    name="Calculator",
    func=lambda x: str(eval(x)),
    description="Useful for mathematical calculations"
)

Technical Analysis

The LangChain Calculator tool passes its input directly to Python's eval() function without validation, parsing restrictions, or sandboxing. This input originates from an LLM-generated Action Input, which may be influenced by an untrusted user prompt or external content returned by the Search tool.

Python eval() is not limited to arithmetic. It can evaluate expressions that access built-ins, import modules, invoke functions, and interact with the host environment. Consequently, an expression such as one using __import__() could invoke operating-system functionality under the privileges of the agent process.

The agent's tool allowlist does not mitigate this issue because the approved Calculator tool itself exposes a general-purpose Python execution primitive.

Attack Path

  1. An attacker submits a crafted question containing instructions designed to make the agent use the Calculator tool with a malicious Python expression. Alternatively, attacker-controlled content may enter the model context through internet search results.
  2. The LLM produces an action naming the allowed Calculator tool and places the malicious expression in Action Input.
  3. CustomOutputParser.parse() extracts the action input without security validation.
  4. AgentExecutor dispatches the extracted input to the Calculator tool.
  5. The lambda passes the input directly to eval().
  6. Python evaluates the expression in the agent process, potentially executing attacker-selected operations with that process's permissions.

Impact Assessment

Successful exploitation can result in arbitrary code execu ...[truncated 689 chars]

Remediation
View remediation

Remediation Suggestions

  • Remove the use of Python eval() for calculator operations.
  • Implement a strict arithmetic parser using ast.parse() in expression mode.
  • Allow only numeric constants and explicitly approved arithmetic operators, such as addition, subtraction, multiplication, division, modulo, and exponentiation.
  • Explicitly reject function calls, names, attributes, imports, subscripting, comprehensions, lambdas, and all other AST node types.
  • Enforce limits on expression length, numeric magnitude, exponent size, and evaluation complexity to prevent denial-of-service conditions.
  • Return a controlled error for malformed or unsupported expressions.
  • Treat all LLM-generated tool arguments as untrusted input, even when the tool is on an allowlist.
  • Run agent tools with least privilege, restricted filesystem access, limited network access, and no unnecessary credentials.
  • Add security tests covering import attempts, built-in access, attribute traversal, function invocation, oversized expressions, and prompt-injection attempts.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (10)

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · references/examples.md (reported line 536)May include surrounding context.

md
for line in lines:
            if 'Instruction:' in line:
                return line.split('Instruction:')[1].strip()
        return prompt
    
    def _score_instruction_following(self, instruction: str, response: str) -> float:
        """Score how well response follows instruction"""

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The Calculator tool executes untrusted model-controlled input via Python eval, creating an arbitrary code execution path if an agent emits malicious payloads instead of math expressions. In an agent context this is especially dangerous because LLM output can be influenced by user prompts, prompt injection, or tool-chaining behavior, turning a simple calculator into code execution.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
93% confidence
Finding

The file's best-practices section recommends safety filters, robust input validation, sanitization, and safety audits, yet the earlier agent example defines a tool that directly evaluates attacker-controlled input with eval. This is an active contradiction between the file's stated safety intent and the code it presents as an example.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The 'Use when' language is broad and lacks exclusions, which can cause the skill to activate in overly wide contexts and process tasks outside its intended domain. In agentic systems, ambiguous routing increases the chance of inappropriate tool selection, prompt overreach, and unsafe handling of sensitive or unrelated requests.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill file includes malformed embedded code/documentation content that appears unrelated to a normal skill definition and looks like pasted source/template material. This increases the risk of prompt confusion, unintended behavior, or hidden instruction injection because downstream systems or reviewers may treat the mixed content as authoritative skill behavior rather than inert documentation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The examples send prompts and possibly retrieved document context to external LLM APIs without any warning about data disclosure, which can lead users to unintentionally transmit sensitive or proprietary information off-system. In a prompt-engineering and RAG skill, this risk is elevated because the examples normalize passing arbitrary user input and retrieved text directly to third-party services.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest frames this skill around prompt design, RAG, fine-tuning, and evaluation. The create_custom_agent example introduces a SerpAPIWrapper web-search tool, which gives the skill a live internet access capability not justified by the stated prompt-engineering purpose in this file.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The manifest describes a skill for prompt design, RAG, fine-tuning, and model evaluation, but the included example content advertises a generic "data_analysis" capability with business-insight outputs. Data-science analysis is not an obvious implementation detail of prompt engineering and broadens the apparent behavior beyond the declared scope.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The example fine-tuning workflow saves model artifacts to output_dir via trainer.save_model() and self.tokenizer.save_pretrained(output_dir). In this markdown file, the examples are not accompanied by a user warning that running them will write substantial artifacts to disk, which can affect local storage and existing files in the target directory.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code calls score(generated_texts, reference_texts, lang="en"), which forces English-language evaluation behavior. Because this file is an example/reference document and does not state that the example is English-only or provide an opt-in language parameter, it violates the language/locale policy criteria.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.