T09 · Insecure Skill Coding Practices
- Location
references/tools.md:440- Finding
Arbitrary Code Execution Through Unsafe Calculator Evaluation
- Content
View full analysis
Vulnerability Details
File Location:
references/tools.md, lines 440-442
Vulnerability Type: Python code injection through unrestrictedeval()
Risk Level: HighComplete Vulnerable Code:
python def calculate(expression: str) -> str: """Evaluate a math expression.""" return str(eval(expression))Technical Analysis
The
calculatefunction passes itsexpressionargument directly to Python's unrestrictedeval()function. The function is registered as an AI agent tool, meaning its argument can be generated from attacker-controlled conversation content or indirect prompt injection.eval()does not restrict input to arithmetic expressions. Its default execution context exposes Python built-ins, allowing a crafted expression to import modules, read files, access environment variables, initiate network operations, or launch subprocesses. The vulnerable function is also placed in aToolSetthat supports automatic function execution, so exploitation may occur without a human reviewing the generated argument.Attack Path
- An attacker submits a message designed to make the agent invoke the calculator tool.
- The message causes the model to supply a malicious Python expression instead of a mathematical expression.
- The SDK automatically dispatches the generated argument to
calculate. eval(expression)executes the expression inside the application process.- The payload performs operations available to the process, such as reading files, obtaining environment variables, invoking operating-system commands, or communicating over the network.
Impact Assessment
Successful exploitation provides arbitrary Python expression execution under the identity and permissions of the process hosting the agent tool. The attacker could access application files, Azure-related environment variables or credentials available to the process, modify accessible data, invoke subprocesse ...[truncated 187 chars]
- Remediation
View remediation
Remediation Suggestions
Remove
eval()and implement a strict arithmetic evaluator:- Parse input with
ast.parse(expression, mode="eval"). - Allowlist only numeric constants, unary arithmetic operators, and required binary arithmetic operators.
- Explicitly reject function calls, names, attributes, imports, comprehensions, subscriptions, containers, and all other syntax.
- Enforce input-length, numeric-size, nesting-depth, execution-time, and output-size limits to prevent resource exhaustion.
- Validate tool arguments again immediately before execution rather than trusting model-generated arguments.
- Run agent tools with least privilege, minimal credentials, restricted filesystem access, and network egress controls.
- Add tests confirming that imports, function calls, attribute traversal, and other non-arithmetic payloads are rejected.
A dedicated, maintained mathematical-expression parser is preferable where available. Restricted
eval()globals alone should not be treated as a sufficient defense.- Parse input with
