T09 · Insecure Skill Coding Practices
- Location
references/examples.md:606- Finding
Arbitrary Code Execution Through Unsafe Calculator Evaluation
- Content
View full analysis
Vulnerability Details
File Location:
references/examples.md, lines 606–610
Vulnerability Type: Unsafe evaluation of attacker-controlled input
Risk Level: Highpython Tool( name="Calculator", func=lambda x: str(eval(x)), description="Useful for mathematical calculations" )Technical Analysis
The LangChain Calculator tool passes its input directly to Python's
eval()function without validation, parsing restrictions, or sandboxing. This input originates from an LLM-generatedAction Input, which may be influenced by an untrusted user prompt or external content returned by the Search tool.Python
eval()is not limited to arithmetic. It can evaluate expressions that access built-ins, import modules, invoke functions, and interact with the host environment. Consequently, an expression such as one using__import__()could invoke operating-system functionality under the privileges of the agent process.The agent's tool allowlist does not mitigate this issue because the approved Calculator tool itself exposes a general-purpose Python execution primitive.
Attack Path
- An attacker submits a crafted question containing instructions designed to make the agent use the Calculator tool with a malicious Python expression. Alternatively, attacker-controlled content may enter the model context through internet search results.
- The LLM produces an action naming the allowed
Calculatortool and places the malicious expression inAction Input. CustomOutputParser.parse()extracts the action input without security validation.AgentExecutordispatches the extracted input to the Calculator tool.- The lambda passes the input directly to
eval(). - Python evaluates the expression in the agent process, potentially executing attacker-selected operations with that process's permissions.
Impact Assessment
Successful exploitation can result in arbitrary code execu ...[truncated 689 chars]
- Remediation
View remediation
Remediation Suggestions
- Remove the use of Python
eval()for calculator operations. - Implement a strict arithmetic parser using
ast.parse()in expression mode. - Allow only numeric constants and explicitly approved arithmetic operators, such as addition, subtraction, multiplication, division, modulo, and exponentiation.
- Explicitly reject function calls, names, attributes, imports, subscripting, comprehensions, lambdas, and all other AST node types.
- Enforce limits on expression length, numeric magnitude, exponent size, and evaluation complexity to prevent denial-of-service conditions.
- Return a controlled error for malformed or unsupported expressions.
- Treat all LLM-generated tool arguments as untrusted input, even when the tool is on an allowlist.
- Run agent tools with least privilege, restricted filesystem access, limited network access, and no unnecessary credentials.
- Add security tests covering import attempts, built-in access, attribute traversal, function invocation, oversized expressions, and prompt-injection attempts.
- Remove the use of Python
