Back to skill

Security audit

Calculator

Security checks for vulnerabilities and agentic risk

Overview

This is a straightforward calculator skill, with a real but limited risk that very large expressions could consume local CPU or memory.

This skill is reasonable for local calculator use. Do not expose it directly to untrusted users or automated public inputs without adding expression length, complexity, numeric magnitude, timeout, and memory limits.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/calculator.py:39
Finding

Unbounded Mathematical Expression Evaluation Enables Resource Exhaustion

Content
View full analysis

Vulnerability Details

File Location: scripts/calculator.py, lines 39-43 and 66-74
Vulnerability Type: Uncontrolled resource consumption through unrestricted mathematical operations
Risk Level: Medium

Vulnerable Code:

python
        'pow': pow,
        'max': max,
        'min': min,
        'pi': math.pi,
        'e': math.e,
python
        'factorial': math.factorial,
        'gcd': math.gcd,
python
        code = compile(expr, '<string>', 'eval')
        
        # Check for disallowed names
        for name in code.co_names:
            if name not in allowed_names:
                return {"error": f"Unknown function or variable: {name}"}
        
        result = eval(code, {"__builtins__": {}}, allowed_names)

Technical Analysis

The evaluator restricts accessible names and removes built-ins, which mitigates straightforward arbitrary code execution. However, it compiles and evaluates a user-controlled expression without enforcing limits on expression length, integer magnitude, exponent size, factorial arguments, nesting depth, execution time, memory consumption, or result size.

Allowed operations such as pow, factorial, exponentiation, and arbitrary-precision integer arithmetic can require excessive CPU time or memory. Name allowlisting does not prevent this class of attack because the resource-intensive operations are intentionally exposed. Exception handling only applies after Python raises an exception; it does not provide a time limit or protect the process from operating-system termination caused by memory exhaustion.

Attack Path

  1. An attacker submits a computationally expensive calculator expression, such as an extremely large factorial or integer exponent.
  2. The Agent invokes the documented command:
    bash
    python3 scripts/calculator.py calc "<attacker-controlled expression>"
    
  3. compile() acce ...[truncated 1041 chars]
Remediation
View remediation

Remediation Suggestions

  1. Replace direct evaluation of compiled input with an AST-based expression interpreter that explicitly permits only required numeric literals, operators, and function calls.
  2. Reject expressions exceeding a small maximum input length and enforce limits on AST node count and nesting depth.
  3. Limit integer literal length, numeric magnitude, exponent values, and factorial arguments before performing calculations.
  4. Reject non-finite values and cap the magnitude and serialized size of results.
  5. Execute calculations in an isolated worker process with strict wall-clock timeout, CPU, and memory limits. Terminate the worker when a limit is exceeded.
  6. Add tests for denial-of-service inputs, including very large factorials, exponents, deeply nested expressions, oversized literals, and expressions with many operations.
  7. Return a generic limit-related JSON error when an input exceeds an enforced constraint rather than attempting the calculation.
Vulnerability Patterns
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (2)

eval() call detected

High
Category
Dangerous Code Execution
Confidence
85% confidence
Finding

Using eval() on user-controlled input is inherently risky, even with builtins removed and a name allowlist. This implementation meaningfully reduces classic code-injection risk, but it still exposes the process to denial-of-service via computationally expensive expressions such as huge exponentiation, deeply nested operations, or expensive math calls that can consume CPU or memory.

Content

Scanner excerpt · scripts/calculator.py (reported line 71)May include surrounding context.

python
if name not in allowed_names:
                return {"error": f"Unknown function or variable: {name}"}
        
        result = eval(code, {"__builtins__": {}}, allowed_names)
        
        # Format result
        if isinstance(result, (int, float)):

compile() call detected

Medium
Category
Dangerous Code Execution
Confidence
65% confidence
Finding

compile() creates code objects from strings. When combined with exec()/eval(), it enables obfuscated code execution.

Content

Scanner excerpt · scripts/calculator.py (reported line 64)May include surrounding context.

python
try:
        # Compile and evaluate safely
        code = compile(expr, '<string>', 'eval')
        
        # Check for disallowed names
        for name in code.co_names:

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/calculator.py:71