Back to skill

Security audit

Azure Ai Agents Py - Microsoft Foundry

Security checks for vulnerabilities and agentic risk

Overview

The skill is a documentation-only Azure AI Agents SDK guide, but it includes unsafe agent-tool examples that could lead users to build automatic function execution around unrestricted Python eval.

Review this skill carefully before using it to generate production agent code. The Azure SDK coverage is useful and mostly purpose-aligned, but do not copy the calculator examples that use eval; replace them with a restricted parser or a safe math library, validate all tool arguments, and require explicit approval for tools that can modify data, call external APIs, spend money, or access sensitive resources. Also verify which Azure identity DefaultAzureCredential will use and avoid uploading sensitive files unless you intend them to be sent to Azure services.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
references/tools.md:440
Finding

Arbitrary Code Execution Through Unsafe Calculator Evaluation

Content
View full analysis

Vulnerability Details

File Location: references/tools.md, lines 440-442
Vulnerability Type: Python code injection through unrestricted eval()
Risk Level: High

Complete Vulnerable Code:

python
def calculate(expression: str) -> str:
    """Evaluate a math expression."""
    return str(eval(expression))

Technical Analysis

The calculate function passes its expression argument directly to Python's unrestricted eval() function. The function is registered as an AI agent tool, meaning its argument can be generated from attacker-controlled conversation content or indirect prompt injection.

eval() does not restrict input to arithmetic expressions. Its default execution context exposes Python built-ins, allowing a crafted expression to import modules, read files, access environment variables, initiate network operations, or launch subprocesses. The vulnerable function is also placed in a ToolSet that supports automatic function execution, so exploitation may occur without a human reviewing the generated argument.

Attack Path

  1. An attacker submits a message designed to make the agent invoke the calculator tool.
  2. The message causes the model to supply a malicious Python expression instead of a mathematical expression.
  3. The SDK automatically dispatches the generated argument to calculate.
  4. eval(expression) executes the expression inside the application process.
  5. The payload performs operations available to the process, such as reading files, obtaining environment variables, invoking operating-system commands, or communicating over the network.

Impact Assessment

Successful exploitation provides arbitrary Python expression execution under the identity and permissions of the process hosting the agent tool. The attacker could access application files, Azure-related environment variables or credentials available to the process, modify accessible data, invoke subprocesse ...[truncated 187 chars]

Remediation
View remediation

Remediation Suggestions

Remove eval() and implement a strict arithmetic evaluator:

  1. Parse input with ast.parse(expression, mode="eval").
  2. Allowlist only numeric constants, unary arithmetic operators, and required binary arithmetic operators.
  3. Explicitly reject function calls, names, attributes, imports, comprehensions, subscriptions, containers, and all other syntax.
  4. Enforce input-length, numeric-size, nesting-depth, execution-time, and output-size limits to prevent resource exhaustion.
  5. Validate tool arguments again immediately before execution rather than trusting model-generated arguments.
  6. Run agent tools with least privilege, minimal credentials, restricted filesystem access, and network egress controls.
  7. Add tests confirming that imports, function calls, attribute traversal, and other non-arithmetic payloads are rejected.

A dedicated, maintained mathematical-expression parser is preferable where available. Restricted eval() globals alone should not be treated as a sufficient defense.

T09 · Insecure Skill Coding Practices

Error
Location
references/async-patterns.md:440
Finding

Arbitrary Code Execution Through Unsafe Async Calculator Evaluation

Content
View full analysis

Vulnerability Details

File Location: references/async-patterns.md, lines 440-445
Vulnerability Type: Python code injection through unrestricted eval()
Risk Level: High

Complete Vulnerable Code:

python
def calculate(expression: str) -> str:
    """Evaluate a math expression."""
    try:
        return str(eval(expression))
    except Exception as e:
        return f"Error: {e}"

Technical Analysis

This asynchronous agent example exposes a calculator function through FunctionTool and evaluates its model-supplied argument with unrestricted Python eval(). The surrounding exception handler only converts errors to text; it does not validate the expression, limit available built-ins, or prevent successful malicious code from executing.

Because the function is added to a ToolSet and supplied to the streaming run, model-generated tool calls can automatically reach the vulnerable function. User prompts and untrusted content processed by the agent can therefore cross the boundary from natural-language input into local code execution.

Attack Path

  1. An attacker provides direct or indirect prompt content that instructs the agent to use the calculator.
  2. The model generates a calculator argument containing executable Python rather than arithmetic.
  3. Automatic function handling invokes calculate with that argument.
  4. eval() executes the malicious expression in the agent application's process.
  5. Any successful side effects occur before the function returns; the try and except block does not prevent them.

Impact Assessment

Exploitation can result in arbitrary Python expression execution with the host application's privileges. Potential consequences include reading or changing accessible files, exposing environment variables and credentials, starting subprocesses, making outbound requests, and compromising data processed by the agent. If the process has cloud credentia ...[truncated 91 chars]

Remediation
View remediation

Remediation Suggestions

Replace eval() with a parser that recognizes only the intended arithmetic grammar:

  1. Use an AST-based evaluator or a dedicated arithmetic-expression library.
  2. Allow only numeric literals and explicitly approved arithmetic operators.
  3. Reject calls, identifiers, attributes, imports, indexing, comprehensions, and unknown AST nodes.
  4. Impose limits on expression length, nesting, exponent size, execution duration, and output size.
  5. Validate all model-generated tool arguments before dispatch.
  6. Do not return raw exception details to untrusted users, because they may disclose implementation information.
  7. Isolate tool execution using least-privilege credentials, filesystem restrictions, and outbound-network controls.
  8. Add negative security tests for code-execution and resource-exhaustion payloads.

T08 · Insecure Dependencies

Note
Location
SKILL.md:14
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 14-16
Vulnerability Type: Uncontrolled dependency resolution
Risk Level: Low

Complete Vulnerable Code:

bash
pip install azure-ai-agents azure-identity
# Or with azure-ai-projects for additional features
pip install azure-ai-projects azure-identity

Technical Analysis

The installation instructions specify package names without exact versions, a lock file, or artifact hashes. Each installation can therefore resolve to different package versions depending on when it is performed and which package index is configured.

The package names are consistent with the Skill's declared Azure SDK purpose, and the reviewed content does not identify a malicious package or unsafe custom registry. Nevertheless, unconstrained resolution weakens reproducibility and supply-chain assurance. A compromised future release, an unexpectedly incompatible update, or an untrusted configured index could introduce vulnerable or unauthorized code.

Attack Path

  1. A user follows the documented pip install command.
  2. pip resolves the latest versions available from its configured indexes.
  3. A compromised, vulnerable, or unexpectedly changed release is selected.
  4. The package is installed and later imported by the application.
  5. The dependency executes with the application's permissions during import or normal use.

Impact Assessment

The impact depends on the behavior of the resolved dependency and the privileges of the installation or application process. A compromised dependency could execute code, access application data and credentials, alter files, or communicate over the network. This finding does not establish that the named Azure packages are malicious; it identifies the absence of controls that would ensure reviewed and reproducible dependency versions.

Remediation
View remediation

Remediation Suggestions

  1. Pin each dependency to a reviewed exact version.
  2. Maintain a generated lock file for transitive dependencies.
  3. Require package hashes, such as through pip install --require-hashes.
  4. Use an explicitly trusted package index or an internally controlled package mirror.
  5. Review dependency provenance and signatures where supported.
  6. Run automated vulnerability and license scanning against the resolved dependency graph.
  7. Update dependencies through a controlled review process rather than automatically consuming the latest release.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (11)

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 134)May include surrounding context.

md
toolset=toolset,
)

# Process run - toolset auto-executes functions
run = client.runs.create_and_process(
    thread_id=thread.id,
    agent_id=agent.id,

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/acceptance-criteria.md (reported line 633)May include surrounding context.

md
toolset=toolset,
)

# Process run - toolset auto-executes functions
run = client.runs.create_and_process(
    thread_id=thread.id,
    agent_id=agent.id,

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/tools.md (reported line 198)May include surrounding context.

md
toolset=toolset,
)

# Process run - toolset auto-executes functions
run = client.runs.create_and_process(
    thread_id=thread.id,
    agent_id=agent.id,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This markdown file includes examples that upload local files via client.files.upload_and_poll and create vector stores from those files, which can transmit user data off-system. The surrounding prose presents these operations as routine workflow patterns but does not warn readers about privacy or data-handling implications.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The example exposes a FunctionTool that evaluates arbitrary Python expressions using eval on model-influenced input. In an agent context, user prompts can steer the model to pass attacker-controlled strings into the function, enabling arbitrary code execution in the host process rather than mere arithmetic evaluation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The example explicitly passes a toolset to streaming with a comment indicating it enables automatic function execution, but it provides no warning about the trust boundary or the risks of model-triggered tool invocation. In an agent SDK context, this can normalize unsafe patterns where LLM output causes local functions or external actions to run without explicit confirmation, increasing the chance of unintended side effects or abuse when developers copy the example into real systems.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

This markdown file documents File Search uploads, Bing web search, and later API-connected tools, but it does not include any warning that uploaded documents or prompts may be transmitted to external services or indexed remotely. Under the markdown-specific SQP-2 criteria, descriptions of behaviors that can affect user data or privacy should disclose those implications.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The example function advertises mathematical evaluation but uses Python eval() on untrusted input, which allows arbitrary code execution rather than restricted arithmetic. In an agent/tooling context, users may copy this pattern directly into callable tools, turning model- or user-supplied strings into executable code with access to the host process.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This markdown file includes explicit deletion examples for vector stores, uploaded files, and agents, but it does not warn users that these operations remove resources and may be irreversible. Under the markdown criteria, skills should disclose behaviors that can affect user data or system integrity.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The file presents agents_client.files.save(...) as a normal pattern but does not mention that it writes files to the local filesystem and may persist generated content. For markdown guidance, user-visible warnings are expected when behavior can affect local data handling.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The examples access PROJECT_ENDPOINT and MODEL_DEPLOYMENT_NAME, and use DefaultAzureCredential, which commonly relies on sensitive local or cloud identity context. In markdown guidance, there is no note reminding users to protect credentials or verify which identity/account will be used before running the examples.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.