Back to skill

Security audit

MAI Transcribe

Security checks for vulnerabilities and agentic risk

Overview

This skill is a straightforward Azure Speech transcription wrapper that uploads user-selected audio to the configured Azure endpoint and writes a transcript locally.

Install only if you are comfortable sending selected audio files to Azure Speech. Set AZURE_SPEECH_ENDPOINT to your exact HTTPS Azure Speech resource, protect and rotate AZURE_SPEECH_KEY like any cloud credential, and avoid transcribing sensitive recordings unless your Azure data-handling setup permits it.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/common.js:13
Finding

Unvalidated Azure Speech Endpoint Can Expose Credentials and Audio

Content
View full analysis

Vulnerability Details

File Location: scripts/common.js:13-15 and scripts/transcribe.js:51-60
Vulnerability Type: Unrestricted transmission of secrets and sensitive audio to a configurable endpoint
Risk Level: Medium

Vulnerable Code

From scripts/common.js:13-15:

js
function apiBase() {
  return requiredEnv('AZURE_SPEECH_ENDPOINT').replace(/\/$/, '');
}

From scripts/transcribe.js:51-60:

js
const url = `${apiBase()}/speechtotext/transcriptions:transcribe?api-version=${encodeURIComponent(args['api-version'] || apiVersion())}`;
const res = await fetch(url, {
  method: 'POST',
  headers: {
    'Ocp-Apim-Subscription-Key': apiKey(),
  },
  body: form,
});

Technical Analysis

The value of AZURE_SPEECH_ENDPOINT is used directly as the request origin without parsing or validating its scheme, hostname, port, embedded credentials, or path. The resulting request includes the Azure Speech subscription key in the Ocp-Apim-Subscription-Key header and the complete audio recording in the multipart request body.

Consequently, anyone able to modify the process environment can redirect the request to an attacker-controlled server. The implementation also accepts a plain HTTP endpoint, which could expose both the credential and audio to network interception.

This issue does not independently grant an attacker control over the environment. Exploitation requires poisoning, misconfiguration, or unauthorized modification of AZURE_SPEECH_ENDPOINT.

Attack Path

  1. An attacker gains the ability to influence the environment used to launch the skill, such as through a compromised configuration, deployment secret, wrapper script, or shell environment.
  2. The attacker sets AZURE_SPEECH_ENDPOINT to an attacker-controlled HTTP or HTTPS origin.
  3. A user or agent invokes scripts/transcribe.js with an audio file.
  4. The script reads the complete audio file and constructs a multipart request.
  5. The script sends the audio and `AZUR ...[truncated 892 chars]
Remediation
View remediation

Remediation Suggestions

  1. Parse the configured endpoint with the standard URL class and reject malformed values.
  2. Require the https: protocol. Permit plain HTTP only behind a clearly named, explicit development-only opt-in.
  3. Restrict production requests to documented Azure Speech hostname suffixes or a deployment-specific trusted-host allowlist.
  4. Reject embedded usernames or passwords, unexpected ports, query strings, fragments, and unapproved base paths.
  5. Construct the API URL from validated URL components rather than concatenating an unchecked string.
  6. Consider separating support for custom endpoints behind an explicit option that warns that the subscription key and audio will be disclosed to that destination.
  7. Apply least privilege to the Azure key, rotate it if endpoint poisoning is suspected, and prefer short-lived identity-based authentication where the service supports it.
  8. Add automated tests covering attacker-controlled hosts, plain HTTP, deceptive suffixes, embedded credentials, unexpected ports, and malformed URLs.

Example validation pattern:

js
function apiBase() {
  const endpoint = new URL(requiredEnv('AZURE_SPEECH_ENDPOINT'));

  if (endpoint.protocol !== 'https:') {
    throw new Error('AZURE_SPEECH_ENDPOINT must use HTTPS');
  }

  const host = endpoint.hostname.toLowerCase();
  const allowedSuffixes = [
    '.cognitiveservices.azure.com',
    '.api.cognitive.microsoft.com',
  ];

  if (!allowedSuffixes.some(suffix => host.endsWith(suffix))) {
    throw new Error('AZURE_SPEECH_ENDPOINT is not an approved Azure Speech host');
  }

  if (
    endpoint.username ||
    endpoint.password ||
    endpoint.port ||
    endpoint.search ||
    endpoint.hash
  ) {
    throw new Error('AZURE_SPEECH_ENDPOINT contains unsupported components');
  }

  return endpoint.origin;
}

The final allowlist should be verified against the Azure Speech endpoint formats officially supported by the deployment, including any required sovereign-clou ...[truncated 31 chars]

Vulnerability Patterns
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (1)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill explicitly requires environment variables containing a cloud API key and performs outbound network requests, but it does not declare any tool scope such as permissions or allowed-tools. This weakens least-privilege controls and makes it easier for an agent runtime to expose secrets or allow unexpected external communication without explicit user/operator approval.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.