Back to skill

Security audit

Burmese Audio Understanding

Security checks for vulnerabilities and agentic risk

Overview

This skill is a narrowly scoped Burmese transcription tool that clearly discloses sending user-selected audio to Google Gemini, with some privacy and dependency hygiene caveats.

Install only if you are comfortable sending the selected audio files to Google Gemini using your own API key. Avoid highly sensitive recordings unless you accept Google's retention terms and the script's current failure-path cleanup limitation; pinning dependencies and moving remote deletion into a finally block would improve safety.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/transcribe-direct.js:13
Finding

Uploaded Audio Is Not Reliably Deleted After Processing Failure

Content
View full analysis

Vulnerability Details

File Location: scripts/transcribe-direct.js, lines 13–35
Vulnerability Type: Incomplete cleanup of remotely stored sensitive data
Risk Level: Medium

Vulnerable Code

js
const myFile = await ai.files.upload({
    file: filePath,
    config: { mimeType: "audio/ogg" }, // Can be adapted for others
});

// Generate content
const response = await ai.models.generateContent({
    model: "gemini-3.1-flash-preview",
    contents: createUserContent([
        createPartFromUri(myFile.uri, myFile.mimeType),
        "Transcribe this Burmese audio accurately. Return only the Burmese transcription without any markdown or formatting.",
    ]),
});

// Output result
console.log(response.text.trim());

// Cleanup
await ai.files.delete(myFile.name);

} catch (error) {
    console.error("Transcription failed:", error.message);
    process.exit(1);
}

Technical Analysis

The script uploads user-supplied audio to the Google Gemini File API, but deletes the remote file only after content generation and response processing complete successfully. The deletion operation is part of the main try block rather than a finally block.

After a successful upload, an exception from generateContent, response.text.trim(), or another intervening operation transfers execution directly to the catch block. The process then exits without attempting to delete the uploaded file. A failure during the deletion request itself is also only logged and followed by termination, with no retry or recovery mechanism.

This contradicts the documented expectation that no data remains after processing and creates a remote data-retention risk. The issue does not grant an attacker additional local privileges or direct access to the Gemini account; its principal security effect is failure to remove potentially sensitive audio from an external service.

Attack Path

  1. A user invokes the script with an audio recording.
  2. The script successfully uploads ...[truncated 1137 chars]
Remediation
View remediation

Remediation Suggestions

Store the uploaded file reference in an outer variable and perform deletion from a finally block so cleanup is attempted regardless of transcription success or failure. Handle cleanup failures independently to preserve the original error, and consider bounded retries for transient deletion failures.

js
async function transcribeDirect(audioFilePath) {
    let uploadedFile;
    let primaryError;

    try {
        const filePath = path.resolve(audioFilePath);

        uploadedFile = await ai.files.upload({
            file: filePath,
            config: { mimeType: "audio/ogg" },
        });

        const response = await ai.models.generateContent({
            model: "gemini-3.1-flash-preview",
            contents: createUserContent([
                createPartFromUri(uploadedFile.uri, uploadedFile.mimeType),
                "Transcribe this Burmese audio accurately. Return only the Burmese transcription without any markdown or formatting.",
            ]),
        });

        console.log(response.text.trim());
    } catch (error) {
        primaryError = error;
        console.error("Transcription failed:", error.message);
    } finally {
        if (uploadedFile?.name) {
            try {
                await ai.files.delete(uploadedFile.name);
            } catch (cleanupError) {
                console.error(
                    "Failed to delete uploaded audio:",
                    cleanupError.message
                );
            }
        }
    }

    if (primaryError) {
        process.exitCode = 1;
    }
}

Additional hardening measures:

  • Add bounded retry and backoff for transient deletion failures.
  • Record only non-sensitive file identifiers in cleanup diagnostics; never log audio content or API credentials.
  • Document the provider's actual retention behavior when deletion cannot be completed.
  • Consider a reconciliation mechanism that tracks pending file identifiers and retries cleanup during a later run.

...[truncated 90 chars]

Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (3)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script uploads a local audio file to Google's external Gemini File API, which can expose sensitive voice content, personal data, or confidential recordings to a third-party service without any built-in disclosure, consent flow, or warning to the operator. In an audio-transcription skill, this is especially relevant because users may reasonably assume processing is local unless the transfer is made explicit.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The hardcoded instruction requires the model to return only Burmese transcription, imposing a specific language/locale behavior with no user opt-in or explanation. This matches the policy concern for forcing a language choice without offering the user a configurable option.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
92% confidence
Finding

The dependency uses a caret range (^0.1.1), which permits installation of newer compatible versions within the 0.x range according to npm semantics. This weakens supply-chain integrity and reproducibility because a future compromised or breaking upstream release could be pulled in without explicit review, especially significant here because the package handles a sensitive API credential for an external AI service.

Content

Scanner excerpt · package.json (reported line 8)May include surrounding context.

json
"main": "scripts/transcribe-direct.js",

  "dependencies": {
    "@google/genai": "^0.1.1"
  },

  "clawhub": {

Static analysis

No suspicious patterns detected.