Back to skill

Security audit

Azure Document OCR

Security checks for vulnerabilities and agentic risk

Overview

This OCR skill appears legitimate, but it needs review because it can send sensitive documents and an Azure API key to network destinations without strong safeguards.

Install only if you are comfortable with documents being processed by Azure Document Intelligence. Use a trusted HTTPS Azure endpoint, avoid sensitive ID, tax, health, or regulated documents unless you have approval, and consider hardening the scripts to validate the endpoint and polling URL before sending the API key.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/ocr_extract.py:32
Finding

Azure API Key and Document Data Can Be Sent to Untrusted Network Destinations

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
scripts/ocr_extract.py:78
Finding

Missing HTTP Timeouts Permit OCR Processing Denial of Service

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (10)

Tainted flow: 'analyze_url' from os.environ.get (line 64, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/ocr_extract.py (reported line 75)May include surrounding context.

python
if url:
        headers["Content-Type"] = "application/json"
        body = {"urlSource": url}
        response = requests.post(analyze_url, params=params, headers=headers, json=body)
    else:
        ext = Path(file_path).suffix.lower()
        content_type = CONTENT_TYPES.get(ext)

Tainted flow: 'analyze_url' from os.environ.get (line 64, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/ocr_extract.py (reported line 87)May include surrounding context.

python
headers["Content-Type"] = content_type
        with open(file_path, "rb") as f:
            body = f.read()
        response = requests.post(analyze_url, params=params, headers=headers, data=body)

    if response.status_code != 202:
        print(f"Error: Failed to submit document (HTTP {response.status_code})", file=sys.stderr)

Tainted flow: 'operation_location' from os.environ.get (line 246, credential/environment) → requests.get (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/ocr_extract.py (reported line 126)May include surrounding context.

python
print(f"Error: Polling timeout after {MAX_POLL_TIME} seconds", file=sys.stderr)
            sys.exit(1)

        response = requests.get(operation_location, headers=headers)

        if response.status_code != 200:
            print(f"Error: Failed to poll status (HTTP {response.status_code})", file=sys.stderr)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explains how to process local files and remote URLs through Azure Document Intelligence but does not warn that document contents and linked files are transmitted to Microsoft's external cloud service. In an OCR/document-processing context, inputs commonly contain sensitive personal, financial, legal, or identity data, so omission of this disclosure creates a significant privacy and compliance risk.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill advertises commands that use environment variables, shell execution, file output, and outbound network access, but it does not declare any tool scope or permissions boundaries. This can cause an agent to invoke the skill with broader capabilities than the user expects, increasing the risk of unintended file writes, credential use, or external data transmission.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The trigger list contains broad phrases like OCR, text extraction, and document parsing, which could match many generic user requests and cause this skill to activate outside a narrowly intended Azure OCR workflow. Over-broad triggering increases the chance that sensitive documents or URLs are routed to this skill and then sent to an external service without clear user intent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This reference file documents processing of highly sensitive personal data such as passports, driver's licenses, W-2 tax forms, and health insurance cards, but it provides no privacy, retention, access-control, or compliance warnings. In an OCR skill, these examples can normalize handling regulated data without safeguards, increasing the risk of insecure adoption, over-collection, or noncompliant downstream storage and sharing.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/batch_ocr.py (reported line 75)May include surrounding context.

python
]

    try:
        result = subprocess.run(
            cmd,
            capture_output=True,
            text=True,

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/ocr_extract.py (reported line 75)May include surrounding context.

python
if url:
        headers["Content-Type"] = "application/json"
        body = {"urlSource": url}
        response = requests.post(analyze_url, params=params, headers=headers, json=body)
    else:
        ext = Path(file_path).suffix.lower()
        content_type = CONTENT_TYPES.get(ext)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This script sends full document contents or document URLs to Azure for OCR, but it provides no explicit user-facing warning or consent mechanism about external data transfer. In a document-processing context, inputs may contain sensitive PII, financial, legal, or identity data, so silent transmission to a third-party cloud service creates a real privacy and compliance risk.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.