Back to skill

Security audit

gemma4-dev

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed Gemma 4 development guide with examples; its main risks are developer-controlled model/cloud use and visible reasoning-trace demos, not hidden or malicious behavior.

Installers should treat this as a developer reference skill. Do not expose the Gradio reasoning inspector to end users or shared environments without access controls, avoid sending sensitive prompts to Vertex AI unless your cloud policy allows it, and pin model/package versions when adapting the examples.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
references/multimodal.md:59
Finding

Mutable Remote Model Artifacts Without Immutable Revision Pinning

Content
View full analysis

Vulnerability Details

File Location: references/multimodal.md, lines 59–64
Vulnerability Type: Unpinned third-party model and processor artifacts
Risk Level: Medium

python
MODEL_ID = "google/gemma-4-12B-it"

processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalCausalLM.from_pretrained(
    MODEL_ID,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

Technical Analysis

The example loads a processor and model from a mutable Hugging Face repository without supplying an immutable revision commit hash. Consequently, two executions of the same example can retrieve different repository contents.

This conflicts with the artifact-pinning guidance in SKILL.md, which instructs users to pin remote Hugging Face artifacts and disable remote code execution. Although the shown calls do not explicitly enable trust_remote_code, mutable model, tokenizer, configuration, and serialization artifacts still create a supply-chain and reproducibility risk.

The reviewed code does not establish that arbitrary code execution is immediately possible. The direct consequences supported by the code are the ingestion of changed model artifacts, altered inference behavior, resource exhaustion from unexpectedly changed artifacts, and loss of build reproducibility. More severe effects would depend on the file formats accepted by the installed Transformers version and its artifact-loading behavior.

Attack Path

  1. An attacker compromises the referenced model repository or obtains authority to change its default branch or revision.
  2. The attacker replaces or modifies model, processor, tokenizer, or configuration artifacts.
  3. A developer follows the documented example without specifying a commit revision.
  4. from_pretrained() resolves and downloads the changed repository contents.
  5. The application loads the unreviewed artifacts, potentially causing manipulated inference outpu ...[truncated 640 chars]
Remediation
View remediation

Remediation Suggestions

  • Pin both from_pretrained() calls to the same verified immutable commit:

    python
    MODEL_ID = "google/gemma-4-12B-it"
    MODEL_REVISION = "VERIFIED_FULL_COMMIT_HASH"
    
    processor = AutoProcessor.from_pretrained(
        MODEL_ID,
        revision=MODEL_REVISION,
        trust_remote_code=False,
    )
    model = AutoModelForMultimodalCausalLM.from_pretrained(
        MODEL_ID,
        revision=MODEL_REVISION,
        trust_remote_code=False,
        torch_dtype=torch.bfloat16,
        device_map="auto",
    )
    
  • Verify that the selected commit belongs to the intended publisher and document how it was reviewed.

  • Prefer safe serialization formats such as Safetensors and reject legacy executable serialization formats where supported.

  • For sensitive deployments, download artifacts in a controlled build stage, verify cryptographic hashes, scan them, and load from a read-only local directory with offline mode enabled.

  • Add an automated documentation or static-analysis check that rejects remote from_pretrained() calls lacking an immutable revision.

T08 · Insecure Dependencies

Note
Location
examples/vertex_ai_deployment.py:57
Finding

Unpinned Package Installation Command in Error Guidance

Content
View full analysis

Vulnerability Details

File Location: examples/vertex_ai_deployment.py, line 57
Vulnerability Type: Unpinned third-party package installation
Risk Level: Low

python
except ImportError:
    raise ImportError("google-cloud-aiplatform is required for Vertex AI deployment. Run: pip install google-cloud-aiplatform")

Technical Analysis

The error message recommends installing google-cloud-aiplatform without a version constraint, lockfile, or package hash. Following this instruction causes pip to resolve the latest available release and its mutable transitive dependency graph.

The package name is plausible, and installation is neither automatic nor silently executed. Therefore, this is not evidence of malicious package use. The issue is that the guidance does not provide a reproducible, reviewed dependency set and exposes users to future compromised, incompatible, or unexpectedly changed releases.

Attack Path

  1. An attacker compromises a future release of the named package or one of its transitive dependencies, or a future release otherwise becomes unsafe.
  2. A user runs the example without the dependency installed.
  3. The example displays the unpinned pip install google-cloud-aiplatform instruction.
  4. The user executes that command.
  5. Pip resolves and installs the then-current package and dependency versions.
  6. Installed package code subsequently executes during import or use inside the user's Python environment.

Impact Assessment

The potential scope is the Python environment and the operating-system permissions of the user performing the installation. A malicious dependency could theoretically access files, environment variables, cloud credentials available to the process, or network resources under those permissions.

That impact is conditional on a future upstream or transitive dependency compromise; no malicious dependency, active compromise, automatic installation, privilege ...[truncated 65 chars]

Remediation
View remediation

Remediation Suggestions

  • Replace the floating installation instruction with a reviewed version constraint, for example:

    python
    except ImportError as exc:
        raise ImportError(
            "Install the audited dependencies from the project's locked requirements file."
        ) from exc
    
  • Provide a committed lockfile or constraints file containing exact versions of google-cloud-aiplatform and relevant transitive dependencies.

  • For high-assurance deployments, publish and require hashes with pip install --require-hashes.

  • Install dependencies inside an isolated virtual environment or container rather than a privileged system Python environment.

  • Use automated dependency scanning and a controlled update process to review new package versions before changing the lockfile.

Vulnerability Patterns
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a broad application-design and implementation aid for the Gemma 4 family, covering several areas: model architecture selection, Thinking Mode formatting, function calling, multimodal pipelines, and deployment references. The supplied code, however, is narrowly focused on a single Gradio demo for inspecting streamed tags from model output. While it does relate to Gemma 4 and Thinking Mode, it does not implement or address most of the stated scope, especially architecture selection, function calling schemas, or multimodal audio/vision handling. Its primary purpose is materially narrower and different: a reasoning-trace inspection UI. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a broad implementation/design skill for Gemma 4 applications, covering architecture choice, thinking mode formatting, function calling, multimodal processing, and deployment. The supplied code does not implement or expose those capabilities. Instead, it is a test file specifically validating a stream parser's handling of tags and ensuring internal reasoning text does not leak into public responses. While this loosely relates to 'Thinking Mode formatting,' the primary purpose is materially narrower and different from the declared scope, so this is a mismatch.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The parsing and UI flow are purpose-built to collect, separate, and reveal internal reasoning traces back to the operator. In skill context, this makes the issue more dangerous because it operationalizes chain-of-thought disclosure as a product feature, increasing the chance of leaking secrets, unsafe hidden content, prompt artifacts, or other non-user-facing internals.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The example is explicitly built to extract and display content inside blocks, which can expose internal reasoning, sensitive intermediate text, or hidden policy/process artifacts to the operator. Even if intended as a developer demo, surfacing chain-of-thought is an information disclosure risk because model internals may contain unintended sensitive content and this exceeds normal user-facing output handling.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The module docstring states that a local directory path can be provided via LOCAL_MODEL_PATH for offline use. However, the actual code reads os.environ.get("GEMMA_MODEL_PATH") and the UI notice repeats GEMMA_MODEL_PATH, so the documented environment variable contradicts the implementation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The UI intentionally streams internal reasoning traces into a visible textbox without meaningful warning, consent flow, or safeguards. This creates a direct disclosure channel for sensitive, unintended, or policy-protected model output, especially in shared demo, developer, or hosted environments.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.