Back to skill

Security audit

Gmail Label Routing

Security checks for vulnerabilities and agentic risk

Overview

This Gmail-routing skill appears purpose-aligned, but it should be reviewed carefully because it can use stored Gmail credentials to create or delete filters and bulk move existing messages out of the Inbox.

Install only if you trust this skill to act on the authenticated Gmail account. Use --dry-run first, verify the exact sender addresses and label, prefer --keep-inbox or --no-retro for safer runs, and avoid --replace-sender-filters unless you intentionally want existing Gmail filters for that sender deleted or replaced.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/gws_gmail_label_workflow.py:274
Finding

Unvalidated sender input permits Gmail search-query injection during retroactive modification

Content
View full analysis

Vulnerability Details

File Location: scripts/gws_gmail_label_workflow.py, lines 233-234 and 274-286
Vulnerability Type: Gmail search-query injection
Risk Level: Medium

The workflow accepts arbitrary text through repeated --sender arguments. It strips surrounding whitespace but does not verify that each value is a single valid mailbox address:

python
senders = [s.strip() for s in args.sender if s and s.strip()]
if not senders:
    raise WorkflowError("Provide at least one --sender")

The unvalidated value is then interpolated directly into several Gmail search queries. The first query determines which message IDs will be modified:

python
retro_applied = 0
if not args.no_retro:
    ids = list_all_message_ids(args.user_id, f"from:{sender}")
    retro_applied = batch_modify(
        user_id=args.user_id,
        ids=ids,
        add_label_ids=[label_id],
        remove_label_ids=["INBOX"] if args.remove_inbox else [],
        dry_run=args.dry_run,
    )

from_count = len(list_all_message_ids(args.user_id, f"from:{sender}"))
label_count = len(list_all_message_ids(args.user_id, f"from:{sender} label:\"{args.label}\""))
inbox_count = len(list_all_message_ids(args.user_id, f"from:{sender} in:inbox"))

Technical Analysis

Gmail search queries support operators, grouping, quotation, and logical expressions. Because sender is concatenated into from:{sender} without validation or safe encoding, a value containing Gmail query syntax can change the meaning or scope of the resulting query.

The IDs returned by the injected query are supplied to batch_modify. Unless --keep-inbox is selected, the workflow adds the chosen label and removes INBOX from every returned message. This makes the injection consequential rather than merely affecting displayed search results.

The later count queries are vulnerable to the same input handling and may produce misleading verification r ...[truncated 1723 chars]

Remediation
View remediation

Remediation Suggestions

  1. Validate every --sender value as exactly one mailbox address before using it:

    • Use a standards-aware address parser.
    • Require the parsed address to consume the entire input.
    • Reject display names, multiple addresses, control characters, whitespace, quotes, parentheses, braces, and Gmail query operators.
    • Apply a conservative length limit.
  2. Construct the query through a dedicated helper that enforces the invariant that only validated addresses can reach Gmail search operations:

    python
    import re
    
    SENDER_RE = re.compile(
        r"^[A-Za-z0-9.!#$%&'*+/=?^_`{|}~-]+@"
        r"[A-Za-z0-9](?:[A-Za-z0-9-]{0,61}[A-Za-z0-9])?"
        r"(?:\.[A-Za-z0-9](?:[A-Za-z0-9-]{0,61}[A-Za-z0-9])?)+$"
    )
    
    def validate_sender(value: str) -> str:
        sender = value.strip()
        if not SENDER_RE.fullmatch(sender):
            raise WorkflowError(f"Invalid sender address: {value!r}")
        return sender
    

    Apply this function while building senders, before filter creation or message searches.

  3. If advanced Gmail from: expressions are a legitimate requirement, expose them through a separate, explicitly named option rather than overloading --sender. Advanced-query mode should:

    • Be disabled by default.
    • Display the complete query and number of matched messages.
    • Require explicit confirmation before any retroactive modification.
    • Avoid default INBOX removal.
  4. Add a maximum affected-message safeguard. Abort or require confirmation when the match count exceeds a configurable threshold.

  5. Query and preview matching IDs before mutation, and report the planned number of affected messages. This is especially important because INBOX removal is enabled by default.

  6. Add tests covering whitespace, quotes, parentheses, logical operators, multiple addresses, control characters, malformed domains, and valid international or plus-address ...[truncated 52 chars]

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (10)

Credential Access

High
Category
Privilege Escalation
Confidence
89% confidence
Finding

The workflow is designed to discover and use credential files from fixed root-owned locations, which embeds credential access behavior directly into the skill. In an automated agent environment, this materially increases risk because the skill can leverage preexisting secrets to obtain Gmail access without requiring the operator to supply credentials at invocation time.

Content

Scanner excerpt · scripts/gws_gmail_label_workflow.py (reported line 17)May include surrounding context.

python
DEFAULT_CREDENTIAL_CANDIDATES = [
    "/root/.config/gws/credentials.new.json",
    "/root/.config/gws/credentials.json",
]

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The skill instructs the agent to execute a local Python workflow that can read local files, invoke shell, and interact with Gmail/networked APIs, but the manifest declares no explicit tool scope or permission boundaries. That creates an authorization and review gap: the skill appears less powerful than it is, making unintended or unsafe execution more likely.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The description contains broad natural-language triggers like 'manda este remitente a esta etiqueta' and 'hazlo para varios remitentes' that can match ordinary conversation without strong disambiguation. In an agentic environment, ambiguous invocation increases the chance the skill runs and changes mailbox rules when the user did not intend to launch this workflow specifically.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill description does not prominently warn that it may retroactively relabel existing mail in bulk and remove messages from the Inbox by default. Because these are high-impact mailbox modifications, insufficient disclosure can lead to unexpected data organization changes, missed messages, or accidental mass processing.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/gws_gmail_label_workflow.py (reported line 26)May include surrounding context.

python
def run_json(cmd: list[str]) -> dict[str, Any]:
    cp = subprocess.run(cmd, capture_output=True, text=True)
    if cp.returncode != 0:
        detail = (cp.stderr or cp.stdout or "").strip()
        raise WorkflowError(detail or f"command failed: {' '.join(cmd)}")

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The workflow automatically reads OAuth refresh-token material from /root/.config/gws/... and exchanges it for an access token without any explicit confirmation, disclosure, or scope validation. In an agent-skill context, this creates a credential-use primitive that can silently act on a user's Gmail account, increasing the risk of unauthorized mailbox access if the skill is triggered unexpectedly or repurposed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script can delete existing filters and retroactively relabel or remove messages from the inbox at scale, yet it performs these state-changing operations without an explicit destructive-action confirmation step. In this skill context, that is especially dangerous because a natural-language request could trigger broad, hard-to-notice mailbox changes affecting organization, visibility, and future delivery behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The headings and instructional text are presented in Spanish, which can amount to a language-policy issue if the skill implicitly forces one language without user opt-in. No alternate language option or explanation of a region-specific requirement is provided in this file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

This markdown file provides command examples that appear to create or replace Gmail filtering behavior, including moving messages out of INBOX and replacing duplicate/conflicting filters. While the examples are clear, the file does not include any warning that these actions can alter mailbox organization or replace existing rules.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
98% confidence
Finding

The argument parser description and help strings are written entirely in Spanish, which imposes a specific language for user-facing interaction without any opt-in or documented locale constraint. This matches the policy category for language or locale restrictions that are not optional or justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.