Back to skill

Security audit

Model Routing Middleware

Security checks for vulnerabilities and agentic risk

Overview

The skill is a plausible model-routing helper, but it logs prompt text by default and can steer requests toward cloud models without a clear consent boundary.

Review before installing or integrating. Disable or remove prompt-preview logging, add redaction if logs are needed, and make cloud routing or escalation opt-in for sensitive workflows. Also verify the package imports and model-key configuration before relying on it operationally.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/router.py:217
Finding

Unredacted Prompt Content Is Logged by Default

Content
View full analysis

Vulnerability Details

File Location: scripts/router.py:217-225
Related Configuration: scripts/config.yaml:157-162
Vulnerability Type: Sensitive information exposure through application logs
Risk Level: Medium

Vulnerable Code

python
log_config = self.config.get("logging", {})
preview_len = log_config.get("prompt_preview_length", 200)
preview = prompt[:preview_len] + "..." if len(prompt) > preview_len else prompt

logger.info(
    f"ROUTE | model={route.model} | think={route.think} | "
    f"task={route.task_type.value} | confidence={route.confidence:.2f} | "
    f"ctx_override={route.context_overridden} | context={context_size} | "
    f"prompt=\"{preview}\""
)

The corresponding default configuration is:

yaml
logging:
  enabled: true
  path: routing.log
  level: INFO
  include_prompt_preview: true
  prompt_preview_length: 200

Technical Analysis

Routing logs are enabled by default, and _log_route() places up to 200 characters of the user-controlled prompt into an INFO log message without redaction. Prompts can contain API keys, passwords, personal information, proprietary source code, confidential business data, or security-sensitive instructions.

Although the configuration defines include_prompt_preview, the implementation never checks that setting. Consequently, changing include_prompt_preview to false does not suppress prompt content as long as routing logging remains enabled.

The prompt is also inserted without escaping carriage returns, line feeds, or other control characters. An attacker can therefore supply a multiline prompt that creates misleading or forged entries in text-based log viewers. This log-injection aspect can impede investigations, although it does not grant code execution.

The configured path is not used by this implementation; output is sent through a logging.StreamHandler. The exposure scope therefore depends on the embedding application's standard-error capture, proces ...[truncated 1449 chars]

Remediation
View remediation

Remediation Suggestions

  1. Disable prompt-content logging by default and log only non-sensitive routing metadata.
  2. Enforce the existing configuration flag before constructing a preview:
python
include_preview = log_config.get("include_prompt_preview", False)
preview = ""

if include_preview:
    preview_len = max(0, min(int(log_config.get("prompt_preview_length", 0)), 200))
    preview = redact_sensitive_data(prompt[:preview_len])
    preview = preview.replace("\r", "\\r").replace("\n", "\\n")
  1. Omit the prompt field entirely when include_prompt_preview is false rather than logging an empty or placeholder value.
  2. Apply tested redaction for authorization headers, API keys, access tokens, passwords, private keys, email addresses, and other organization-specific sensitive values.
  3. Escape or remove control characters to prevent multiline log injection and forged records.
  4. Prefer structured logging with separate fields and ensure the logging backend safely serializes untrusted text.
  5. Restrict log access, encrypt logs in transit and at rest, define short retention periods, and prevent unnecessary forwarding to third-party analytics systems.
  6. Add regression tests verifying that prompts are absent when preview logging is disabled, common secret formats are redacted, and CR/LF characters cannot create additional log records.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (13)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

This code chunk is a test suite for classifiers.py, validating how text prompts are categorized into task types and checking confidence scores. That behavior is related to task detection, which could be a supporting component of a routing system, but the chunk itself does not perform the declared primary functions of model selection middleware: no routing to models, no context handling, no API interaction, and no cost-management logic are shown. Because the actual code’s primary purpose is classifier testing rather than the declared middleware functionality, this is a material description-to-behavior mismatch for the supplied chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description emphasizes broad intelligent model selection, context management, and cost optimization. The actual code chunk does not implement or test those behaviors directly. Instead, it focuses narrowly on escalation: detecting uncertain phrases like "I don't know," assigning confidence scores, and escalating to the next model in a fixed chain. There is no evidence here of context management, cost-cutting logic, or generalized task routing to the best model. While escalation between models is adjacent to model selection, the primary purpose of this code is materially different and significantly narrower than the declared description.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger phrase is broad enough to activate on many ordinary discussions about model routing, LLM selection, or cost optimization. Over-broad activation can cause the skill to run in unintended contexts, which is risky for middleware that may influence downstream model choice, context handling, or tool invocation in an agent pipeline.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/config.yaml (reported line 51)May include surrounding context.

yaml
# Example: Cloud large-context model
  cloud-large:
    provider: openai
    endpoint: https://api.openai.com/v1
    model_id: gpt-4o
    context_limit: 128000
    supports_think: true

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/config.yaml (reported line 64)May include surrounding context.

yaml
# Example: Cloud large-context model
  cloud-large:
    provider: openai
    endpoint: https://api.openai.com/v1
    model_id: gpt-4o
    context_limit: 128000
    supports_think: true

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The file explicitly defines an escalation chain that upgrades to a cloud model (glm-5-1-cloud) and is designed to retry based on model response content. In practice, that means prior prompt/response material may be forwarded to a cloud provider during escalation, but this file contains no consent gate, sensitivity check, redaction step, or user-facing warning before that boundary is crossed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The function builds a prompt containing prior message contents so they can be sent to a model for summarization, which is a network/data-transmission relevant operation in the context of this skill. While the docstring explains the technical purpose, there is no user-facing warning, confirmation, or disclosure that potentially sensitive conversation history will be shared for summarization.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The inline comment on L151 states that think mode is preserved from the original task, but the variable rule was looked up after task_type was reset to TaskType.CHAT on L137-L145. As written, think = rule.get("think", False) uses the overridden task's rule rather than the original task's rule, directly contradicting the comment.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The router logs a preview of the raw prompt, which can include secrets, personal data, proprietary text, or regulated content unrelated to the middleware's routing purpose. Because this component sits centrally in request handling, it may systematically capture sensitive user inputs into logs that are retained, exported, or viewed by operators.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

Prompt text is written to logs without any explicit consent boundary or warning, creating a direct data exposure path for sensitive user input. In an AI routing middleware, prompts often contain API keys, credentials, customer data, source code, or internal business information, so even truncated previews can leak material secrets.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The configuration routes vision tasks directly to a cloud-hosted model endpoint, but there is no natural-language indication here that users can opt out of remote processing or choose an alternative. While this is not a strict security flaw by itself, it can conflict with organizational policy expectations around user choice for externally processed requests.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/router.py (reported line 84)May include surrounding context.

python
# Set up logging
        log_config = self.config.get("logging", {})
        if log_config.get("enabled", True):
            log_level = getattr(logging, log_config.get("level", "INFO"), logging.INFO)
            logger.setLevel(log_level)
            if not logger.handlers:
                handler = logging.StreamHandler()

Static analysis

No suspicious patterns detected.