Back to skill

Security audit

Token Decision

Security checks for vulnerabilities and agentic risk

Overview

This skill has a coherent model-routing purpose, but it defaults to sending prompts to external model APIs, stores prompt text locally, and makes undisclosed telemetry calls.

Install only if you are comfortable with prompt content being sent to external model providers and with prompt snippets being stored locally in plaintext. Prefer explicit /token决策 use, run with --dry-run or --no-exec when you only want a recommendation, and remove or disable CountAPI telemetry and raw prompt logging before using it with confidential data.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
token_eval.py:556
Finding

Automatic Plaintext Retention of Potentially Sensitive User Prompts

Content
View full analysis
Remediation
View remediation

other

Note
Location
token_eval.py:574
Finding

Undisclosed Third-Party Usage Telemetry Through CountAPI

Content
View full analysis
int: try: import urllib.request, json resp = urllib.request.urlopen("https://api.countapi.xyz/get/token-decision/total-calls", timeout=5) return json.loads(resp.read()).get("value", 0) except: return -1 ``` `stats.py:99-106`: ```python # Global invocation count try: import urllib.request, json resp = urllib.request.urlopen("https://api.countapi.xyz/get/token-decision/total-calls", timeout=5) data = json.loads(resp.read()) global_calls = data.get("value", 0) print(f"\n 🌍 Global calls: {global_calls} total across all users") except: pass ``` ### Technical Analysis After recording a recommendation, the Skill contacts `api.countapi.xyz` to increment a global invocation counter. The main and statistics modules can also contact the same service to retrieve the counter. The request URLs do not contain the prompt, API credentials, environment variables, or database content. Therefore, the static pre-scan concern that these requests directly transmit sensitive prompt data is not substantiated by the reviewed code. Nevertheless, the requests disclose network metadata to a third party, including the originating IP address and approximate Skill invocation timing. The counter increment occurs as part of `log_usage()` and is therefore made even when model execution is disabled with `--dry-run`. This telemetry is not necessary for recommendation, local budget tracking, model execution, f ...[truncated 1332 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (27)

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill is configured to auto-trigger on very common phrases such as '推荐模型', '最便宜', '性价比', and similar wording that can appear in ordinary conversation. Because the skill can then proceed to model selection and direct API execution, this creates a real risk of unintended invocation, unnecessary data disclosure to external APIs, and unplanned spend.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 88)May include surrounding context.

cluster.py # 聚类分析+关键词自动更新 stats.py # 使用统计面板(含累计节省) benchmark.db # 9模型×8类实测数据库 .env # API Keys(不纳入版本管理)

text

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The file is presented as a token/cost decision helper, but the CLI path also executes the user’s prompt against third-party LLM APIs. That mismatch defeats user expectations and can cause sensitive prompts to be transmitted off-host when a user believed they were only getting a recommendation.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This skill contains full remote prompt execution capability even though its apparent purpose is evaluation/recommendation. In this context, hidden execution substantially increases risk because user inputs may contain secrets, proprietary text, or regulated data that get sent to external vendors unexpectedly.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · token_eval.py (reported line 474)May include surrounding context.

python
try:
        from openai import OpenAI
        from dotenv import load_dotenv
        env_path = os.path.join(os.path.dirname(__file__), ".env")
        if os.path.exists(env_path):
            load_dotenv(env_path)
        else:

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly states that recommendation is immediately followed by calling the model API and returning the result, but it does not mention consent, data handling, or billing consequences at the point of execution. In context, this is more dangerous because the same document also advertises automatic triggering, so a casual request could silently cause external transmission of user prompts and paid usage.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The script reads historical user prompts from benchmark.db and processes them without any visible notice, consent check, minimization, or access control in this code path. Because prompts may contain sensitive or proprietary information, this creates a privacy risk if the tool is run by someone unaware that user-generated content is being mined for secondary analysis.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

When invoked with --update or --force, the script can automatically rewrite token_eval.py based on derived keywords, with no approval gate, backup, or integrity validation. This creates a supply-chain and safety risk because untrusted prompt data can indirectly influence source code behavior, potentially degrading classifications or introducing unsafe logic changes into a production skill file.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The statistics display function performs an outbound HTTP request to a third-party analytics service even though its primary purpose is local reporting from a SQLite database. This creates undisclosed data flow and a dependency on an external service, which can leak usage metadata such as the fact that the skill was run, the user's IP address, and timing information.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The code contacts api.countapi.xyz without any in-band disclosure to the user at execution time, which is inconsistent with a local statistics viewer and can surprise users in restricted or privacy-sensitive environments. Even if no local database contents are uploaded, the request still exposes metadata to a third party and introduces tracking and reliability risks.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

This line performs direct external transmission to a third-party endpoint from a script that otherwise operates on local statistics. The network call can reveal execution metadata and user/network identifiers, and it expands the attack surface by relying on an external service during a local reporting operation.

Content

Scanner excerpt · stats.py (reported line 101)May include surrounding context.

python
# 全量调用
    try:
        import urllib.request, json
        resp = urllib.request.urlopen("https://api.countapi.xyz/get/token-decision/total-calls", timeout=5)
        data = json.loads(resp.read())
        global_calls = data.get("value", 0)
        print(f"\n  🌍 全量调用: {global_calls} 次(所有用户累计)")

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This endpoint definition shows that user prompts may be transmitted to DeepSeek’s external API during execution. External transmission is risky in this skill because the tool’s primary description does not make remote data transfer obvious, so users may expose sensitive content unintentionally.

Content

Scanner excerpt · token_eval.py (reported line 449)May include surrounding context.

python
# ============================================================
MODEL_API = {
    # DeepSeek 系列: 统一使用 deepseek-chat endpoint
    "deepseek-v4-pro":   ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
    "deepseek-v4-flash": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
    "deepseek-v3.2":     ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
    # 智谱系列: 注意 GLM-5.1 对应 glm-4-plus endpoint(智谱 API 命名滞后于产品名)

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This configuration routes prompts to another external DeepSeek-backed execution path. The risk is not the mere presence of a URL, but that the skill can send user content to third parties without prominent consent aligned with the tool’s stated purpose.

Content

Scanner excerpt · token_eval.py (reported line 450)May include surrounding context.

python
MODEL_API = {
    # DeepSeek 系列: 统一使用 deepseek-chat endpoint
    "deepseek-v4-pro":   ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
    "deepseek-v4-flash": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
    "deepseek-v3.2":     ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
    # 智谱系列: 注意 GLM-5.1 对应 glm-4-plus endpoint(智谱 API 命名滞后于产品名)
    "glm-5.1":          ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4-plus", "ZHIPU_KEY"),

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This external API endpoint enables transmission of prompt content to a third-party service. In context, that is a real privacy/security concern because the skill appears to be an evaluator but also functions as a remote executor.

Content

Scanner excerpt · token_eval.py (reported line 451)May include surrounding context.

python
# DeepSeek 系列: 统一使用 deepseek-chat endpoint
    "deepseek-v4-pro":   ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
    "deepseek-v4-flash": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
    "deepseek-v3.2":     ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
    # 智谱系列: 注意 GLM-5.1 对应 glm-4-plus endpoint(智谱 API 命名滞后于产品名)
    "glm-5.1":          ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4-plus", "ZHIPU_KEY"),
    "glm-5.0-turbo":    ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4-flash", "ZHIPU_KEY"),

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This Moonshot/Kimi endpoint is another external transmission destination for prompt data. Multiple providers amplify the exposure surface because fallback behavior may route the same sensitive prompt to more than one vendor on failure.

Content

Scanner excerpt · token_eval.py (reported line 457)May include surrounding context.

python
"glm-5.0-turbo":    ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4-flash", "ZHIPU_KEY"),
    "glm-5v-turbo":     ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4v-plus", "ZHIPU_KEY"),
    # Kimi/Moonshot: 使用 auto 自动选择最佳模型版本
    "kimi-k2.6":        ("kimi", "https://api.moonshot.cn/v1", "moonshot-v1-auto", "KIMI_KEY"),
    # MiniMax
    "minimax-m2.5":     ("minimax", "https://api.minimaxi.com/v1", "abab6.5s-chat", "MINIMAX_KEY"),
    # OpenRouter (Hy3 免费)

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This MiniMax endpoint is part of the external execution surface that can receive user prompts. The danger is increased by automatic routing/fallback, which can broaden data sharing beyond the user’s expectations.

Content

Scanner excerpt · token_eval.py (reported line 459)May include surrounding context.

python
# Kimi/Moonshot: 使用 auto 自动选择最佳模型版本
    "kimi-k2.6":        ("kimi", "https://api.moonshot.cn/v1", "moonshot-v1-auto", "KIMI_KEY"),
    # MiniMax
    "minimax-m2.5":     ("minimax", "https://api.minimaxi.com/v1", "abab6.5s-chat", "MINIMAX_KEY"),
    # OpenRouter (Hy3 免费)
    "hy3-preview":      ("openrouter", "https://openrouter.ai/api/v1", "tencent/hy3-preview:free", "OPENROUTER_KEY"),
}

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The execution path sends prompt contents to multiple external model providers, but the skill description does not clearly disclose that prompts leave the local environment. This is dangerous because users may submit confidential or regulated text assuming the tool performs only local analysis.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The logging section does more than local usage logging: it silently performs outbound telemetry to countapi.xyz. Hidden network telemetry is dangerous because it creates undisclosed data flows and undermines informed consent and auditability.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code stores up to 1000 characters of the user prompt plus task metadata in a local SQLite database without any notice, consent, retention limit, or sanitization policy. Prompt logs frequently contain credentials, internal documents, or personal data, so silent persistence creates a tangible confidentiality and compliance risk.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The combination of persistent local prompt logging and outbound telemetry increases the risk of sensitive natural-language data being retained or indirectly exposed. In a prompt-handling tool, users often paste proprietary or personal content, so silent retention materially raises privacy, insider, and compliance risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

A telemetry call is made to an external analytics endpoint without user disclosure or consent. Even if only a counter is incremented, hidden analytics traffic creates an undisclosed external dependency and can reveal usage patterns or violate policy in restricted environments.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
96% confidence
Finding

This telemetry call sends outbound traffic to countapi.xyz as part of ordinary usage logging. Undisclosed analytics beacons are problematic in enterprise or privacy-sensitive settings because they create hidden network dependencies and usage leakage.

Content

Scanner excerpt · token_eval.py (reported line 579)May include surrounding context.

python
# 埋点
    try:
        import urllib.request
        urllib.request.urlopen("https://api.countapi.xyz/hit/token-decision/total-calls", timeout=3)
    except:
        pass

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This function reads analytics data from countapi.xyz, which is another undisclosed external network dependency. It is lower impact than prompt transmission because it does not itself send prompt contents, but it still creates unneeded outbound traffic and operational/privacy concerns.

Content

Scanner excerpt · token_eval.py (reported line 587)May include surrounding context.

python
def get_total_calls() -> int:
    try:
        import urllib.request, json
        resp = urllib.request.urlopen("https://api.countapi.xyz/get/token-decision/total-calls", timeout=5)
        return json.loads(resp.read()).get("value", 0)
    except:
        return -1

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The file's user-facing natural-language instructions and labels are entirely in Chinese, with no indication that users may opt into another language or locale. The policy calls for flagging language or locale constraints when they are imposed without user choice or explicit justification.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The header documents python cluster.py as analysis-only, but the same usage block also exposes --update and --force modes that automatically rewrite token_eval.py by updating TASK_KEYWORDS. This is an intent-level contradiction within the file documentation because the tool is not purely analytical; it includes code-modification behavior.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.