T09 · Insecure Skill Coding Practices
- Location
token_eval.py:556- Finding
Automatic Plaintext Retention of Potentially Sensitive User Prompts
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill has a coherent model-routing purpose, but it defaults to sending prompts to external model APIs, stores prompt text locally, and makes undisclosed telemetry calls.
Install only if you are comfortable with prompt content being sent to external model providers and with prompt snippets being stored locally in plaintext. Prefer explicit /token决策 use, run with --dry-run or --no-exec when you only want a recommendation, and remove or disable CountAPI telemetry and raw prompt logging before using it with confidential data.
token_eval.py:556Automatic Plaintext Retention of Potentially Sensitive User Prompts
token_eval.py:574Undisclosed Third-Party Usage Telemetry Through CountAPI
The skill is configured to auto-trigger on very common phrases such as '推荐模型', '最便宜', '性价比', and similar wording that can appear in ordinary conversation. Because the skill can then proceed to model selection and direct API execution, this creates a real risk of unintended invocation, unnecessary data disclosure to external APIs, and unplanned spend.
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
cluster.py # 聚类分析+关键词自动更新 stats.py # 使用统计面板(含累计节省) benchmark.db # 9模型×8类实测数据库 .env # API Keys(不纳入版本管理)
The file is presented as a token/cost decision helper, but the CLI path also executes the user’s prompt against third-party LLM APIs. That mismatch defeats user expectations and can cause sensitive prompts to be transmitted off-host when a user believed they were only getting a recommendation.
This skill contains full remote prompt execution capability even though its apparent purpose is evaluation/recommendation. In this context, hidden execution substantially increases risk because user inputs may contain secrets, proprietary text, or regulated data that get sent to external vendors unexpectedly.
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
try:
from openai import OpenAI
from dotenv import load_dotenv
env_path = os.path.join(os.path.dirname(__file__), ".env")
if os.path.exists(env_path):
load_dotenv(env_path)
else:
The skill explicitly states that recommendation is immediately followed by calling the model API and returning the result, but it does not mention consent, data handling, or billing consequences at the point of execution. In context, this is more dangerous because the same document also advertises automatic triggering, so a casual request could silently cause external transmission of user prompts and paid usage.
The script reads historical user prompts from benchmark.db and processes them without any visible notice, consent check, minimization, or access control in this code path. Because prompts may contain sensitive or proprietary information, this creates a privacy risk if the tool is run by someone unaware that user-generated content is being mined for secondary analysis.
When invoked with --update or --force, the script can automatically rewrite token_eval.py based on derived keywords, with no approval gate, backup, or integrity validation. This creates a supply-chain and safety risk because untrusted prompt data can indirectly influence source code behavior, potentially degrading classifications or introducing unsafe logic changes into a production skill file.
The statistics display function performs an outbound HTTP request to a third-party analytics service even though its primary purpose is local reporting from a SQLite database. This creates undisclosed data flow and a dependency on an external service, which can leak usage metadata such as the fact that the skill was run, the user's IP address, and timing information.
The code contacts api.countapi.xyz without any in-band disclosure to the user at execution time, which is inconsistent with a local statistics viewer and can surprise users in restricted or privacy-sensitive environments. Even if no local database contents are uploaded, the request still exposes metadata to a third party and introduces tracking and reliability risks.
This line performs direct external transmission to a third-party endpoint from a script that otherwise operates on local statistics. The network call can reveal execution metadata and user/network identifiers, and it expands the attack surface by relying on an external service during a local reporting operation.
# 全量调用
try:
import urllib.request, json
resp = urllib.request.urlopen("https://api.countapi.xyz/get/token-decision/total-calls", timeout=5)
data = json.loads(resp.read())
global_calls = data.get("value", 0)
print(f"\n 🌍 全量调用: {global_calls} 次(所有用户累计)")
This endpoint definition shows that user prompts may be transmitted to DeepSeek’s external API during execution. External transmission is risky in this skill because the tool’s primary description does not make remote data transfer obvious, so users may expose sensitive content unintentionally.
# ============================================================
MODEL_API = {
# DeepSeek 系列: 统一使用 deepseek-chat endpoint
"deepseek-v4-pro": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
"deepseek-v4-flash": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
"deepseek-v3.2": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
# 智谱系列: 注意 GLM-5.1 对应 glm-4-plus endpoint(智谱 API 命名滞后于产品名)
This configuration routes prompts to another external DeepSeek-backed execution path. The risk is not the mere presence of a URL, but that the skill can send user content to third parties without prominent consent aligned with the tool’s stated purpose.
MODEL_API = {
# DeepSeek 系列: 统一使用 deepseek-chat endpoint
"deepseek-v4-pro": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
"deepseek-v4-flash": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
"deepseek-v3.2": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
# 智谱系列: 注意 GLM-5.1 对应 glm-4-plus endpoint(智谱 API 命名滞后于产品名)
"glm-5.1": ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4-plus", "ZHIPU_KEY"),
This external API endpoint enables transmission of prompt content to a third-party service. In context, that is a real privacy/security concern because the skill appears to be an evaluator but also functions as a remote executor.
# DeepSeek 系列: 统一使用 deepseek-chat endpoint
"deepseek-v4-pro": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
"deepseek-v4-flash": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
"deepseek-v3.2": ("deepseek", "https://api.deepseek.com/v1", "deepseek-chat", "DEEPSEEK_KEY"),
# 智谱系列: 注意 GLM-5.1 对应 glm-4-plus endpoint(智谱 API 命名滞后于产品名)
"glm-5.1": ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4-plus", "ZHIPU_KEY"),
"glm-5.0-turbo": ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4-flash", "ZHIPU_KEY"),
This Moonshot/Kimi endpoint is another external transmission destination for prompt data. Multiple providers amplify the exposure surface because fallback behavior may route the same sensitive prompt to more than one vendor on failure.
"glm-5.0-turbo": ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4-flash", "ZHIPU_KEY"),
"glm-5v-turbo": ("zhipu", "https://open.bigmodel.cn/api/paas/v4", "glm-4v-plus", "ZHIPU_KEY"),
# Kimi/Moonshot: 使用 auto 自动选择最佳模型版本
"kimi-k2.6": ("kimi", "https://api.moonshot.cn/v1", "moonshot-v1-auto", "KIMI_KEY"),
# MiniMax
"minimax-m2.5": ("minimax", "https://api.minimaxi.com/v1", "abab6.5s-chat", "MINIMAX_KEY"),
# OpenRouter (Hy3 免费)
This MiniMax endpoint is part of the external execution surface that can receive user prompts. The danger is increased by automatic routing/fallback, which can broaden data sharing beyond the user’s expectations.
# Kimi/Moonshot: 使用 auto 自动选择最佳模型版本
"kimi-k2.6": ("kimi", "https://api.moonshot.cn/v1", "moonshot-v1-auto", "KIMI_KEY"),
# MiniMax
"minimax-m2.5": ("minimax", "https://api.minimaxi.com/v1", "abab6.5s-chat", "MINIMAX_KEY"),
# OpenRouter (Hy3 免费)
"hy3-preview": ("openrouter", "https://openrouter.ai/api/v1", "tencent/hy3-preview:free", "OPENROUTER_KEY"),
}
The execution path sends prompt contents to multiple external model providers, but the skill description does not clearly disclose that prompts leave the local environment. This is dangerous because users may submit confidential or regulated text assuming the tool performs only local analysis.
The logging section does more than local usage logging: it silently performs outbound telemetry to countapi.xyz. Hidden network telemetry is dangerous because it creates undisclosed data flows and undermines informed consent and auditability.
The code stores up to 1000 characters of the user prompt plus task metadata in a local SQLite database without any notice, consent, retention limit, or sanitization policy. Prompt logs frequently contain credentials, internal documents, or personal data, so silent persistence creates a tangible confidentiality and compliance risk.
The combination of persistent local prompt logging and outbound telemetry increases the risk of sensitive natural-language data being retained or indirectly exposed. In a prompt-handling tool, users often paste proprietary or personal content, so silent retention materially raises privacy, insider, and compliance risks.
A telemetry call is made to an external analytics endpoint without user disclosure or consent. Even if only a counter is incremented, hidden analytics traffic creates an undisclosed external dependency and can reveal usage patterns or violate policy in restricted environments.
This telemetry call sends outbound traffic to countapi.xyz as part of ordinary usage logging. Undisclosed analytics beacons are problematic in enterprise or privacy-sensitive settings because they create hidden network dependencies and usage leakage.
# 埋点
try:
import urllib.request
urllib.request.urlopen("https://api.countapi.xyz/hit/token-decision/total-calls", timeout=3)
except:
pass
This function reads analytics data from countapi.xyz, which is another undisclosed external network dependency. It is lower impact than prompt transmission because it does not itself send prompt contents, but it still creates unneeded outbound traffic and operational/privacy concerns.
def get_total_calls() -> int:
try:
import urllib.request, json
resp = urllib.request.urlopen("https://api.countapi.xyz/get/token-decision/total-calls", timeout=5)
return json.loads(resp.read()).get("value", 0)
except:
return -1
The file's user-facing natural-language instructions and labels are entirely in Chinese, with no indication that users may opt into another language or locale. The policy calls for flagging language or locale constraints when they are imposed without user choice or explicit justification.
The header documents python cluster.py as analysis-only, but the same usage block also exposes --update and --force modes that automatically rewrite token_eval.py by updating TASK_KEYWORDS. This is an intent-level contradiction within the file documentation because the tool is not purely analytical; it includes code-modification behavior.
No suspicious patterns detected.