Back to skill

Security audit

Token Usage Dashboard

Security checks for vulnerabilities and agentic risk

Overview

This skill is not clearly malicious, but it ships a broader local dashboard and reporting tool than its model-usage description and can expose or persist detailed prompt and user usage data.

Install only if you intend to use the broader local dashboard/reporting features, not just a simple model-cost summary. Avoid feeding it logs that contain prompts, secrets, customer data, or sensitive user/project identifiers unless you control the output files. Use private output directories, avoid the default /tmp paths, and do not rely on its viewer/RBAC settings as a strong security boundary.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/token_usage_dashboard.py:277
Finding

RBAC Filtering Fails to Remove Sensitive Call-Level Data

Content
View full analysis
List[Dict[str, Any]]: """Build normalized per-call records from row-level payloads. Supported nested keys (per day): llmCalls/apiCalls/requests/events. """ out: List[Dict[str, Any]] = [] for row in rows: day = str(row.get("date") or "") for key in ("llmCalls", "apiCalls", "requests", "events"): raw_calls = row.get(key) if not isinstance(raw_calls, list): continue for c in raw_calls: if not isinstance(c, dict): continue prompt_tokens = _safe_float(c.get("promptTokens") or c.get("inputTokens")) completion_tokens = _safe_float(c.get("completionTokens") or c.get("outputTokens")) total_tokens = _safe_float(c.get("totalTokens"), prompt_tokens + completion_tokens) cost = _safe_float(c.get("cost") or c.get("costUSD")) latency_ms = _safe_float(c.get("latencyMs") or c.get("responseMs") or c.get("durationMs"), default=-1.0) model = c.get("modelName") or c.get("model") or "unknown" model_type = c.get("modelType") or str(model).split("-")[0] task = c.get("useCase") or c.get("task") or c.get("scenario") or "unspecified" out.append( { "date": day, ...[truncated 7964 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
run_dashboard.sh:17
Finding

Predictable Temporary Output Files Permit Symlink-Based Overwrite and Local Disclosure

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (29)

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The README describes a much broader dashboard and governance platform than the skill metadata's narrow purpose of per-model usage summarization. This scope drift is dangerous because an agent may invoke or expose capabilities such as tenant isolation, administration, forecasting, and policy controls that were not expected, reviewed, or permissioned for this skill.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

User, role, and dashboard-view administration are privileged control-plane capabilities that are unjustified for a read-oriented model-usage summary skill. If exposed through an agent, these functions could enable unauthorized access changes, privilege assignment, or cross-tenant data exposure far beyond simple reporting.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

Real-time cost-control actions such as degrade, switch_model, or stop_calls move from analytics into active operational control. In an agent context, this creates a materially higher-risk capability because a prompt or misuse could disrupt service, change runtime behavior, or cause denial of service under the guise of reporting.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill is described as a narrow per-model usage summarizer, but the documented commands also invoke a dashboard generator that writes files and can open a browser. This mismatch can cause an agent or user to authorize broader actions than expected, creating room for unintended file creation, UI launching, or other side effects outside the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is described as a narrow per-model usage summarizer, but the documented commands also invoke a dashboard generator that writes files and can open a browser. This mismatch can cause an agent or user to authorize broader actions than expected, creating room for unintended file creation, UI launching, or other side effects outside the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is described as a narrow per-model usage summarizer, but the documented commands also invoke a dashboard generator that writes files and can open a browser. This mismatch can cause an agent or user to authorize broader actions than expected, creating room for unintended file creation, UI launching, or other side effects outside the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is described as a narrow per-model usage summarizer, but the documented commands also invoke a dashboard generator that writes files and can open a browser. This mismatch can cause an agent or user to authorize broader actions than expected, creating room for unintended file creation, UI launching, or other side effects outside the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is described as a narrow per-model usage summarizer, but the documented commands also invoke a dashboard generator that writes files and can open a browser. This mismatch can cause an agent or user to authorize broader actions than expected, creating room for unintended file creation, UI launching, or other side effects outside the declared purpose.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The test file imports and exercises a far broader capability set than the skill metadata advertises, including dashboard generation, prompt analysis, anomaly detection, alerting, access control, tenant handling, and scheduling. In an agent-skill context, this scope expansion is dangerous because it can hide unexpected data processing and privileged behaviors behind a seemingly simple 'model usage summary' interface, increasing the chance of unauthorized access, data exposure, or unsafe invocation paths.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The tested capabilities include tenant administration, role/view management, multi-tenant context resolution, and scheduled report generation with recipients and artifact history, none of which fit a local cost-summary skill. These features materially expand the trust boundary: if exposed through the agent, they could enable unauthorized configuration changes, cross-tenant data access, persistence of sensitive usage data, and unapproved report delivery.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill includes persistent report scheduling and filesystem-backed report/history generation that substantially exceed a model-usage summarization role. In an agent context, this scope expansion increases the chance of unauthorized data retention, broader data exposure, and unintended side effects on the host system from a skill the user would expect to be read-only and analytical.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

Tenant and user management functionality is unrelated to simple usage summarization and allows the script to modify persistent access-control configuration. In an agent setting, that creates a dangerous privilege boundary crossing: a reporting skill can become an administrative tool capable of changing who can access what data.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document begins in English, then switches to Chinese for multiple key feature bullets and notes. This creates a language/locale policy issue because the skill content forces part of the user-facing description into a specific language without opt-in or explanation.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

Scheduled report automation and report history management exceed the stated purpose of on-demand scriptable per-model summaries and introduce persistence and background execution behavior. While not inherently malicious, these features enlarge the attack surface by storing historical artifacts and enabling unattended processing beyond a simple local summary task.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill advertises shell execution, file read/write, and likely network-adjacent installation behavior, but the manifest does not declare any explicit tool scope such as permissions or allowed-tools. That makes the effective capability boundary ambiguous, which increases the risk of over-privileged execution or unsafe invocation by an agent runtime.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documentation positions the skill as a per-model usage summary tool, yet it instructs users to run a dashboard script that generates HTML and JSON files and can open a browser via --open. That hidden expansion of behavior creates unexpected side effects and may trigger shell, file-write, and local application launch actions that a caller did not intend to authorize.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill manifest says this skill summarizes local CodexBar cost JSON into per-model usage for Codex or Claude, optionally current-model or full breakdown. This roadmap instead documents a much broader platform including multi-tenancy, user/role management, forecasting, anomaly alerts, prompt-content analysis, cloud-cost integrations, report scheduling, and policy controls, which materially exceeds the manifest's narrow scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The roadmap content is written entirely in Traditional Chinese, which can constitute a language/locale policy violation when the skill documentation implicitly requires a specific language without offering alternatives. There is no indication that this locale choice is optional or justified as region-specific.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The checklist marks numerous enterprise features as implemented, such as multi-tenant org management, anomaly warning infrastructure, cost attribution, and scheduled reporting. For a skill documented as a local per-model CodexBar usage summarizer, these statements actively convey a substantially different implemented intent and would mislead reviewers about what the skill does.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest describes a skill focused on summarizing CodexBar cost usage per model, either for the current model or a full model breakdown. This script is documented and configured as a 'token usage dashboard' launcher that produces dashboard HTML and summary JSON, with browser-opening behavior, which is broader than a simple scriptable per-model summary.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/model_usage.py (reported line 37)May include surrounding context.

python
def run_codexbar_cost(provider: str) -> List[Dict[str, Any]]:
    cmd = ["codexbar", "cost", "--format", "json", "--provider", provider]
    try:
        output = subprocess.check_output(cmd, text=True)
    except FileNotFoundError:
        raise RuntimeError("codexbar not found on PATH. Install CodexBar CLI first.")
    except subprocess.CalledProcessError as exc:

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/token_usage_dashboard.py (reported line 74)May include surrounding context.

python
def run_codexbar_cost(provider: str) -> List[Dict[str, Any]]:
    cmd = ["codexbar", "cost", "--format", "json", "--provider", provider]
    try:
        output = subprocess.check_output(cmd, text=True)
    except FileNotFoundError:
        raise RuntimeError("codexbar not found on PATH. Install CodexBar CLI first.")
    except subprocess.CalledProcessError as exc:

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Prompt-content analysis and optimization-engine behavior goes beyond summarizing model-level usage and introduces processing of prompt text, keywords, attribution dimensions, and recommendation logic. Even with some anonymization coverage in tests, this broader inspection of prompt content raises privacy and data-minimization concerns because prompt text may contain secrets, personal data, or business-sensitive inputs that a usage-summary skill should not need to analyze.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The top-level docstring presents the script as a dashboard generator, yet later code supports tenant config management, report scheduler execution, and persistent history/report file creation. That documentation materially understates and mischaracterizes the script's operational intent.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code performs deep call-level analytics, prompt-family analysis, attribution, and optimization planning beyond stated per-model usage summarization. That broader processing materially increases data sensitivity because it touches prompt text, users, projects, sessions, and workflows that are not necessary for the declared function of the skill.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.