Back to skill

Security audit

Tavily Quota Router

Security checks for vulnerabilities and agentic risk

Overview

The skill appears purpose-built for Tavily routing, but it needs Review because it stores live API keys locally and its rate-limit safety promises conflict with actual failover behavior.

Review before installing. Use this only if you intentionally want a multi-key Tavily router, store keys outside the skill directory when possible, add your own gitignore and file-permission protections, and do not rely on the documented 'stop on first 429' safety rule unless the code is fixed to enforce it.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/tavily_multi_key.py:11
Finding

Plaintext API Key Storage with Incomplete Repository Leak Protection

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (25)

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The quickstart instructs users that the router will continue operating after a 429 by rotating to other keys, which conflicts with the stated safety rule to stop on the first 429 and avoid retries during cooldown. In a quota-aware multi-key search tool, this guidance can cause an agent to bypass provider rate limiting and continue external requests when it was supposed to halt, increasing the risk of policy violations and quota exhaustion.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This usage guidance explicitly tells users that when key A gets a 429, the next search automatically uses key B, directly undermining the critical rule to stop on the first 429. That creates dangerous operator expectations and could drive implementations or agent behavior that evade rate-limit backoff by cycling credentials.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The troubleshooting table reinforces the same unsafe operating model by describing post-429 reuse behavior without emphasizing the required immediate stop. In a skill specifically designed to route around quota issues, such contradictory documentation materially increases the chance of agents or users treating rate limits as something to route around rather than respect.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill promises strict safety behavior such as stopping on first 429 and limiting searches, yet the finding says the actual search loop can continue failing over across multiple keys after a 429. That mismatch can amplify rate-limit cascades, consume multiple API keys in one invocation, and violate the operational guarantees users rely on.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill promises strict safety behavior such as stopping on first 429 and limiting searches, yet the finding says the actual search loop can continue failing over across multiple keys after a 429. That mismatch can amplify rate-limit cascades, consume multiple API keys in one invocation, and violate the operational guarantees users rely on.

Content

No source excerpt is available for this finding.

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · SKILL.md (reported line 229)May include surrounding context.

md
- **`last_sync_at` ≠ real-time health.** The state file snapshots when `/usage` was last called. If it's months old, key health may have changed in either direction (recovered or newly banned). Always refresh before reasoning about a key's status.
**老大 (2026-07-14) "你的检查机制是什么" — answer:** when web_search fails or router reports "no available key", the FIRST move is always `cat state/quota.json | python3 -m json.tool` and compute `now() - cooldown_until`. Reporting "quota exhausted" or "Tavily down" without this is the same class of error as "I assumed it was Ollama" — verify before stating. The router's error message is a **local cache state**, not a Tavily API verdict.

- **`cooldown_until` ≠ backend dead.** A cooldown marker was written after a single 429 burst, but the upstream may have recovered seconds later. **Never report "all keys exhausted / Tavily down" based on `state/quota.json` alone.** Before declaring the backend dead, parallel-curl each key against `https://api.tavily.com/search` directly (bypassing the router). If all return 200, the state is stale — run `reset-month` or `test-keys` to refresh, then proceed. The router state is a local cache, not ground truth. Verified 2026-07-10: four keys with `cooldown_until` 10 min in the future were all 200 OK on direct curl; reporting "exhausted" without testing is the same class of error as "I assumed it was Ollama." Test before reporting.
- **Router 错误消息 ≠ quota 真实状态**(2026-07-14 老大原话 "怎么可能额度用完了")。router 报 "no available key (all disabled, cooled, or quota exhausted)" 时**第一步永远是看 quota.json**,**不**直接信 router 报错。**诊断顺序**:
  1. `cat state/quota.json | python3 -c "import json,sys; d=json.load(sys.stdin); [print(k['cooldown_until'], k.get('disabled'), k.get('plan_usage'),'/',k.get('plan_limit')) for k in d['keys']]"`
  2. 对比当前时间 vs `cooldown_until` —— 如果 cooldown_until < now → **cooldown 已过**,**不是真死**

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Although honoring Retry-After is normally good practice, in this context the documentation frames 429 handling as part of continued routing behavior despite the manifest requiring an immediate stop on first 429. The mismatch is dangerous because users may infer the router can absorb rate limits transparently instead of halting searches when required.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill describes and instructs operations that read local credential/state files, modify configuration, and send authenticated requests to an external API, but it declares no explicit tool scope or permissions boundary. In an agent ecosystem, that omission weakens least-privilege controls and can cause the skill to be loaded with broader file and network access than users expect.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

The documented workflow relies on live calls to an external service to verify rate-limit behavior, which transmits operational metadata and authenticated requests off-host. In a skill that also stores plaintext keys locally, this raises the security significance because local secrets are routinely used in outbound commands.

Content

Scanner excerpt · SKILL.md (reported line 155)May include surrounding context.

md
2. **Code path (wired through search path on 2026-07-17)**: `mark_error(cfg, state, idx, msg, disable=False, retry_after_seconds=None)` — new optional kwarg. The `cmd_search` HTTPError branch now parses `e.headers.get('Retry-After')` (integer seconds, per Tavily's actual 429 response) and passes it through. **Verified end-to-end on 2026-07-17**: hitting key[0] with 30 concurrent requests triggered 429, router captured `Retry-After: 60`, set `cooldown_until = now + 60s` (not 600s), and the key was queryable again after a single `sleep 50` — matching the documented behavior. Before this fix, SKILL.md claimed "applied 2026-07-15" but the search call path was NOT actually wired through — see `references/tavily-429-retry-after-2026-07-17.md` for the full diff and verification transcript.  **Operational rule (mandatory after every fix like this):** do not trust SKILL.md claims about "applied / wired / patched" without re-running the 4-step verification (`py_compile` + `reset-month` + trigger 429 + read quota.json `cooldown_until`). The bug pattern is: someone adds the parameter, writes the doc, but never calls the parameter from the search path. If the verification step is skipped, you end up with a skill that lies about its own behavior.

  **Operational rule**: when user reports "I just need it to retry in 10-20s, not 10 minutes", first check if the actual recovery time is sub-minute by directly `curl https://api.tavily.com/search -H "Authorization: Bearer $KEY" -d '{"query":"ping"}'`. If it returns 200, the router is over-conservative — set `cooldown_minutes: 0.5` (or lower) and confirm.

**Workaround to change it:** edit `~/.hermes/skills/openclaw-imports/tavily-quota-router/config/keys.json` → `cooldown_minutes` field. Takes effect on next router run.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

The documented workflow relies on live calls to an external service to verify rate-limit behavior, which transmits operational metadata and authenticated requests off-host. In a skill that also stores plaintext keys locally, this raises the security significance because local secrets are routinely used in outbound commands.

Content

Scanner excerpt · SKILL.md (reported line 155)May include surrounding context.

md
2. **Code path (wired through search path on 2026-07-17)**: `mark_error(cfg, state, idx, msg, disable=False, retry_after_seconds=None)` — new optional kwarg. The `cmd_search` HTTPError branch now parses `e.headers.get('Retry-After')` (integer seconds, per Tavily's actual 429 response) and passes it through. **Verified end-to-end on 2026-07-17**: hitting key[0] with 30 concurrent requests triggered 429, router captured `Retry-After: 60`, set `cooldown_until = now + 60s` (not 600s), and the key was queryable again after a single `sleep 50` — matching the documented behavior. Before this fix, SKILL.md claimed "applied 2026-07-15" but the search call path was NOT actually wired through — see `references/tavily-429-retry-after-2026-07-17.md` for the full diff and verification transcript.  **Operational rule (mandatory after every fix like this):** do not trust SKILL.md claims about "applied / wired / patched" without re-running the 4-step verification (`py_compile` + `reset-month` + trigger 429 + read quota.json `cooldown_until`). The bug pattern is: someone adds the parameter, writes the doc, but never calls the parameter from the search path. If the verification step is skipped, you end up with a skill that lies about its own behavior.

  **Operational rule**: when user reports "I just need it to retry in 10-20s, not 10 minutes", first check if the actual recovery time is sub-minute by directly `curl https://api.tavily.com/search -H "Authorization: Bearer $KEY" -d '{"query":"ping"}'`. If it returns 200, the router is over-conservative — set `cooldown_minutes: 0.5` (or lower) and confirm.

**Workaround to change it:** edit `~/.hermes/skills/openclaw-imports/tavily-quota-router/config/keys.json` → `cooldown_minutes` field. Takes effect on next router run.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

L219 says disabled: true is no longer set by any error path and that 401/403 keys auto-recover after a temporary cooldown. L225 directly contradicts that by warning that 401'd keys stay disabled across restarts until test-keys clears them, which is an active contradiction in the skill's own intent/documentation.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Advising operators to probe each key directly against the external API encourages repeated authenticated egress and broadens who may handle live credentials. While not malicious, this can normalize insecure debugging practices and increase leak surface through terminals, history files, and transcripts.

Content

Scanner excerpt · SKILL.md (reported line 229)May include surrounding context.

md
- **`last_sync_at` ≠ real-time health.** The state file snapshots when `/usage` was last called. If it's months old, key health may have changed in either direction (recovered or newly banned). Always refresh before reasoning about a key's status.
**老大 (2026-07-14) "你的检查机制是什么" — answer:** when web_search fails or router reports "no available key", the FIRST move is always `cat state/quota.json | python3 -m json.tool` and compute `now() - cooldown_until`. Reporting "quota exhausted" or "Tavily down" without this is the same class of error as "I assumed it was Ollama" — verify before stating. The router's error message is a **local cache state**, not a Tavily API verdict.

- **`cooldown_until` ≠ backend dead.** A cooldown marker was written after a single 429 burst, but the upstream may have recovered seconds later. **Never report "all keys exhausted / Tavily down" based on `state/quota.json` alone.** Before declaring the backend dead, parallel-curl each key against `https://api.tavily.com/search` directly (bypassing the router). If all return 200, the state is stale — run `reset-month` or `test-keys` to refresh, then proceed. The router state is a local cache, not ground truth. Verified 2026-07-10: four keys with `cooldown_until` 10 min in the future were all 200 OK on direct curl; reporting "exhausted" without testing is the same class of error as "I assumed it was Ollama." Test before reporting.
- **Router 错误消息 ≠ quota 真实状态**(2026-07-14 老大原话 "怎么可能额度用完了")。router 报 "no available key (all disabled, cooled, or quota exhausted)" 时**第一步永远是看 quota.json**,**不**直接信 router 报错。**诊断顺序**:
  1. `cat state/quota.json | python3 -c "import json,sys; d=json.load(sys.stdin); [print(k['cooldown_until'], k.get('disabled'), k.get('plan_usage'),'/',k.get('plan_limit')) for k in d['keys']]"`
  2. 对比当前时间 vs `cooldown_until` —— 如果 cooldown_until < now → **cooldown 已过**,**不是真死**

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

The step-by-step diagnostic path explicitly includes external requests using bearer credentials, so the skill is encouraging outbound use of sensitive tokens during troubleshooting. In context this is functionally related, but still a real security concern because it expands exposure paths beyond the main application flow.

Content

Scanner excerpt · SKILL.md (reported line 234)May include surrounding context.

md
1. `cat state/quota.json | python3 -c "import json,sys; d=json.load(sys.stdin); [print(k['cooldown_until'], k.get('disabled'), k.get('plan_usage'),'/',k.get('plan_limit')) for k in d['keys']]"`
  2. 对比当前时间 vs `cooldown_until` —— 如果 cooldown_until < now → **cooldown 已过**,**不是真死**
  3. 对比 `plan_usage / plan_limit` —— < 80% 不算满
  4. 直接 `curl https://api.tavily.com/search -H "Authorization: Bearer $KEY" -d '{"query":"ping"}'` 实测 → 200 = quota 正常
  5. 都过了才报告"quota 真死";不然重试 1 次再判定

  **真实陷阱**(2026-07-14):router 报 "all disabled, cooled, or quota exhausted",实际 quota.json 里 4 个 key 都 `cooldown_until: 21:54`,当前 22:22 —— **cooldown 早过了 28 分钟**,月度配额只用 37/1000 (3.7%)。router **没刷新缓存**导致误判。**老大原话"你的检查机制是什么"** —— 答案是先看 quota.json,**不**信 router 报错。

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

The step-by-step diagnostic path explicitly includes external requests using bearer credentials, so the skill is encouraging outbound use of sensitive tokens during troubleshooting. In context this is functionally related, but still a real security concern because it expands exposure paths beyond the main application flow.

Content

Scanner excerpt · SKILL.md (reported line 234)May include surrounding context.

md
1. `cat state/quota.json | python3 -c "import json,sys; d=json.load(sys.stdin); [print(k['cooldown_until'], k.get('disabled'), k.get('plan_usage'),'/',k.get('plan_limit')) for k in d['keys']]"`
  2. 对比当前时间 vs `cooldown_until` —— 如果 cooldown_until < now → **cooldown 已过**,**不是真死**
  3. 对比 `plan_usage / plan_limit` —— < 80% 不算满
  4. 直接 `curl https://api.tavily.com/search -H "Authorization: Bearer $KEY" -d '{"query":"ping"}'` 实测 → 200 = quota 正常
  5. 都过了才报告"quota 真死";不然重试 1 次再判定

  **真实陷阱**(2026-07-14):router 报 "all disabled, cooled, or quota exhausted",实际 quota.json 里 4 个 key 都 `cooldown_until: 21:54`,当前 22:22 —— **cooldown 早过了 28 分钟**,月度配额只用 37/1000 (3.7%)。router **没刷新缓存**导致误判。**老大原话"你的检查机制是什么"** —— 答案是先看 quota.json,**不**信 router 报错。

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

L243 states that cmd_search calls sync_all_usage before every search and then says the fix was applied on 2026-07-13 by removing that call. Elsewhere, the skill consistently describes official usage-based routing and current behavior, so this line is self-contradictory documentation about whether the fix is present now, indicating the docs may not match the actual implemented search path.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The verification snippet directly overwrites config/keys.json with a fake API key, which changes live credential configuration in-place and can break subsequent searches if the restore step is missed or interrupted. In this skill context, that is more dangerous because the document is operational guidance for a quota-routing component, so users may copy-paste the commands into a real environment and accidentally cause a self-inflicted denial of service or credential loss state.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This markdown file contains a code comment written as an instruction in Chinese ('429 时解析...'), and the document also includes Chinese-only phrasing later. Under the policy criteria, forcing a specific language or locale without user choice is a natural-language policy violation unless clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill manifest's critical rule says to stop on first 429 and not retry within cooldown. In this file, lines L128-L139 prescribe stopping immediately on the first 429, yet the transcript at L103-L118 documents a router search that records 429 results for multiple keys, indicating continued failover attempts after encountering a 429. That is an active contradiction in the skill's own documentation about intended behavior.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · QUICKSTART.md (reported line 9)May include surrounding context.

md
BASE = Path(__file__).resolve().parents[1]
CONFIG = BASE / 'config' / 'keys.json'
STATE = BASE / 'state' / 'quota.json'
SEARCH_URL = 'https://api.tavily.com/search'
USAGE_URL = 'https://api.tavily.com/usage'
CONFIG_FORMAT_VERSION = 2
# 401/403 临时冷却时间(秒):不再永久 disable,过期后自动重新探测 key

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/tavily_multi_key.py (reported line 13)May include surrounding context.

python
BASE = Path(__file__).resolve().parents[1]
CONFIG = BASE / 'config' / 'keys.json'
STATE = BASE / 'state' / 'quota.json'
SEARCH_URL = 'https://api.tavily.com/search'
USAGE_URL = 'https://api.tavily.com/usage'
CONFIG_FORMAT_VERSION = 2
# 401/403 临时冷却时间(秒):不再永久 disable,过期后自动重新探测 key

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/tavily_multi_key.py (reported line 14)May include surrounding context.

python
BASE = Path(__file__).resolve().parents[1]
CONFIG = BASE / 'config' / 'keys.json'
STATE = BASE / 'state' / 'quota.json'
SEARCH_URL = 'https://api.tavily.com/search'
USAGE_URL = 'https://api.tavily.com/usage'
CONFIG_FORMAT_VERSION = 2
# 401/403 临时冷却时间(秒):不再永久 disable,过期后自动重新探测 key

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file includes operationally meaningful instructions and user-facing quoted guidance in Chinese, such as the quoted user statements and terms used to explain required behavior. Because the skill does not state that Chinese is required or offer a language/locale choice, this can conflict with a language/locale policy that avoids forcing a specific language without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

These lines contain policy-relevant quoted guidance in Chinese that affects retry behavior and cooldown interpretation. Since the file does not give users or operators a language/locale choice, the instructions effectively assume comprehension of a specific language.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This section contains a mixed-language diagnostic procedure where key explanatory text is in Chinese. For a general-purpose skill document, embedding required operational instructions in a single language without opt-in or clear locale scoping can violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The document records a real search execution and ties the routed request to a specific named account and masked key index without placing a clear privacy warning adjacent to the example command. While not exposing full credentials, this creates unnecessary linkage between operational behavior, account identity, and search activity that could aid profiling or correlate usage if combined with other leaks.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.