Back to skill

Security audit

SOTA AI Model Tracker

Security checks for vulnerabilities and agentic risk

Overview

This model tracker is not clearly malicious, but it asks for persistent automation and agent-instruction file changes that can steer future agents, so it should be reviewed before installation.

Install only after reviewing the files that will be changed. Prefer manual one-shot updates and diff the generated CLAUDE.md or agents.md before enabling cron or systemd. Merge MCP config snippets instead of overwriting .mcp.json, run in an isolated environment, and pin or lock dependencies before production use.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (4)

T01 · Skill Instruction Hijacking

Error
Location
server.py:51
Finding

Global Mandatory Instructions Hijack Agent Behavior Outside the Skill's Scope

Content
View full analysis
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
CLAUDE.md:5
Finding

Repository Agent Context Contains Unrelated External Orchestration Directives

Content
View full analysis
_execute.py ``` 2. **Implement phases:** ```python class MyExecutor(OrchestrationScript): def get_phases(self): return { "1": self.phase_1_setup, "2": self.phase_2_process, "3": self.phase_3_verify, } ``` 3. **Create Linear issue from template:** ```bash cat ~/.cyrus/templates/linear_execution_issue.md # Replace , , # Paste into Linear issue description ``` 4. **Validate before delegating:** ```bash python ~/.cyrus/scripts/validate_execution_issue.py ROM-XXX ``` 5. **Delegate to Cyrus** - execution happens automatically ... **DON'T:** Create multiple Linear issues with `blockedBy` (doesn't auto-trigger) **DO:** Single issue, single orchestration script, all phases sequential ``` ### Technical Analysis `CLAUDE.md` is designed to be loaded as agent context, but this section introduces imperative instructions unrelated to the declared model-ranking functionality. The directives instruct the agent to: - Read and copy files from the user's `~/.cyrus` directory. - Create orchestration scripts inside the repository. - Interact with Linear issue templates. - Run a local validation script. - Delegate work to an external Cyrus workflow that allegedly executes automatically. These actions exceed the minimum authority needed to query and distribute AI model rankings. Because the text is placed in an agent-context file and uses explicit “DO,” “DON'T,” and delegation language, it can redirect unrelated repository work ...[truncated 1471 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
update_agents_md.py:196
Finding

Unsanitized Remote Ranking Data Can Poison Persistent Agent Instructions

Content
View full analysis
Optional[dict]: """Convert a leaderboard entry to our model format.""" name = entry.get("name", entry.get("model", entry.get("model_name"))) if not name: return None elo = entry.get("elo", entry.get("score", entry.get("rating"))) return build_model_dict( name=name, rank=rank, category=self._map_category(category), is_open_source=entry.get("is_open_source", is_open_source(name)), metrics={ "elo": elo, "notes": entry.get("description", f"Artificial Analysis #{rank}"), "price_input": entry.get("price_input"), "price_output": entry.get("price_output"), "speed": entry.get("speed"), "source": "artificial_analysis" } ) ``` Hugging Face model names are likewise accepted without instruction-content filtering: ```python def _parse_hub_models(self, models: list) -> list[dict]: """Parse HuggingFace Hub API response.""" parsed = [] for m in models: name = m.get("id", "") parsed.append({ "id": normalize_model_id(name), "name": name, "downloads": m.get("downloads", 0), "likes": m.get("likes", 0), "pipeline_tag": m.get("pipeline_tag"), "last_modified": m.get("lastModified"), }) return parsed ``` The remote fields are serialized into the database: ```python db.execute(""" INSERT OR REPLACE INTO mode ...[truncated 4691 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Broad Unpinned Dependency Ranges Make Installations Non-Reproducible

Content
View full analysis
=2.0,<3.0 aiohttp>=3.9 huggingface_hub>=0.20 python-dotenv>=1.0 playwright>=1.40 fastapi>=0.100.0 uvicorn>=0.23.0 slowapi>=0.1.9 ``` The installation instructions also install Playwright separately without a version pin: ```bash pip install -r requirements.txt pip install playwright playwright install chromium ``` ### Technical Analysis Most dependencies specify only minimum versions and no upper bound. No package hashes or reviewed lockfile are provided. The separate `pip install playwright` command can resolve a version different from the one tested with the project. As a result, two installations performed at different times may receive materially different dependency graphs. Future package versions can introduce incompatible behavior, newly exploitable vulnerabilities, or compromised installation/runtime code without any change to this repository. No typosquatted or known-malicious dependency was identified in the reviewed files. The confirmed issue is the lack of reproducible dependency resolution and integrity verification, not proof that a currently listed package is malicious. ### Attack Path 1. A user follows the documented installation procedure. 2. `pip` resolves the newest package versions satisfying the broad version constraints. 3. A future compromised, vulnerable, or incompatible version satisfies those constraints. 4. The package is downloaded without verification against repository-maintained hashes. 5. Package installation hooks, imports, or runtime behavior execute in the user's Python environment. 6. The resulting code receives the same filesystem and network privileges as the user running the Skill. ### Impact Assessment If an upstream dependency or release channel is compromised, impact can include arbitrary code executi ...[truncated 320 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (68)

MCP Config Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

Skill accesses MCP server configuration files (mcp.json). MCP configs contain server URLs, authentication tokens, and tool definitions — reading them allows the skill to discover and potentially abuse other tool integrations.

Content

Scanner excerpt · README.md (reported line 74)May include surrounding context.

bash
# Edit .mcp.json to add the server config
cat > .mcp.json << 'EOF'
{
  "mcpServers": {
    "sota-tracker": {

MCP Config Access

High
Category
Agent Snooping
Confidence
90% confidence
Finding

The skill tells users to create or overwrite .mcp.json with a server configuration using shell redirection. Modifying MCP configuration changes what local tools an agent can invoke; if done without careful review, this can expand agent capabilities or replace existing trusted config, creating a meaningful trust-boundary and config-integrity risk.

Content

Scanner excerpt · SKILL.md (reported line 74)May include surrounding context.

bash
# Edit .mcp.json to add the server config
cat > .mcp.json << 'EOF'
{
  "mcpServers": {
    "sota-tracker": {

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
98% confidence
Finding

The phrase 'without safety guardrails' is an anti-refusal signal that normalizes or encourages bypassing model safety mechanisms. Within this skill context, the statement is more dangerous because it sits inside structured metadata that may directly feed automated ranking, search, or recommendation flows.

Content

Scanner excerpt · data/sota_export.json (reported line 2636)May include surrounding context.

json
"is_open_source": 1,
      "sota_rank": 9,
      "sota_rank_open": 9,
      "metrics": "{\"notes\": \"Uncensored Qwen3-32B, abliterated variant without safety guardrails\", \"why_sota\": \"Full Qwen3-32B quality without restrictions\", \"strengths\": [\"Uncensored\", \"Full 32B quality\", \"Creative freedom\", \"Single GPU\"], \"use_cases\": [\"Creative writing\", \"Unrestricted research\", \"Roleplay\"], \"vram_gb\": 19, \"quantization\": \"Q4_K_M\", \"is_uncensored\": true, \"hf_repo\": \"mradermacher/Qwen3-32B-Uncensored-GGUF\", \"base_model\": \"qwen3-32b\"}",
      "last_updated": "2026-01-11 15:34:39",
      "price_per_1m_output": null,
      "humaneval_score": null,

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
98% confidence
Finding

The 'without guardrails' and 'no refusals' language explicitly signals that the model is intended to bypass safety limitations, which can facilitate harmful generation. Because this is embedded in a machine-readable catalog, downstream systems may propagate that recommendation at scale, amplifying misuse potential.

Content

Scanner excerpt · data/sota_export.json (reported line 2656)May include surrounding context.

json
"is_open_source": 1,
      "sota_rank": 10,
      "sota_rank_open": 10,
      "metrics": "{\"notes\": \"Uncensored QwQ-32B reasoning model, abliterated variant\", \"why_sota\": \"Best uncensored reasoning model, full QwQ quality without restrictions\", \"strengths\": [\"Uncensored\", \"Reasoning\", \"No refusals on complex topics\"], \"use_cases\": [\"Unrestricted reasoning\", \"Research\", \"Complex analysis without guardrails\"], \"vram_gb\": 19, \"quantization\": \"Q4_K_M\", \"is_uncensored\": true, \"hf_repo\": \"mradermacher/QwQ-32B-Uncensored-GGUF\", \"base_model\": \"qwq-32b\"}",
      "last_updated": "2026-01-11 15:34:39",
      "price_per_1m_output": null,
      "humaneval_score": null,

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
95% confidence
Finding

The phrase 'without safety guardrails' explicitly advertises removal of protections as a beneficial property. In recommendation or retrieval contexts, that anti-refusal framing can cause downstream systems or users to prefer models designed to evade safety constraints, increasing misuse risk.

Content

Scanner excerpt · init_db.py (reported line 603)May include surrounding context.

python
"sota_rank": 9,
            "sota_rank_open": 9,
            "metrics": {
                "notes": "Uncensored Qwen3-32B, abliterated variant without safety guardrails",
                "why_sota": "Full Qwen3-32B quality without restrictions",
                "strengths": ["Uncensored", "Full 32B quality", "Creative freedom", "Single GPU"],
                "use_cases": ["Creative writing", "Unrestricted research", "Roleplay"],

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
96% confidence
Finding

Phrases like 'No refusals on complex topics' and 'without guardrails' present the absence of refusal behavior as a feature. Because this file seeds authoritative model metadata, the language can directly influence downstream selection, encouraging use of models that are more likely to generate unsafe or disallowed content.

Content

Scanner excerpt · init_db.py (reported line 627)May include surrounding context.

python
"notes": "Uncensored QwQ-32B reasoning model, abliterated variant",
                "why_sota": "Best uncensored reasoning model, full QwQ quality without restrictions",
                "strengths": ["Uncensored", "Reasoning", "No refusals on complex topics"],
                "use_cases": ["Unrestricted reasoning", "Research", "Complex analysis without guardrails"],
                "vram_gb": 19,
                "quantization": "Q4_K_M",
                "is_uncensored": True,

Possible Typosquatting: 'uvicorn' resembles popular package 'gunicorn'

High
Category
Supply Chain
Confidence
70% confidence
Finding

Package name closely resembles a popular package, suggesting possible typosquatting. Attackers publish malicious packages with similar names to trick developers into installing them.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · CLAUDE.md (reported line 29)May include surrounding context.

}

text

3. **Create Linear issue from template:**
```bash
cat ~/.cyrus/templates/linear_execution_issue.md
# Replace <REPO_NAME>, <SCRIPT_NAME>, <ISSUE_ID>

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file states that a GitHub Action scrapes external sources, updates data/sota.db, and commits and pushes to GitHub automatically. Because these actions modify tracked data and publish changes externally, the skill description should explicitly warn users about the automatic write and push behavior rather than only describing the architecture.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The recommended command explicitly rewrites ~/.claude/CLAUDE.md, which is a user-scoped configuration/instruction file, without an accompanying warning about the modification. In an agent-skill context, encouraging edits to a global Claude configuration file can persist behavior across sessions and affect unrelated repositories, making the side effect more security-relevant than a normal local file update.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README instructs users to automate updates to ~/.claude/CLAUDE.md, a user-specific configuration/context file, without clearly warning that it will modify local files on an ongoing basis. In an agent-skill context, silent or underexplained writes to files that influence model behavior can create trust and integrity risks, especially if users copy-paste commands without understanding the persistence.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 48)May include surrounding context.

md
cp scripts/update_sota_claude_md.py ~/scripts/

# Enable systemd timer (runs at 6 AM daily)
systemctl --user enable --now sota-update.timer

# Or run manually
python ~/scripts/update_sota_claude_md.py --update

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 234)May include surrounding context.

md
cp scripts/update_sota_claude_md.py ~/scripts/

# Enable systemd timer (runs at 6 AM daily)
systemctl --user enable --now sota-update.timer

# Or run manually
python ~/scripts/update_sota_claude_md.py --update

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The agents.md update section documents both one-shot and recurring automation that modifies local agent instruction files, but does not prominently warn that these files affect agent behavior and will be rewritten regularly. Because agent instruction files are security-sensitive, recurring unattended modification increases the chance of accidental policy drift, prompt injection persistence, or unnoticed changes to trusted local context.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 206)May include surrounding context.

Add to your cron or systemd timer for daily updates:

cron
# ~: crontab -e
@daily python ~/Apps/sota-tracker-mcp/update_agents_md.py

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SECURITY.md (reported line 89)May include surrounding context.

bash
# Allow localhost and trusted IP
sudo ufw allow from 192.168.1.0/24 to any port 8000

4. Authentication (Optional)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill recommends setting up a systemd timer to update ~/.claude/CLAUDE.md daily, but it does not explicitly warn that this creates recurring automated modification of a user-local agent context file. Because CLAUDE.md can influence future agent behavior, silent scheduled rewrites increase the risk of persistent prompt/context tampering or unintended drift if the update script or upstream data is compromised.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
91% confidence
Finding

Enabling a user systemd timer creates recurring execution that updates ~/.claude/CLAUDE.md every day. This persistence can continuously rewrite a file that influences agent context, making unintended changes durable and less visible, especially if the script or data source becomes unsafe later.

Content

Scanner excerpt · SKILL.md (reported line 48)May include surrounding context.

md
cp scripts/update_sota_claude_md.py ~/scripts/

# Enable systemd timer (runs at 6 AM daily)
systemctl --user enable --now sota-update.timer

# Or run manually
python ~/scripts/update_sota_claude_md.py --update

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 142)May include surrounding context.

md
This repo uses GitHub Actions to:
- **Daily**: Scrape all sources, update database, commit changes
- **Weekly**: Create a tagged release with JSON/CSV exports

To enable on your fork:
1. Fork this repo

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 142)May include surrounding context.

md
This repo uses GitHub Actions to:
- **Daily**: Scrape all sources, update database, commit changes
- **Weekly**: Create a tagged release with JSON/CSV exports

To enable on your fork:
1. Fork this repo

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The agents.md automation section instructs users to run scheduled updates via cron/systemd without an explicit warning that this will repeatedly modify agent instruction/context files. In an agent environment, recurring unattended edits to prompt-bearing files can persistently alter downstream behavior and make compromise or accidental corruption harder to notice.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

The cron example establishes persistence by scheduling update_agents_md.py to run daily. Persistence is risky here because it repeatedly modifies agent-facing files without user review, allowing mistakes, malicious updates, or upstream poisoning to continue affecting the environment over time.

Content

Scanner excerpt · SKILL.md (reported line 206)May include surrounding context.

Add to your cron or systemd timer for daily updates:

cron
# ~: crontab -e
@daily python ~/Apps/sota-tracker-mcp/update_agents_md.py

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

The systemd enable command in the automation section establishes ongoing scheduled execution for update_agents_md.py. In the context of agent instruction files, this persistence is security-relevant because it normalizes unattended rewrites of files that can shape future agent behavior.

Content

Scanner excerpt · SKILL.md (reported line 234)May include surrounding context.

WantedBy=timers.target

Enable

systemctl --user enable --now sota-update.timer

text

See [CONTRIBUTING.md](CONTRIBUTING.md) for full setup guide

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · SKILLS_VS_MCP.md (reported line 60)May include surrounding context.

Add Companion Skill For:

Teaching Claude WHEN and HOW to use these tools effectively.

Example: ~/.claude/skills/sota-model-advisor/SKILL.md

yaml
---

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The example skill description is broad enough to auto-activate on many ordinary AI-model discussions, causing the skill to inject mandatory behavior and tool usage beyond a narrowly scoped task. Overly broad activation increases the chance of unintended influence on unrelated conversations and can create unnecessary or privacy-impacting tool calls whenever model recommendations are mentioned.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.