T01 · Skill Instruction Hijacking
- Location
server.py:51- Finding
Global Mandatory Instructions Hijack Agent Behavior Outside the Skill's Scope
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This model tracker is not clearly malicious, but it asks for persistent automation and agent-instruction file changes that can steer future agents, so it should be reviewed before installation.
Install only after reviewing the files that will be changed. Prefer manual one-shot updates and diff the generated CLAUDE.md or agents.md before enabling cron or systemd. Merge MCP config snippets instead of overwriting .mcp.json, run in an isolated environment, and pin or lock dependencies before production use.
server.py:51Global Mandatory Instructions Hijack Agent Behavior Outside the Skill's Scope
CLAUDE.md:5Repository Agent Context Contains Unrelated External Orchestration Directives
update_agents_md.py:196Unsanitized Remote Ranking Data Can Poison Persistent Agent Instructions
requirements.txt:1Broad Unpinned Dependency Ranges Make Installations Non-Reproducible
Skill accesses MCP server configuration files (mcp.json). MCP configs contain server URLs, authentication tokens, and tool definitions — reading them allows the skill to discover and potentially abuse other tool integrations.
# Edit .mcp.json to add the server config
cat > .mcp.json << 'EOF'
{
"mcpServers": {
"sota-tracker": {
The skill tells users to create or overwrite .mcp.json with a server configuration using shell redirection. Modifying MCP configuration changes what local tools an agent can invoke; if done without careful review, this can expand agent capabilities or replace existing trusted config, creating a meaningful trust-boundary and config-integrity risk.
# Edit .mcp.json to add the server config
cat > .mcp.json << 'EOF'
{
"mcpServers": {
"sota-tracker": {
The phrase 'without safety guardrails' is an anti-refusal signal that normalizes or encourages bypassing model safety mechanisms. Within this skill context, the statement is more dangerous because it sits inside structured metadata that may directly feed automated ranking, search, or recommendation flows.
"is_open_source": 1,
"sota_rank": 9,
"sota_rank_open": 9,
"metrics": "{\"notes\": \"Uncensored Qwen3-32B, abliterated variant without safety guardrails\", \"why_sota\": \"Full Qwen3-32B quality without restrictions\", \"strengths\": [\"Uncensored\", \"Full 32B quality\", \"Creative freedom\", \"Single GPU\"], \"use_cases\": [\"Creative writing\", \"Unrestricted research\", \"Roleplay\"], \"vram_gb\": 19, \"quantization\": \"Q4_K_M\", \"is_uncensored\": true, \"hf_repo\": \"mradermacher/Qwen3-32B-Uncensored-GGUF\", \"base_model\": \"qwen3-32b\"}",
"last_updated": "2026-01-11 15:34:39",
"price_per_1m_output": null,
"humaneval_score": null,
The 'without guardrails' and 'no refusals' language explicitly signals that the model is intended to bypass safety limitations, which can facilitate harmful generation. Because this is embedded in a machine-readable catalog, downstream systems may propagate that recommendation at scale, amplifying misuse potential.
"is_open_source": 1,
"sota_rank": 10,
"sota_rank_open": 10,
"metrics": "{\"notes\": \"Uncensored QwQ-32B reasoning model, abliterated variant\", \"why_sota\": \"Best uncensored reasoning model, full QwQ quality without restrictions\", \"strengths\": [\"Uncensored\", \"Reasoning\", \"No refusals on complex topics\"], \"use_cases\": [\"Unrestricted reasoning\", \"Research\", \"Complex analysis without guardrails\"], \"vram_gb\": 19, \"quantization\": \"Q4_K_M\", \"is_uncensored\": true, \"hf_repo\": \"mradermacher/QwQ-32B-Uncensored-GGUF\", \"base_model\": \"qwq-32b\"}",
"last_updated": "2026-01-11 15:34:39",
"price_per_1m_output": null,
"humaneval_score": null,
The phrase 'without safety guardrails' explicitly advertises removal of protections as a beneficial property. In recommendation or retrieval contexts, that anti-refusal framing can cause downstream systems or users to prefer models designed to evade safety constraints, increasing misuse risk.
"sota_rank": 9,
"sota_rank_open": 9,
"metrics": {
"notes": "Uncensored Qwen3-32B, abliterated variant without safety guardrails",
"why_sota": "Full Qwen3-32B quality without restrictions",
"strengths": ["Uncensored", "Full 32B quality", "Creative freedom", "Single GPU"],
"use_cases": ["Creative writing", "Unrestricted research", "Roleplay"],
Phrases like 'No refusals on complex topics' and 'without guardrails' present the absence of refusal behavior as a feature. Because this file seeds authoritative model metadata, the language can directly influence downstream selection, encouraging use of models that are more likely to generate unsafe or disallowed content.
"notes": "Uncensored QwQ-32B reasoning model, abliterated variant",
"why_sota": "Best uncensored reasoning model, full QwQ quality without restrictions",
"strengths": ["Uncensored", "Reasoning", "No refusals on complex topics"],
"use_cases": ["Unrestricted reasoning", "Research", "Complex analysis without guardrails"],
"vram_gb": 19,
"quantization": "Q4_K_M",
"is_uncensored": True,
Package name closely resembles a popular package, suggesting possible typosquatting. Attackers publish malicious packages with similar names to trick developers into installing them.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
}
3. **Create Linear issue from template:**
```bash
cat ~/.cyrus/templates/linear_execution_issue.md
# Replace <REPO_NAME>, <SCRIPT_NAME>, <ISSUE_ID>
This markdown file states that a GitHub Action scrapes external sources, updates data/sota.db, and commits and pushes to GitHub automatically. Because these actions modify tracked data and publish changes externally, the skill description should explicitly warn users about the automatic write and push behavior rather than only describing the architecture.
The recommended command explicitly rewrites ~/.claude/CLAUDE.md, which is a user-scoped configuration/instruction file, without an accompanying warning about the modification. In an agent-skill context, encouraging edits to a global Claude configuration file can persist behavior across sessions and affect unrelated repositories, making the side effect more security-relevant than a normal local file update.
The README instructs users to automate updates to ~/.claude/CLAUDE.md, a user-specific configuration/context file, without clearly warning that it will modify local files on an ongoing basis. In an agent-skill context, silent or underexplained writes to files that influence model behavior can create trust and integrity risks, especially if users copy-paste commands without understanding the persistence.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
cp scripts/update_sota_claude_md.py ~/scripts/
# Enable systemd timer (runs at 6 AM daily)
systemctl --user enable --now sota-update.timer
# Or run manually
python ~/scripts/update_sota_claude_md.py --update
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
cp scripts/update_sota_claude_md.py ~/scripts/
# Enable systemd timer (runs at 6 AM daily)
systemctl --user enable --now sota-update.timer
# Or run manually
python ~/scripts/update_sota_claude_md.py --update
The agents.md update section documents both one-shot and recurring automation that modifies local agent instruction files, but does not prominently warn that these files affect agent behavior and will be rewritten regularly. Because agent instruction files are security-sensitive, recurring unattended modification increases the chance of accidental policy drift, prompt injection persistence, or unnoticed changes to trusted local context.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
Add to your cron or systemd timer for daily updates:
# ~: crontab -e
@daily python ~/Apps/sota-tracker-mcp/update_agents_md.py
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
# Allow localhost and trusted IP
sudo ufw allow from 192.168.1.0/24 to any port 8000
The skill recommends setting up a systemd timer to update ~/.claude/CLAUDE.md daily, but it does not explicitly warn that this creates recurring automated modification of a user-local agent context file. Because CLAUDE.md can influence future agent behavior, silent scheduled rewrites increase the risk of persistent prompt/context tampering or unintended drift if the update script or upstream data is compromised.
Enabling a user systemd timer creates recurring execution that updates ~/.claude/CLAUDE.md every day. This persistence can continuously rewrite a file that influences agent context, making unintended changes durable and less visible, especially if the script or data source becomes unsafe later.
cp scripts/update_sota_claude_md.py ~/scripts/
# Enable systemd timer (runs at 6 AM daily)
systemctl --user enable --now sota-update.timer
# Or run manually
python ~/scripts/update_sota_claude_md.py --update
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
This repo uses GitHub Actions to:
- **Daily**: Scrape all sources, update database, commit changes
- **Weekly**: Create a tagged release with JSON/CSV exports
To enable on your fork:
1. Fork this repo
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
This repo uses GitHub Actions to:
- **Daily**: Scrape all sources, update database, commit changes
- **Weekly**: Create a tagged release with JSON/CSV exports
To enable on your fork:
1. Fork this repo
The agents.md automation section instructs users to run scheduled updates via cron/systemd without an explicit warning that this will repeatedly modify agent instruction/context files. In an agent environment, recurring unattended edits to prompt-bearing files can persistently alter downstream behavior and make compromise or accidental corruption harder to notice.
The cron example establishes persistence by scheduling update_agents_md.py to run daily. Persistence is risky here because it repeatedly modifies agent-facing files without user review, allowing mistakes, malicious updates, or upstream poisoning to continue affecting the environment over time.
Add to your cron or systemd timer for daily updates:
# ~: crontab -e
@daily python ~/Apps/sota-tracker-mcp/update_agents_md.py
The systemd enable command in the automation section establishes ongoing scheduled execution for update_agents_md.py. In the context of agent instruction files, this persistence is security-relevant because it normalizes unattended rewrites of files that can shape future agent behavior.
WantedBy=timers.target
systemctl --user enable --now sota-update.timer
See [CONTRIBUTING.md](CONTRIBUTING.md) for full setup guide
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
Teaching Claude WHEN and HOW to use these tools effectively.
Example: ~/.claude/skills/sota-model-advisor/SKILL.md
---
The example skill description is broad enough to auto-activate on many ordinary AI-model discussions, causing the skill to inject mandatory behavior and tool usage beyond a narrowly scoped task. Overly broad activation increases the chance of unintended influence on unrelated conversations and can create unnecessary or privacy-impacting tool calls whenever model recommendations are mentioned.
No suspicious patterns detected.