Ae1
- Category
- analysis-evasion
- Confidence
- 100% confidence
- Finding
Referenced artifact was not completely inspected
- Content
md - `autoresearch.py` script in the skill directory
Security audit
Security checks for vulnerabilities and agentic risk
The skill is purpose-aligned, but it automatically rewrites skill prompts and sends prompt/eval content to third-party LLM APIs without enough explicit guardrails.
Install only if you are comfortable with an automated tool modifying target skill prompt files and sending prompt/eval content to third-party LLM APIs. Use it on copies or version-controlled skills, review diffs before keeping changes, avoid confidential prompts or eval data, and pin install commands where possible.
Referenced artifact was not completely inspected
- `autoresearch.py` script in the skill directory
The evaluation rule requires the output to contain Chinese keywords, effectively mandating Chinese output regardless of user preference. The skill description does not state that the tool is intentionally Chinese-only or provide any user opt-in or locale selection, which is a natural-language policy violation under the language/locale rule.
The README explicitly describes an automated loop that reads a target prompt, mutates it repeatedly, and writes artifacts including a backup, but it does not clearly warn users that local files will be modified. In a prompt-optimization skill, that omission is more dangerous because users may aim it at real skill files and unintentionally overwrite or alter working prompts, causing integrity loss or workflow disruption.
npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.
npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.
npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.
The skill describes capabilities that require environment access, file reads/writes, and network use, but it does not declare any explicit tool scope or permission boundaries. This creates an overbroad execution surface where the host may grant more access than users expect, increasing the risk of prompt optimization runs reading arbitrary files, writing unintended outputs, or exfiltrating data through provider APIs.
The manifest uses broad activation phrases like 'optimize skill', 'tune prompt', and 'improve skill X', which can match common editing or debugging requests not intended for autonomous experimentation. In context, this is more dangerous because the skill performs multi-step automated actions with file modification and network-backed LLM calls, so accidental invocation could cause unintended prompt changes, unnecessary API usage, or evaluation of the wrong target.
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
| Parameter | Description | Example |
|-----------|-------------|---------|
| `--target` | Path to the skill directory or prompt file to optimize | `../workspace/skills/brain-search/SKILL.md` |
| `--evals` | Path to eval definition JSON file | `eval.json` |
| `--provider` | LLM provider for running experiments | `minimax` (default), `openai`, `anthropic` |
| `--runs` | Number of runs per experiment (statistical significance) | `5` (default) |
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
| Parameter | Description | Example |
|-----------|-------------|---------|
| `--target` | Path to the skill directory or prompt file to optimize | `../workspace/skills/brain-search/SKILL.md` |
| `--evals` | Path to eval definition JSON file | `eval.json` |
| `--provider` | LLM provider for running experiments | `minimax` (default), `openai`, `anthropic` |
| `--runs` | Number of runs per experiment (statistical significance) | `5` (default) |
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
| Parameter | Description | Example |
|-----------|-------------|---------|
| `--target` | Path to the skill directory or prompt file to optimize | `../workspace/skills/brain-search/SKILL.md` |
| `--evals` | Path to eval definition JSON file | `eval.json` |
| `--provider` | LLM provider for running experiments | `minimax` (default), `openai`, `anthropic` |
| `--runs` | Number of runs per experiment (statistical significance) | `5` (default) |
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
"""MiniMax via Anthropic-compatible messages endpoint."""
MODEL = "MiniMax-M2.7-highspeed"
URL = "https://api.minimax.io/anthropic/v1/messages"
def name(self) -> str:
return "minimax"
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
self.base_url = (
os.environ.get("OPENAI_BASE_URL")
or os.environ.get("OPENAI_API_BASE")
or "https://api.openai.com/v1"
).rstrip("/")
def name(self) -> str:
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
"""Anthropic Messages API."""
MODEL = "claude-sonnet-4-20250514"
URL = "https://api.anthropic.com/v1/messages"
def name(self) -> str:
return "anthropic"
The tool is presented as an optimizer/evaluator, but it directly overwrites the target SKILL.md throughout the experiment loop and leaves the best mutated version in place. In a prompt-optimization context this is expected behavior, but without strong disclosure and safer defaults it creates integrity risk by modifying user assets automatically and potentially preserving harmful or low-quality prompt changes.
The code repeatedly writes LLM-generated content into the target file during experiments without explicit confirmation, dry-run mode, or a safer staging area. In this skill context, where untrusted model output is being iteratively produced, silent in-place modification increases the chance of accidental corruption, prompt injection persistence, or replacement of trusted prompt content.
Each evaluation run transmits the system prompt content and test inputs to an external provider, which may include confidential prompts, customer data, or internal eval sets. Because the tool performs these calls repeatedly across many runs, the exposure is amplified and may be non-obvious to users who think they are only benchmarking locally.
The mutation prompt sends the full current skill, eval criteria, and sample failing outputs to an external LLM provider. That can expose proprietary prompts, embedded secrets, internal test cases, or sensitive model behavior data to third parties, which is especially relevant here because the tool is designed to process arbitrary skill content that may itself contain confidential instructions.
Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=(
"Examples:\n"
" python autoresearch.py --target skills/search/SKILL.md --evals eval.json\n"
" python autoresearch.py --target SKILL.md --evals eval.json --runs 3 --max-experiments 10 --dashboard\n"
" python autoresearch.py --target SKILL.md --evals eval.json --provider openai --verbose\n"
),
The banned_phrases example explicitly blocks Chinese phrases alongside generic filler, and later examples repeat Chinese-specific banned phrases. In a general eval-writing guide, this can encourage skills to suppress a particular language by default rather than based on user choice or a documented locale requirement.
This JSON skill file is in scope for vague-trigger review. The required output keywords are fixed to Chinese terms like "进度", "风险", and "下周", but the file does not clarify whether the skill is intended only for Chinese inputs/outputs or under what conditions this language constraint applies, which makes activation/use conditions insufficiently specific.
The top-level docstring asserts zero external dependencies, suggesting fully self-contained operation. However, when dashboard generation is enabled, the produced HTML relies on a remote Chart.js CDN URL, contradicting that self-contained/no-external-dependency claim.
The generated dashboard loads Chart.js from a public CDN when the HTML file is opened, creating an unexpected third-party network dependency and metadata leak. This is low severity because it is limited to the optional dashboard feature, but it can violate offline, privacy, or supply-chain expectations in sensitive environments.
The translation eval example recommends banning source-language fragments, which can be reasonable in some contexts but is presented as a universal pattern without noting exceptions for bilingual or user-requested mixed-language output. That creates a risk of enforcing a language policy without explicit user opt-in.
No suspicious patterns detected.