Back to skill

Security audit

BrainHarness Autoresearch

Security checks for vulnerabilities and agentic risk

Overview

The skill is purpose-aligned, but it automatically rewrites skill prompts and sends prompt/eval content to third-party LLM APIs without enough explicit guardrails.

Install only if you are comfortable with an automated tool modifying target skill prompt files and sending prompt/eval content to third-party LLM APIs. Use it on copies or version-controlled skills, review diffs before keeping changes, avoid confidential prompts or eval data, and pin install commands where possible.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (24)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 52)May include surrounding context.

md
- `autoresearch.py` script in the skill directory

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The evaluation rule requires the output to contain Chinese keywords, effectively mandating Chinese output regardless of user preference. The skill description does not state that the tool is intentionally Chinese-only or provide any user opt-in or locale selection, which is a natural-language policy violation under the language/locale rule.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README explicitly describes an automated loop that reads a target prompt, mutates it repeatedly, and writes artifacts including a backup, but it does not clearly warn users that local files will be modified. In a prompt-optimization skill, that omission is more dangerous because users may aim it at real skill files and unintentionally overwrite or alter working prompts, causing integrity loss or workflow disruption.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill describes capabilities that require environment access, file reads/writes, and network use, but it does not declare any explicit tool scope or permission boundaries. This creates an overbroad execution surface where the host may grant more access than users expect, increasing the risk of prompt optimization runs reading arbitrary files, writing unintended outputs, or exfiltrating data through provider APIs.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest uses broad activation phrases like 'optimize skill', 'tune prompt', and 'improve skill X', which can match common editing or debugging requests not intended for autonomous experimentation. In context, this is more dangerous because the skill performs multi-step automated actions with file modification and network-backed LLM calls, so accidental invocation could cause unintended prompt changes, unnecessary API usage, or evaluation of the wrong target.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · SKILL.md (reported line 66)May include surrounding context.

md
| Parameter | Description | Example |
|-----------|-------------|---------|
| `--target` | Path to the skill directory or prompt file to optimize | `../workspace/skills/brain-search/SKILL.md` |
| `--evals` | Path to eval definition JSON file | `eval.json` |
| `--provider` | LLM provider for running experiments | `minimax` (default), `openai`, `anthropic` |
| `--runs` | Number of runs per experiment (statistical significance) | `5` (default) |

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · SKILL.md (reported line 153)May include surrounding context.

md
| Parameter | Description | Example |
|-----------|-------------|---------|
| `--target` | Path to the skill directory or prompt file to optimize | `../workspace/skills/brain-search/SKILL.md` |
| `--evals` | Path to eval definition JSON file | `eval.json` |
| `--provider` | LLM provider for running experiments | `minimax` (default), `openai`, `anthropic` |
| `--runs` | Number of runs per experiment (statistical significance) | `5` (default) |

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · SKILL.md (reported line 217)May include surrounding context.

md
| Parameter | Description | Example |
|-----------|-------------|---------|
| `--target` | Path to the skill directory or prompt file to optimize | `../workspace/skills/brain-search/SKILL.md` |
| `--evals` | Path to eval definition JSON file | `eval.json` |
| `--provider` | LLM provider for running experiments | `minimax` (default), `openai`, `anthropic` |
| `--runs` | Number of runs per experiment (statistical significance) | `5` (default) |

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · autoresearch.py (reported line 130)May include surrounding context.

python
"""MiniMax via Anthropic-compatible messages endpoint."""

    MODEL = "MiniMax-M2.7-highspeed"
    URL = "https://api.minimax.io/anthropic/v1/messages"

    def name(self) -> str:
        return "minimax"

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · autoresearch.py (reported line 163)May include surrounding context.

python
self.base_url = (
            os.environ.get("OPENAI_BASE_URL")
            or os.environ.get("OPENAI_API_BASE")
            or "https://api.openai.com/v1"
        ).rstrip("/")

    def name(self) -> str:

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · autoresearch.py (reported line 199)May include surrounding context.

python
"""Anthropic Messages API."""

    MODEL = "claude-sonnet-4-20250514"
    URL = "https://api.anthropic.com/v1/messages"

    def name(self) -> str:
        return "anthropic"

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The tool is presented as an optimizer/evaluator, but it directly overwrites the target SKILL.md throughout the experiment loop and leaves the best mutated version in place. In a prompt-optimization context this is expected behavior, but without strong disclosure and safer defaults it creates integrity risk by modifying user assets automatically and potentially preserving harmful or low-quality prompt changes.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code repeatedly writes LLM-generated content into the target file during experiments without explicit confirmation, dry-run mode, or a safer staging area. In this skill context, where untrusted model output is being iteratively produced, silent in-place modification increases the chance of accidental corruption, prompt injection persistence, or replacement of trusted prompt content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Each evaluation run transmits the system prompt content and test inputs to an external provider, which may include confidential prompts, customer data, or internal eval sets. Because the tool performs these calls repeatedly across many runs, the exposure is amplified and may be non-obvious to users who think they are only benchmarking locally.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The mutation prompt sends the full current skill, eval criteria, and sample failing outputs to an external LLM provider. That can expose proprietary prompts, embedded secrets, internal test cases, or sensitive model behavior data to third parties, which is especially relevant here because the tool is designed to process arbitrary skill content that may itself contain confidential instructions.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · autoresearch.py (reported line 1119)May include surrounding context.

python
formatter_class=argparse.RawDescriptionHelpFormatter,
        epilog=(
            "Examples:\n"
            "  python autoresearch.py --target skills/search/SKILL.md --evals eval.json\n"
            "  python autoresearch.py --target SKILL.md --evals eval.json --runs 3 --max-experiments 10 --dashboard\n"
            "  python autoresearch.py --target SKILL.md --evals eval.json --provider openai --verbose\n"
        ),

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The banned_phrases example explicitly blocks Chinese phrases alongside generic filler, and later examples repeat Chinese-specific banned phrases. In a general eval-writing guide, this can encourage skills to suppress a particular language by default rather than based on user choice or a documented locale requirement.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

This JSON skill file is in scope for vague-trigger review. The required output keywords are fixed to Chinese terms like "进度", "风险", and "下周", but the file does not clarify whether the skill is intended only for Chinese inputs/outputs or under what conditions this language constraint applies, which makes activation/use conditions insufficiently specific.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The top-level docstring asserts zero external dependencies, suggesting fully self-contained operation. However, when dashboard generation is enabled, the produced HTML relies on a remote Chart.js CDN URL, contradicting that self-contained/no-external-dependency claim.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The generated dashboard loads Chart.js from a public CDN when the HTML file is opened, creating an unexpected third-party network dependency and metadata leak. This is low severity because it is limited to the optional dashboard feature, but it can violate offline, privacy, or supply-chain expectations in sensitive environments.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
71% confidence
Finding

The translation eval example recommends banning source-language fragments, which can be reasonable in some contexts but is presented as a universal pattern without noting exceptions for bilingual or user-requested mixed-language output. That creates a risk of enforcing a language policy without explicit user opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.