T01 · Skill Instruction Hijacking
- Location
SKILL.md:218- Finding
Unbounded autonomous execution through instruction hijacking
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill matches its autonomous experiment purpose, but it needs review because it can repeatedly run commands, modify code indefinitely, and use destructive git resets that may erase work.
Install only if you are comfortable with an agent repeatedly editing code and running commands. Use it in a clean, disposable branch or worktree, set explicit experiment/time/cost limits, review any autoresearch.config.md before running, and avoid repositories with uncommitted work or accessible secrets.
SKILL.md:218Unbounded autonomous execution through instruction hijacking
SKILL.md:159Destructive Git rollback may erase unrelated repository changes
SKILL.md:6Overly broad tool permissions and repeated unrestricted command execution
The skill begins autonomous experimentation with shell execution, file edits, git branch creation, and destructive rollback behavior, but does not prominently warn about these actions in the user-facing description or setup flow. This omission is dangerous because users may invoke it without understanding that it can repeatedly execute commands and discard local changes via hard resets.
The instruction to use git reset --hard HEAD~1 after crashes introduces a destructive operation that can permanently discard working tree and index changes, including unrelated user work if repository state is not perfectly isolated. In an autonomous loop, repeated hard resets amplify the risk of accidental data loss and make recovery difficult.
Using git reset --hard HEAD~1 for non-improving experiments is a destructive rollback pattern that can erase changes beyond the intended experiment if the repository contains concurrent edits or the commit boundary does not fully isolate the modification. Because the skill is designed to run indefinitely, this creates sustained risk of repeated accidental loss of user data.
→ Log: "KEEP: <description> (<metric>: <old> → <new>)"
ELIF metric equal or worse:
→ DISCARD: git reset --hard HEAD~1
→ Log: "DISCARD: <description> (<metric>: <value> vs best <best>)"
ELIF crashed or timed out:
The crash-handling path again mandates git reset --hard HEAD~1, which is especially risky during failure scenarios because repository state may already be inconsistent or partially modified. Performing a hard reset under those conditions can destroy evidence needed for debugging and wipe unrelated local work.
→ Log: "DISCARD: (: vs best )"
ELIF crashed or timed out: → CRASH: git reset --hard HEAD~1 → Log: "CRASH: (error: )"
The trigger list includes broad terms like "optimize" and "run experiments," which can cause this skill to activate in many ordinary contexts where the user did not intend autonomous code modification and command execution. Because the skill is user-invocable and grants exec/write/git capabilities, overbroad triggering materially increases the chance of unsafe invocation.
No suspicious patterns detected.