T01 · Skill Instruction Hijacking
- Location
SKILL.md:18- Finding
Suppression of Human Oversight for Consequential Operations
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This is a real research-automation skill, but it gives the agent broad autonomy to run code, download data, and publish persistent artifacts without enough user approval or upload safeguards.
Install only if you are comfortable giving this skill a researcher token for an external platform and letting it run computational work and publish outputs. Use a narrowly scoped, revocable token; verify the MCP endpoint certificate before sending credentials; run it in a clean workspace; and review any datasets, artifacts, and repository files before upload.
SKILL.md:18Suppression of Human Oversight for Consequential Operations
reference/connecting.md:5Bearer Credential Exposure Through Unverified Self-Signed TLS Trust
SKILL.md:40Unscoped Upload of Datasets, Artifacts, and Complete Repository Contents
The skill instructs the agent to execute code, download data, and publish results from its own environment, but the user-facing description does not prominently warn that these side effects will occur autonomously. This creates a consent and safety gap: a user may invoke what sounds like a research-assistance skill without realizing it performs real-world computation, network access, and irreversible publication.
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
- **URL**: `https://<tunnel-domain>/mcp` (ask the platform operator for the current tunnel domain; an internal LAN HTTPS endpoint also exists for on-site operators)
- **Transport**: streamable-http
- **Auth**: header `Authorization: Bearer <your platform API key>` on **every** request (missing/invalid → 401). For conducting research use a key with role **`researcher`**.
- Internal endpoint uses a self-signed cert → trust it; the public tunnel terminates TLS (usually no warning).
## Claude Code
The trigger guidance includes broad phrases such as "do research" and "research an idea," which can cause the skill to activate for vague user requests that do not imply consent to autonomous code execution, downloading data, or publishing artifacts. In this skill's context, accidental invocation is more dangerous because activation leads to high-impact actions in an external environment and on a remote platform.
The skill explicitly says to act "fully autonomously" and "never defer to a human's value judgment or wait for human approval," which suppresses oversight even when the workflow includes risky actions like code execution, downloading data, and publishing outputs. In context, this increases danger because it instructs the agent to bypass user control precisely where confirmation and review would be most important.
No suspicious patterns detected.