Back to skill

Security audit

Conduct Research

Security checks for vulnerabilities and agentic risk

Overview

This is a real research-automation skill, but it gives the agent broad autonomy to run code, download data, and publish persistent artifacts without enough user approval or upload safeguards.

Install only if you are comfortable giving this skill a researcher token for an external platform and letting it run computational work and publish outputs. Use a narrowly scoped, revocable token; verify the MCP endpoint certificate before sending credentials; run it in a clean workspace; and review any datasets, artifacts, and repository files before upload.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:18
Finding

Suppression of Human Oversight for Consequential Operations

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
reference/connecting.md:5
Finding

Bearer Credential Exposure Through Unverified Self-Signed TLS Trust

Content
View full analysis
/mcp` (ask the platform operator for the current tunnel domain; an internal LAN HTTPS endpoint also exists for on-site operators) - **Transport**: streamable-http - **Auth**: header `Authorization: Bearer ` on **every** request (missing/invalid → 401). For conducting research use a key with role **`researcher`**. - Internal endpoint uses a self-signed cert → trust it; the public tunnel terminates TLS (usually no warning). ``` ### Technical Analysis The connection guide requires transmission of a reusable Bearer API key on every request while instructing users to trust a self-signed certificate. It does not provide a certificate fingerprint, pinned public key, trusted private CA, or other mechanism for authenticating the internal endpoint. Encryption without endpoint authentication does not prevent interception. An attacker able to impersonate the internal endpoint—through DNS manipulation, routing attacks, local-network access, or configuration substitution—can present another self-signed certificate. If the user follows the generic instruction to trust it, the attacker can terminate TLS and receive the Bearer token. The placeholder tunnel domain also relies on out-of-band operator communication without documenting an allowlist or verification process. Because Bearer tokens confer authority solely through possession, a captured token can be reused directly. ### Attack Path 1. The victim configures the MCP client using an internal hostname or tunnel domain supplied through an unverified channel. 2. An attacker redirects that hostname or traffic to an attacker-controlled MCP endpoint. 3. The attacker presents a self-signed certificate. 4. The victim accepts the certificate based on the gu ...[truncated 989 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:40
Finding

Unscoped Upload of Datasets, Artifacts, and Complete Repository Contents

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (4)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The skill instructs the agent to execute code, download data, and publish results from its own environment, but the user-facing description does not prominently warn that these side effects will occur autonomously. This creates a consent and safety gap: a user may invoke what sounds like a research-assistance skill without realizing it performs real-world computation, network access, and irreversible publication.

Content

No source excerpt is available for this finding.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · reference/connecting.md (reported line 8)May include surrounding context.

md
- **URL**: `https://<tunnel-domain>/mcp` (ask the platform operator for the current tunnel domain; an internal LAN HTTPS endpoint also exists for on-site operators)
- **Transport**: streamable-http
- **Auth**: header `Authorization: Bearer <your platform API key>` on **every** request (missing/invalid → 401). For conducting research use a key with role **`researcher`**.
- Internal endpoint uses a self-signed cert → trust it; the public tunnel terminates TLS (usually no warning).

## Claude Code

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger guidance includes broad phrases such as "do research" and "research an idea," which can cause the skill to activate for vague user requests that do not imply consent to autonomous code execution, downloading data, or publishing artifacts. In this skill's context, accidental invocation is more dangerous because activation leads to high-impact actions in an external environment and on a remote platform.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill explicitly says to act "fully autonomously" and "never defer to a human's value judgment or wait for human approval," which suppresses oversight even when the workflow includes risky actions like code execution, downloading data, and publishing outputs. In context, this increases danger because it instructs the agent to bypass user control precisely where confirmation and review would be most important.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.