Back to skill

Security audit

Wei Devils Advocate

Security checks for vulnerabilities and agentic risk

Overview

This skill is not clearly malicious, but it needs Review because it uses risky install instructions and automatically stores sensitive prompts and raw model outputs in plaintext.

Before installing, review the setup path and avoid running the documented curl-to-bash command; prefer a package-manager or verified Bun install. Do not submit secrets, personal data, regulated data, or confidential business material unless you accept transmission to the configured LLM providers and automatic local plaintext retention. Check or delete intermediate/ and reports/ after use, and wait for the publisher to align package metadata and update dependency hygiene if provenance matters to your environment.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T03 · Remote Payload Retrieval and Execution

Error
Location
SKILL.md:49
Finding

Unverified Remote Installer Is Piped Directly into a Shell

Content
View full analysis
Remediation
View remediation
' echo ' bun-installer.sh' | sha256sum -c - less bun-installer.sh bash bun-installer.sh ``` The checksum must be obtained from a trusted, version-specific source rather than from the same unauthenticated delivery path. ]]>

T01 · Skill Instruction Hijacking

Error
Location
scripts/agent.ts:451
Finding

Second-Order Prompt Injection Can Manipulate the Judge Evaluation

Content
View full analysis
` Model: ${r.model} Key Assumptions: ${r.keyAssumptions.join(', ')} Counterarguments: ${r.counterarguments.join(', ')} Failure Scenarios: ${r.failureScenarios.join(', ')} `).join('\n---\n'); const judgePrompt = `Thesis to evaluate: ${query} Counterarguments from multiple models: ${counterargumentsText} ${JUDGE_PROMPT_TEMPLATE}`; const messages: ChatMessage[] = [ { role: 'user', content: judgePrompt }, ]; const response = await this.callModel( this.judgeModel, messages, { maxTokens: appConfig.max_tokens_judge } ); ``` ### Technical Analysis The application interpolates the user thesis and content generated by debater models into the same user-role message that contains the judge's operational instructions. Debater responses are untrusted because they can be influenced by: - A deliberately crafted user thesis. - Prompt injection contained in retrieved online content. - Unexpected or adversarial behavior from an external model provider. The judge receives no trusted system message defining the supplied thesis and model responses as inert data. Consequently, instructions emitted inside fields such as `Counterarguments` or `Failure Scenarios` may be interpreted as instructions for the judge rather than evidence to evaluate. The input blacklist in `sanitizeInput()` does not adequately mitigate this weakness. It only replaces a few literal phrases in the original query: ```ts const injectionPatterns = [ /ignore previous instructions/gi, /system prompt/gi, /you are now/gi, /disregard/gi, /forget/gi, //g, ]; ``` These expressions are defensive rather than malicious, but phrase blacklists are readily bypassed through paraphrasing, Unicode substitutions, inserted whitespa ...[truncated 1742 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/agent.ts:518
Finding

Sensitive Queries and Raw LLM Responses Are Persisted in Plaintext

Content
View full analysis
{ if (response) { const filename = `${response.model}-${timestamp}.txt`; const filepath = join(dir, filename); const content = [ `Model: ${response.model}`, `Timestamp: ${timestamp}`, `Query: ${query}`, ``, `=== Key Assumptions ===`, ...response.keyAssumptions.map(p => `- ${p}`), ``, `=== Counterarguments ===`, ...response.counterarguments.map(p => `- ${p}`), ``, `=== Failure Scenarios ===`, ...response.failureScenarios.map(p => `- ${p}`), ``, `=== What Would Prove You Wrong ===`, ...response.whatWouldProveWrong.map(p => `- ${p}`), ``, `=== Raw Response ===`, response.rawResponse, ].join('\n'); writeFileSync(filepath, content, 'utf-8'); console.log(`[DevilsAdvocate] Saved: intermediate/${filename}`); } }); } ``` The same file also persists the raw judge output and consolidated report: ```ts const filename = `${name}-${timestamp}.txt`; const filepath = join(dir, filename); writeFileSync(filepath, content, 'utf-8'); ``` ```ts writeFileSync(filepath, content, 'utf-8'); console.log(`[DevilsAdvocate] Saved: reports/report-${timestamp}.txt`); return filepath; ``` ### Technical Analysis Every successful run writes the sanitized query, parsed analysis, raw debater output, raw judge output, and final report to plaintext files under `intermediate/` and ...[truncated 2135 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (64)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · .gitignore (reported line 2)May include surrounding context.

text
# Environment variables
.env
.env.local
.env.*.local

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · .gitignore (reported line 3)May include surrounding context.

text
# Environment variables
.env
.env.local
.env.*.local

# Node.js

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill is presented as a specialized devil's-advocate workflow, but the finding indicates it can act as a generic OpenRouter chat client and write debug output locally. That broader behavior can expose sensitive prompts, outputs, or metadata to third parties or disk without clear user expectation or consent.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is presented as a specialized devil's-advocate workflow, but the finding indicates it can act as a generic OpenRouter chat client and write debug output locally. That broader behavior can expose sensitive prompts, outputs, or metadata to third parties or disk without clear user expectation or consent.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill is presented as a specialized devil's-advocate workflow, but the finding indicates it can act as a generic OpenRouter chat client and write debug output locally. That broader behavior can expose sensitive prompts, outputs, or metadata to third parties or disk without clear user expectation or consent.

Content

No source excerpt is available for this finding.

External Script Fetching

High
Category
Supply Chain
Confidence
98% confidence
Finding

The installation instructions tell users to fetch a remote script and execute it immediately via the shell. This bypasses normal review and integrity checks, so if the upstream server, CDN, or connection is compromised, arbitrary code runs on the user's machine.

Content

Scanner excerpt · SKILL.md (reported line 49)May include surrounding context.

Install Bun

bash
curl -fsSL https://bun.sh/install | bash

Or on macOS with Homebrew:

Chaining Abuse

High
Category
Tool Misuse
Confidence
98% confidence
Finding

The use of '| bash' is a classic command-chaining pattern that executes untrusted network content directly, turning a documentation snippet into an arbitrary-code-execution vector. In skill context, this is more dangerous because users may follow setup steps verbatim to enable the skill, increasing the chance of compromise during installation rather than during normal use.

Content

Scanner excerpt · SKILL.md (reported line 49)May include surrounding context.

Install Bun

bash
curl -fsSL https://bun.sh/install | bash

Or on macOS with Homebrew:

Known Vulnerable Dependency: axios==1.13.6 — 16 advisory(ies): CVE-2026-44494 (axios Vulnerable to Full Man-in-the-Middle via Prototype Pollution Gadget in `co); CVE-2026-44495 (axios Vulnerable to Credential Theft and Response Hijacking via Prototype Pollut); CVE-2025-62718 (Axios has a NO_PROXY Hostname Normalization Bypass that Leads to SSRF) +13 more

High
Category
Supply Chain
Confidence
91% confidence
Finding

axios is a runtime dependency and the lockfile pins version 1.13.6, which the finding reports as affected by multiple advisories including SSRF-related proxy bypass and prototype-pollution-adjacent issues. In a skill that orchestrates multiple LLMs and likely performs outbound network calls, a vulnerable HTTP client materially increases risk of server-side request forgery, credential leakage, redirect abuse, or unsafe request handling depending on how requests are constructed.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: form-data==4.0.5 — 1 advisory(ies): CVE-2026-12143 (form-data: CRLF injection in form-data via unescaped multipart field names and f)

High
Category
Supply Chain
Confidence
86% confidence
Finding

form-data 4.0.5 is a runtime dependency and the cited CRLF injection issue in multipart construction can enable request smuggling or header/body manipulation when attacker-controlled field names or filenames are included. If this skill uploads files or relays untrusted content to external services, the vulnerability could be used to tamper with outbound requests or bypass downstream security assumptions.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: axios==1.13.6 — 16 advisory(ies): CVE-2026-44494 (axios Vulnerable to Full Man-in-the-Middle via Prototype Pollution Gadget in `co); CVE-2026-44495 (axios Vulnerable to Credential Theft and Response Hijacking via Prototype Pollut); CVE-2025-62718 (Axios has a NO_PROXY Hostname Normalization Bypass that Leads to SSRF) +13 more

High
Category
Supply Chain
Confidence
98% confidence
Finding

The dependency set includes axios at a version identified by the scanner as having multiple advisories, including SSRF-related and prototype-pollution-assisted man-in-the-middle/credential-theft scenarios. Because this skill is described as querying multiple LLMs in parallel and likely performs outbound HTTP requests, a vulnerable HTTP client is more dangerous in context: it may process attacker-influenced URLs, proxy settings, or responses, increasing the chance of credential leakage, SSRF, or response tampering.

Content

No source excerpt is available for this finding.

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · scripts/agent.ts (reported line 5)May include surrounding context.

ts
/**
 * Devil's Advocate Agent
 *
 * Core implementation of the devil's advocate skill.
 * Challenges ideas by generating strong counterarguments from multiple LLMs.
 */

import { readFileSync, mkdirSync, writeFileSync } from 'fs';
import { dirname, join } from 'path';
import { fileURLToPath } from 'url';
import { BailianClient, OpenRouterClient, OpenAICompliantClient } from './clients/index.js';
import type { ChatMessage, ChatCompletionResponse } from './clients/bailian.js';

/** Configuration file structure */
interface ConfigFile {
  judge_model: string;
  max_models: number;
  max_tokens: number;
  max_tokens_judge: number;
  depth: string;
  models: Record<string, {
    provider: string;
    model_id: string;
    api_base: str

Instruction Override

High
Category
Prompt Injection
Confidence
80% confidence
Finding

This pattern attempts to override system instructions or ignore safety constraints. Without LLM analysis, manual review is recommended.

Content

Scanner excerpt · scripts/agent.ts (reported line 286)May include surrounding context.

ts
sanitized = sanitized.replace(/[\x00-\x08\x0B-\x0C\x0E-\x1F\x7F]/g, '');

    const injectionPatterns = [
      /ignore previous instructions/gi,
      /system prompt/gi,
      /you are now/gi,
      /disregard/gi,

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/clients/bailian.ts (reported line 7)May include surrounding context.

ts
* A reusable HTTP client for making requests to the OpenRouter API.
 * Supports rate limiting, retries, and configurable timeouts.
 *
 * Environment variables are automatically loaded from .env file by Bun.
 * See: https://bun.sh/docs/runtime/env
 */

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/clients/openai_compliant.ts (reported line 10)May include surrounding context.

ts
* A reusable HTTP client for making requests to the OpenRouter API.
 * Supports rate limiting, retries, and configurable timeouts.
 *
 * Environment variables are automatically loaded from .env file by Bun.
 * See: https://bun.sh/docs/runtime/env
 */

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/clients/openrouter.ts (reported line 7)May include surrounding context.

ts
* A reusable HTTP client for making requests to the OpenRouter API.
 * Supports rate limiting, retries, and configurable timeouts.
 *
 * Environment variables are automatically loaded from .env file by Bun.
 * See: https://bun.sh/docs/runtime/env
 */

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · scripts/index.ts (reported line 442)May include surrounding context.

ts
Examples:
  bun run scripts/index.ts "What are the economic impacts of AI?"
  bun run scripts/index.ts -m glm-5,gpt-5.4 "Explain quantum computing"
  bun run scripts/index.ts -t financial "Will the Fed cut rates in 2026?"
  bun run scripts/index.ts -t technical "How do I implement a distributed transaction?"
  bun run scripts/index.ts --json "Latest AI breakthroughs"

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README explicitly instructs users to store live API keys in a project-root .env file but does not warn that such files must be excluded from version control, protected in local environments, and never shared. This creates a realistic secret-exposure risk through accidental commits, packaging, screenshots, support bundles, or multi-user development environments.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill declares environment variables and, per the static findings, relies on networked execution, but the manifest does not declare any explicit tool scope such as permissions or allowed-tools. This weakens least-privilege controls and makes the skill's external data access and secret usage less transparent to operators reviewing SKILL.md.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill routes user queries to multiple third-party LLM providers, yet the documentation does not warn users that their prompts may be transmitted externally. This is a privacy and compliance issue because users may provide confidential or regulated information without realizing it leaves the local environment and may reach multiple vendors.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The configuration routes prompts to external OpenRouter-hosted models and relies on an API key from the environment, which implies user data may be transmitted to third-party services. In a multi-model adversarial-analysis skill, prompts may contain sensitive business ideas, internal documents, or proprietary reasoning, so the lack of any explicit user-facing disclosure or consent mechanism creates a real data-exposure risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This manifest/config file sets the region to "cn" directly, which is a locale/region constraint expressed in natural-language-like configuration. There is no indication in this file that the user can opt into this locale or that the constraint is documented as justified for a region-specific deployment.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This JSON config sets "region": "cn", which imposes a specific locale/region behavior at the configuration level. Under the policy rules, forcing a specific language/locale or regional setting without opt-in or clear justification is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest context says this skill is wei-devils-advocate, intended for adversarial multi-LLM idea stress-testing, but the lockfile's top-level package name is wei-research. That indicates the shipped code/dependency set belongs to a differently named skill, which is a concrete mismatch between declared identity/purpose and the actual packaged artifact.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The package metadata names and describes a different skill ('wei-research' / 'Multi-Model Researcher') than the provided manifest context ('wei-devils-advocate'). This mismatch can mislead reviewers, users, and automated tooling, and in a skill ecosystem it raises supply-chain and provenance concerns because the packaged artifact may not correspond to the declared skill.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The entire requirements file is written in Chinese and provides no indication that other languages are supported or that Chinese is an intentional, documented locale constraint. Under the language/locale policy rule, this can be a natural-language policy violation when a specific language is effectively forced without user opt-in or justification.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.env_credential_access, suspicious.exposed_secret_literal

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
scripts/clients/bailian.ts:136

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
scripts/clients/openai_compliant.ts:152

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
scripts/clients/openrouter.ts:120

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/clients/openai_compliant.ts:212

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/index.ts:216