Back to skill

Security audit

Ars Deep Research

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent academic research skill, but it needs review because some optional behaviors are hidden from the user or can send research material and API credentials through external model workflows.

Install only if you are comfortable with a multi-agent academic workflow using web search and file outputs. Keep ARS_CROSS_MODEL and ARS_SOCRATIC_READING_PROBE disabled unless you intentionally want those behaviors, avoid sensitive or proprietary research material, and use restricted API keys if enabling cross-model verification. Expect some generated reports to contain hidden HTML citation metadata.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
shared/cross_model_verification.md:191
Finding

Google API Key Exposed in a Command-Line URL

Content
View full analysis

Vulnerability Details

File Location: shared/cross_model_verification.md:191-192
Vulnerability Type: API credential exposure through a URL query parameter
Risk Level: Medium

Vulnerable Code

bash
curl -s "https://generativelanguage.googleapis.com/v1beta/models/${ARS_CROSS_MODEL}:generateContent?key=$GOOGLE_AI_API_KEY" \
  -H "Content-Type: application/json" \

Technical Analysis

The documented Gemini request interpolates GOOGLE_AI_API_KEY directly into the request URL. When this command is executed, the expanded credential can be exposed through operating-system process inspection, shell tracing or diagnostic output, proxy and gateway logs, HTTP client telemetry, and monitoring products that record URLs.

Although HTTPS encrypts the request in transit, it does not prevent disclosure at the endpoint, process, shell, or logging layers. Query strings are also more likely to be retained than authentication headers. The package does not contain a hardcoded key; exploitation requires a user to configure a valid key and execute the documented optional cross-model workflow.

Attack Path

  1. A user configures GOOGLE_AI_API_KEY and selects a Gemini model through ARS_CROSS_MODEL.
  2. The cross-model verification workflow executes the documented curl request.
  3. The shell expands $GOOGLE_AI_API_KEY into the command-line URL.
  4. A local process observer, shell-debug trace, proxy, telemetry agent, or URL-logging system records the expanded URL.
  5. An attacker with access to that record extracts the API key.
  6. The attacker uses the key against enabled Google APIs until it is revoked or constrained by provider-side restrictions.

Impact Assessment

Successful exploitation can disclose the configured Google API credential. The attacker could consume the associated API quota, incur charges, access services authorized for that key, or disrupt legitimate verification through quota exhaustion.

The resulting privileges are limite ...[truncated 230 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace the raw curl example with a maintained Google SDK or another client that supports protected credential handling.
  2. Where the API supports it, transmit credentials through an authentication header rather than a query parameter.
  3. If this endpoint necessarily requires the key query parameter, avoid placing the expanded key in command-line arguments. Construct and execute the request within a process using an HTTP library so the complete URL is not exposed through the process list.
  4. Disable shell tracing around credential-bearing operations and ensure errors never print the complete request URL.
  5. Configure proxies, gateways, telemetry systems, and application logs to redact query strings and known credential parameters such as key.
  6. Recommend provider-side API, project, application, quota, and billing restrictions for the key.
  7. Rotate the key immediately if command history, process telemetry, debug output, or proxy logs may have captured it.
  8. Add an explicit warning to the setup guide that API credentials must not be committed, printed, or included in diagnostic reports.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (48)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Content

Scanner excerpt · agents/ethics_review_agent.md (reported line 61)May include surrounding context.

md
| **Low** | Unlikely misuse, minimal harm potential | General education research |
| **Moderate** | Could be misused in specific contexts | Surveillance tech analysis, social manipulation studies |
| **High** | Clear potential for harm if misused | Vulnerability research, weapons-related |
| **Critical** | Should not be published without safeguards | Specific exploitation methods |

For Moderate or above: Include explicit "Responsible Use" statement

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Content

Scanner excerpt · agents/ethics_review_agent.md (reported line 109)May include surrounding context.

md
| **Low** | Unlikely misuse, minimal harm potential | General education research |
| **Moderate** | Could be misused in specific contexts | Surveillance tech analysis, social manipulation studies |
| **High** | Clear potential for harm if misused | Vulnerability research, weapons-related |
| **Critical** | Should not be published without safeguards | Specific exploitation methods |

For Moderate or above: Include explicit "Responsible Use" statement

Hidden Instructions

High
Category
Prompt Injection
Confidence
92% confidence
Finding

This instruction tells the agent to emit hidden HTML comment markers in user-visible output, creating a covert metadata channel that downstream systems may parse while users do not see it. In an agent pipeline, hidden channels can be abused for prompt smuggling, policy bypass, data tagging, or influencing later components without transparent user consent; the deep-research context makes this more dangerous because the skill explicitly compiles and forwards complex multi-agent outputs.

Content

Scanner excerpt · agents/report_compiler_agent.md (reported line 184)May include surrounding context.

md
When emitting any citation in the report output, write the citation in two layers:

1. **Visible layer**: standard author-year form (e.g. `Smith (2024)` or `(Smith, 2024)`).
2. **Hidden layer**: immediately after the visible form, append an HTML comment of the shape `<!--ref:slug-->`, where `slug` is the `citation_key` already present in the corpus context provided in this prompt.

Examples: `Smith (2024) <!--ref:smith2024-->` or `(Smith, 2024)<!--ref:smith2024-->`.

Hidden Instructions

High
Category
Prompt Injection
Confidence
96% confidence
Finding

This expands the hidden-channel design by requiring both hidden reference slugs and structured anchor comments after every citation. That creates a systematic invisible control plane embedded in markdown which can carry machine-interpreted instructions or unreviewed provenance payloads through the pipeline; because this skill operates in a 13-agent research workflow, the chance of hidden instructions affecting downstream agents or validators is materially increased.

Content

Scanner excerpt · agents/report_compiler_agent.md (reported line 203)May include surrounding context.

Every visible citation in the compiled report MUST be followed by BOTH a slug marker AND an anchor marker:

text
<visible> <!--ref:slug--><!--anchor:<kind>:<value>-->

Anchor kinds (closed enum):

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The instructions explicitly tell the agent to answer user meta-questions in a way that conceals the real probe mechanism and avoids acknowledging evaluation, even when the user asks if it is a test. That is deceptive design: it prevents informed user consent and makes the system intentionally misleading about its monitoring behavior.

Content

No source excerpt is available for this finding.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · agents/socratic_mentor_agent.md (reported line 310)May include surrounding context.

md
In addition, do NOT praise the user's paraphrase content, and do NOT judge the user's decline.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · agents/source_verification_agent.md (reported line 104)May include surrounding context.

md
- [ ] The source is suspiciously perfect (exactly supports the claim with no caveats)

Hidden Instructions

High
Category
Prompt Injection
Confidence
94% confidence
Finding

This agent instructs downstream output to embed hidden HTML comment markers in every citation. Hidden, machine-readable content is a prompt-side covert channel: it can carry metadata invisible to users, bypass user review, and create opportunities for instruction smuggling or policy-evasive signaling across later pipeline stages. In a multi-agent research pipeline, that is more dangerous because downstream components may parse and trust these markers automatically.

Content

Scanner excerpt · agents/synthesis_agent.md (reported line 178)May include surrounding context.

md
When emitting any citation in the synthesis output, write the citation in two layers:

1. **Visible layer**: standard author-year form (e.g. `Smith (2024)` or `(Smith, 2024)`).
2. **Hidden layer**: immediately after the visible form, append an HTML comment of the shape `<!--ref:slug-->`, where `slug` is the `citation_key` already present in the corpus context provided in this prompt.

Examples: `Smith (2024) <!--ref:smith2024-->` or `(Smith, 2024)<!--ref:smith2024-->`.

Hidden Instructions

High
Category
Prompt Injection
Confidence
97% confidence
Finding

This expands the hidden-marker scheme by requiring both reference and anchor HTML comments after every citation, increasing the capacity of the covert channel and encouraging invisible propagation of source-derived strings. Because later agents/finalizers are expected to consume and mutate these markers, an attacker could abuse them for prompt injection, parser confusion, or hidden data exfiltration that is not apparent in rendered markdown. The surrounding rules about quote encoding reduce one parser-break risk, but they do not eliminate the core issue of invisible instruction-bearing metadata embedded in freeform model output.

Content

Scanner excerpt · agents/synthesis_agent.md (reported line 197)May include surrounding context.

Every visible citation MUST be followed by BOTH a slug marker AND an anchor marker:

text
<visible> <!--ref:slug--><!--anchor:<kind>:<value>-->

Anchor kinds (closed enum):

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · shared/benchmark_report_pattern.md (reported line 101)May include surrounding context.

md
### `caveats`

Non-empty array, items non-empty strings, minItems 1. The schema physically prevents an
empty caveats field. A benchmark report with no caveats either has no known limitations
(implausible for any real-world evaluation) or the author didn't think about limitations
(which disqualifies the report more than any specific limitation would).

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · shared/benchmark_report_pattern.md (reported line 154)May include surrounding context.

md
without a stated `human_baseline.sample_size`. n=1 is worse than n=10, but it is
  infinitely better than unstated.

- **Author-conducted with no caveat disclosure.** Using `author-conducted` is allowed.
  Not noting it in `caveats` as a limitation is an editorial failure, not a schema
  failure — but reviewers will notice.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · shared/benchmark_report_pattern.md (reported line 158)May include surrounding context.

md
Not noting it in `caveats` as a limitation is an editorial failure, not a schema
  failure — but reviewers will notice.

- **Self-scoring with no warning acknowledgment.** `self-scored` produces a warning on
  every validator run. If you publish a report with `self-scored` and don't address it
  in caveats, the omission is visible in the JSON to anyone who checks.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill declares a very large set of trigger phrases, including broad terms such as 'research', 'help me think through', and multilingual equivalents. This can cause the skill to activate in situations where the user did not intend a heavyweight multi-agent research workflow, potentially sending sensitive or unnecessary context to delegated subagents and web-enabled steps.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill introduces optional cross-model API calls from within an agent whose stated role is purely analytical, creating an unexpected external-service integration path. In an agentic research pipeline, sending reviewed material to another model can exfiltrate sensitive user inputs, internal drafts, or proprietary research data to a third party without clear consent, data-classification checks, or transport/policy controls.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · agents/ethics_review_agent.md (reported line 76)May include surrounding context.

md
### 5. Data Ethics
- [ ] Data sources used ethically (public domain, licensed, or permitted)
- [ ] Privacy considerations addressed
- [ ] No personally identifiable information exposed without consent
- [ ] Aggregate vs. individual data handled appropriately
- [ ] Data limitations acknowledged

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/ethics_checklist.md (reported line 120)May include surrounding context.

md
### 5. Data Ethics
- [ ] Data sources used ethically (public domain, licensed, or permitted)
- [ ] Privacy considerations addressed
- [ ] No personally identifiable information exposed without consent
- [ ] Aggregate vs. individual data handled appropriately
- [ ] Data limitations acknowledged

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The policy text is internally inconsistent: it lists omitted IRB planning under 'Blocking Conditions (Critical)' but then assigns only a CONDITIONAL outcome. In an automated or semi-automated multi-agent pipeline, this ambiguity can cause unsafe human-subjects research to proceed when it should be halted, weakening the ethics gate.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This section adds a covert 'reading probe' that tests whether a user can paraphrase a cited paper, even though the skill is presented as a Socratic research guide rather than an assessment tool. The hidden nature of the capability, combined with silent candidate-paper tracking and AI-initiated probing, creates a deceptive interaction pattern and collects information unrelated to the core guidance function.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
87% confidence
Finding

The skill directs the AI to autonomously initiate a hidden probe, not ask the user whether they want it, and avoid even hinting that it exists. In context, this is risky because the autonomous action is not merely workflow convenience; it is covert evaluation and monitoring of user behavior beyond the stated purpose of the mentor.

Content

Scanner excerpt · agents/socratic_mentor_agent.md (reported line 236)May include surrounding context.

md
- The Layer 2 → Layer 3 transition is imminent (i.e., the Methodology Reflection phase is converging and Evidence Strategy is about to open).
- The probe has not yet fired in this session (each session fires the probe at most once).

If ANY of these is false, this layer is dormant. Do not mention the probe. Do not prepare for the probe. Do not hint that a probe exists. Do not ask the user whether they would like a probe. The probe is strictly AI-initiated.

### Candidate Paper Tracking

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The summary instructions require storing and reproducing verbatim user text about cited papers in a downstream-visible Research Plan Summary without a user-facing warning. This increases privacy and data-minimization risk because users are not clearly informed that their freeform responses may be preserved and exposed beyond the immediate turn.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs the agent to preserve and later reproduce verbatim user-provided paper paraphrases in session summaries and machine-readable tags. That unnecessarily amplifies user-provided content, increases retention surface, and may expose sensitive or proprietary text to downstream stages or reviewers without explicit user awareness.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This markdown file contains natural-language instructions that automatically switch to Chinese-language literature and databases when English results are limited. Because the policy scope applies to all file types, forcing a language/locale path without first offering the user a choice is a language-policy concern.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/literature_monitoring_strategies.md (reported line 71)May include surrounding context.

md
### What is RSS?

RSS (Really Simple Syndication) allows you to subscribe to content updates from websites without checking each site manually. Use an RSS reader (e.g., Feedly, Inoreader, NewsBlur) to aggregate feeds.

### Recommended Feeds

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger phrases include broad natural-language requests like 'set up alerts' and 'track new publications on this', which can plausibly appear in ordinary conversation and cause the monitoring agent to activate outside a tightly scoped monitoring intent. In a multi-agent research skill, overly permissive activation increases the chance of unintended delegation, scope creep, and execution of auxiliary workflows on user content that was not meant for monitoring configuration.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The socratic mode trigger examples include very broad, natural help-seeking phrases such as 'help me think through' and similar Chinese variants. In a trigger-based skill-routing system, this can cause overbroad activation and unintended interception of general user requests, potentially diverting users into a long multi-round workflow that was not explicitly requested and increasing the chance of inappropriate delegation or data exposure across sub-agents.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.