Back to skill

Security audit

Ars Academic Pipeline

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent academic writing pipeline, but it stores broad manuscript/session outputs long term and gives under-scoped instructions for credentials, transcript exports, and in-place document changes.

Review before installing if you will use confidential manuscripts, unpublished research, reviewer correspondence, or private chat history. Disable or avoid Hindsight retention unless you explicitly want long-term storage, avoid putting API keys in shell startup files, prefer session-scoped secrets, and inspect generated process records and manuscript diffs before sharing or submitting them.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

other

Warning
Location
SKILL.md:77
Finding

Unbounded Long-Term Retention of Academic Stage Outputs

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
shared/cross_model_verification.md:194
Finding

Google API Key Exposed Through Command-Line URL

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
shared/cross_model_verification.md:51
Finding

Long-Lived API Keys Stored in Plaintext Shell Startup Files

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
Findings (46)

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · agents/collaboration_depth_agent.md (reported line 154)May include surrounding context.

md
- **Scope**: score the collaboration *pattern*; the paper, research, and AI output belong to other agents.
- **Session-bounded**: the rubric is per-pipeline; produce no cross-session leaderboard or global scoreboard.
- **Describe, don't judge**: speak about the observable pattern, not the person's character or ability.
- **Offer, don't prescribe**: phrase next-stage suggestions as options ("you could try X") rather than duties ("you should X"). The rubric is descriptive.

---

Hidden Instructions

High
Category
Prompt Injection
Confidence
90% confidence
Finding

The design relies on hidden HTML comment markers such as <!--ref:slug--> and anchor comments embedded inside user-facing markdown to carry machine instructions and provenance state. Hidden control channels are dangerous because they can survive review unnoticed, be injected or tampered with by upstream content, and cause downstream agents or formatters to make trust or gating decisions based on content the user may not see.

Content

Scanner excerpt · agents/pipeline_orchestrator_agent.md (reported line 626)May include surrounding context.

md
**Inputs (read-only):**

- The current draft markdown containing `<!--ref:slug-->` HTML-comment markers (one per emitted citation, per Step 3a's two-layer form).
- The Material Passport `literature_corpus[]` entries (each carries `citation_key`, `source_acquired`, `source_verified_against_original`).
- The peer-file `<session>_human_read_log.yaml` (path computed as `<passport-path-parent>/<passport-stem>_human_read_log.yaml` per §3.6 round-5 R5-003 amend) — provides `human_read_source: true` for every `citation_key` the user has explicitly marked via `/ars-mark-read`.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · agents/state_tracker_agent.md (reported line 514)May include surrounding context.

md
---

## Dashboard Output Rules

1. Produce full version when user explicitly requests it
2. **Append simplified version to checkpoint notification after each stage completion**

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · agents/pipeline_orchestrator_agent.md (reported line 221)May include surrounding context.

md
## Interaction with existing features

- **Collaboration Depth Observer (v3.5.0):** fires on FULL/SLIM as before. Observer output is included in the checkpoint notification regardless of reset state. Observer state does NOT carry across resets; each fresh session observes only its own stage.
- **Compliance agent (v3.4.0):** `compliance_history[]` remains append-only and is consumed from the passport on resume. No change to Schema 12.
- **Sprint contract (v3.6.2):** reviewer sprint contracts load from the passport on resume (Phase 1 paper-content-blind stage remains valid across the reset boundary because the contract + paper metadata are carried in the passport).
- **Socratic reading probe (v3.5.1):** reading probe fires at most once per session. Across a reset boundary, the probe counter resets — the next session may fire its own probe. This is by design: each session is its own Socratic unit.

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · references/passport_as_reset_boundary.md (reported line 112)May include surrounding context.

md
## Interaction with existing features

- **Collaboration Depth Observer (v3.5.0):** fires on FULL/SLIM as before. Observer output is included in the checkpoint notification regardless of reset state. Observer state does NOT carry across resets; each fresh session observes only its own stage.
- **Compliance agent (v3.4.0):** `compliance_history[]` remains append-only and is consumed from the passport on resume. No change to Schema 12.
- **Sprint contract (v3.6.2):** reviewer sprint contracts load from the passport on resume (Phase 1 paper-content-blind stage remains valid across the reset boundary because the contract + paper metadata are carried in the passport).
- **Socratic reading probe (v3.5.1):** reading probe fires at most once per session. Across a reset boundary, the probe counter resets — the next session may fire its own probe. This is by design: each session is its own Socratic unit.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · shared/benchmark_report_pattern.md (reported line 101)May include surrounding context.

md
### `caveats`

Non-empty array, items non-empty strings, minItems 1. The schema physically prevents an
empty caveats field. A benchmark report with no caveats either has no known limitations
(implausible for any real-world evaluation) or the author didn't think about limitations
(which disqualifies the report more than any specific limitation would).

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · shared/benchmark_report_pattern.md (reported line 154)May include surrounding context.

md
without a stated `human_baseline.sample_size`. n=1 is worse than n=10, but it is
  infinitely better than unstated.

- **Author-conducted with no caveat disclosure.** Using `author-conducted` is allowed.
  Not noting it in `caveats` as a limitation is an editorial failure, not a schema
  failure — but reviewers will notice.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · shared/benchmark_report_pattern.md (reported line 158)May include surrounding context.

md
Not noting it in `caveats` as a limitation is an editorial failure, not a schema
  failure — but reviewers will notice.

- **Self-scoring with no warning acknowledgment.** `self-scored` produces a warning on
  every validator run. If you publish a report with `self-scored` and don't address it
  in caveats, the omission is visible in the JSON to anyone who checks.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest includes many generic trigger phrases such as "full pipeline," "publish paper," and "research to paper," which are broad enough to match ordinary academic-assistance requests outside a tightly scoped invocation context. Overbroad activation can cause this orchestration skill to run unexpectedly, chaining multiple delegated stages and data flows when a user may have intended a narrower task.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 43)May include surrounding context.

md
## Hermes Execution

The orchestrator (main Hermes agent):
1. **Stage 1**: Clarify user intent → select modes → create Project Plan
2. **Stage 2-6**: Sequential delegate_task calls, each loading the relevant skill
3. **Stage 7-9**: Lightweight parallel or sequential based on dependencies
4. **Stage 10**: Write final package to disk

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description limits the skill to verifying references, citations, and data for factual accuracy. However, the documented behavior includes Phase D originality verification, including paragraph-level plagiarism screening and self-plagiarism checks, which is a distinct function beyond factual verification of citations and data.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The skill hardcodes verification sources and workflows that are biased toward English-language and globally indexed systems without user opt-in or locale-sensitive fallback. In an academic pipeline that may process multilingual or regional scholarship, this can systematically misclassify legitimate non-English or locally indexed references as NOT_FOUND or lower-confidence, causing false integrity failures and disproportionate exclusion of certain sources.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest describes a 10-stage end-to-end pipeline: planning → deep-research → academic-paper → academic-paper-reviewer → revision → polish → ethics → disclosure → format → deliver. This file instead defines coordination around research, writing, integrity, review, revision, finalize, and completion, with no concrete planning, polish, ethics, disclosure, or deliver stages matching the manifest sequence.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The routing triggers use broad natural-language keywords like 'research', 'review', 'draft', and 'format' without strong exclusion logic. In an agent environment, this can cause accidental invocation or misrouting of user requests into a complex pipeline, potentially leading to unintended processing of user files and autonomous stage transitions.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file states the orchestrator should only coordinate, but later instructs it to mutate draft content in place and participate in document-production behavior. This creates a dangerous authority mismatch: operators and downstream systems may trust the orchestrator as non-substantive while it is actually empowered to alter user materials, which can lead to unexpected content modification and reduced oversight.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The orchestrator instructs in-place draft mutation during cite-time provenance finalization but does not require explicit user-facing notice or consent before modifying user-authored content. Silent or unexpected edits to manuscripts can corrupt user work, introduce compliance text the user did not approve, or change publication-ready artifacts without clear accountability.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown skill defines the agent's role and behavior but does not state when it should be invoked, what exact trigger phrases apply, or any exclusion conditions. For markdown files, missing specificity on trigger scope can cause unintended invocation by an orchestrator or user because the boundary between applicable and non-applicable requests is unclear.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list for entering external review mode includes broad natural-language examples plus an open-ended "etc.", which makes the activation boundary ambiguous. In an orchestrated multi-stage agent, ambiguous triggers can cause unintended mode switching, leading the system to parse arbitrary user text as reviewer feedback and perform the wrong workflow, which is a real prompt-scope/control issue even if not directly exploitable for code execution.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · agents/pipeline_orchestrator_agent.md (reported line 92)May include surrounding context.

md
Orchestrator obligations on resume:
- Locate the target `kind: boundary` entry by matching `hash`. Hard error if no match, or if a later `kind: resume` entry already carries `consumes_hash == <hash>` (double-resume is forbidden).
- Do NOT ask the user to re-summarize prior stages; the passport is authoritative. Load artifacts by reference (paths or IDs recorded in the entry).
- Honor the `verification_status` field. If `STALE` or `UNVERIFIED`, display a warning and prompt the user to re-verify before continuing. If `VERIFIED`, proceed without prompting.
- If `pending_decision` is set on the ledger entry, re-prompt the user for that decision BEFORE invoking any downstream stage. Display `pending_decision.question` and each option's `value`. After the user picks, look up the matching entry in `options[]` by `value`, then use that entry's `next_stage` and `next_mode` to determine actual routing. Record the chosen `value` as `chosen_branch` on the new `resume` entry. `next` on the boundary entry is advisory and is superseded by the matched option's `next_stage`. A user-supplied `stage=<n>` override on the resume command does NOT satisfy `pending_decision` — the decision prompt always fires when `pending_decision` is present. CLI `stage=`/`mode=` overrides still win over option routing if the user supplies them after the decision prompt.
- Emit a `### Resume Acknowledged` section at the start of the new session with: hash, source session `session_marker` + `generated_at`, recovered stage, and next-stage plan.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/passport_as_reset_boundary.md (reported line 61)May include surrounding context.

md
Orchestrator obligations on resume:
- Locate the target `kind: boundary` entry by matching `hash`. Hard error if no match, or if a later `kind: resume` entry already carries `consumes_hash == <hash>` (double-resume is forbidden).
- Do NOT ask the user to re-summarize prior stages; the passport is authoritative. Load artifacts by reference (paths or IDs recorded in the entry).
- Honor the `verification_status` field. If `STALE` or `UNVERIFIED`, display a warning and prompt the user to re-verify before continuing. If `VERIFIED`, proceed without prompting.
- If `pending_decision` is set on the ledger entry, re-prompt the user for that decision BEFORE invoking any downstream stage. Display `pending_decision.question` and each option's `value`. After the user picks, look up the matching entry in `options[]` by `value`, then use that entry's `next_stage` and `next_mode` to determine actual routing. Record the chosen `value` as `chosen_branch` on the new `resume` entry. `next` on the boundary entry is advisory and is superseded by the matched option's `next_stage`. A user-supplied `stage=<n>` override on the resume command does NOT satisfy `pending_decision` — the decision prompt always fires when `pending_decision` is present. CLI `stage=`/`mode=` overrides still win over option routing if the user supplies them after the decision prompt.
- Emit a `### Resume Acknowledged` section at the start of the new session with: hash, source session `session_marker` + `generated_at`, recovered stage, and next-stage plan.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

Forcing English-language confirmation without user choice can cause users to misunderstand gating decisions, branch selection, or integrity-check acknowledgments. In a pipeline with mandatory checkpoints and resume semantics, language mismatch increases the chance of erroneous consent or incorrect transitions, which is a workflow safety issue even if it is not a direct code-execution flaw.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The pause transition is triggered by broad natural-language phrases like "stop here," which can easily appear in normal academic discussion or quoted content. In an orchestrator that controls multi-stage execution and resume behavior, this creates a prompt/command ambiguity issue that can unintentionally alter pipeline state, disrupt execution, or cause incorrect checkpoint handling.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a 10-stage pipeline ending with deliver, but this file documents an added 'Stage 6' whose purpose is generating process-history records and collaboration evaluations. That is a substantive behavior expansion beyond the manifest description, not merely an implementation detail, because it adds a distinct output artifact and analysis function over user/AI behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The protocol directs the agent to document the complete human-AI collaboration history for sharing/reporting, but does not require consent review, sensitive-data screening, or minimization. This can turn transient chat content into a durable disclosure artifact containing personal, confidential, or regulated information.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The protocol instructs the system to compile comprehensive user-provided content, including full session history and verbatim quotes, for later sharing/export. This creates a built-in data exfiltration channel from private conversation context into portable documents, even when those details are not necessary for the publication workflow.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.