other
- Location
SKILL.md:77- Finding
Unbounded Long-Term Retention of Academic Stage Outputs
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is a coherent academic writing pipeline, but it stores broad manuscript/session outputs long term and gives under-scoped instructions for credentials, transcript exports, and in-place document changes.
Review before installing if you will use confidential manuscripts, unpublished research, reviewer correspondence, or private chat history. Disable or avoid Hindsight retention unless you explicitly want long-term storage, avoid putting API keys in shell startup files, prefer session-scoped secrets, and inspect generated process records and manuscript diffs before sharing or submitting them.
SKILL.md:77Unbounded Long-Term Retention of Academic Stage Outputs
shared/cross_model_verification.md:194Google API Key Exposed Through Command-Line URL
shared/cross_model_verification.md:51Long-Lived API Keys Stored in Plaintext Shell Startup Files
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
- **Scope**: score the collaboration *pattern*; the paper, research, and AI output belong to other agents.
- **Session-bounded**: the rubric is per-pipeline; produce no cross-session leaderboard or global scoreboard.
- **Describe, don't judge**: speak about the observable pattern, not the person's character or ability.
- **Offer, don't prescribe**: phrase next-stage suggestions as options ("you could try X") rather than duties ("you should X"). The rubric is descriptive.
---
The design relies on hidden HTML comment markers such as <!--ref:slug--> and anchor comments embedded inside user-facing markdown to carry machine instructions and provenance state. Hidden control channels are dangerous because they can survive review unnoticed, be injected or tampered with by upstream content, and cause downstream agents or formatters to make trust or gating decisions based on content the user may not see.
**Inputs (read-only):**
- The current draft markdown containing `<!--ref:slug-->` HTML-comment markers (one per emitted citation, per Step 3a's two-layer form).
- The Material Passport `literature_corpus[]` entries (each carries `citation_key`, `source_acquired`, `source_verified_against_original`).
- The peer-file `<session>_human_read_log.yaml` (path computed as `<passport-path-parent>/<passport-stem>_human_read_log.yaml` per §3.6 round-5 R5-003 amend) — provides `human_read_source: true` for every `citation_key` the user has explicitly marked via `/ars-mark-read`.
Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.
---
## Dashboard Output Rules
1. Produce full version when user explicitly requests it
2. **Append simplified version to checkpoint notification after each stage completion**
Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.
## Interaction with existing features
- **Collaboration Depth Observer (v3.5.0):** fires on FULL/SLIM as before. Observer output is included in the checkpoint notification regardless of reset state. Observer state does NOT carry across resets; each fresh session observes only its own stage.
- **Compliance agent (v3.4.0):** `compliance_history[]` remains append-only and is consumed from the passport on resume. No change to Schema 12.
- **Sprint contract (v3.6.2):** reviewer sprint contracts load from the passport on resume (Phase 1 paper-content-blind stage remains valid across the reset boundary because the contract + paper metadata are carried in the passport).
- **Socratic reading probe (v3.5.1):** reading probe fires at most once per session. Across a reset boundary, the probe counter resets — the next session may fire its own probe. This is by design: each session is its own Socratic unit.
Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.
## Interaction with existing features
- **Collaboration Depth Observer (v3.5.0):** fires on FULL/SLIM as before. Observer output is included in the checkpoint notification regardless of reset state. Observer state does NOT carry across resets; each fresh session observes only its own stage.
- **Compliance agent (v3.4.0):** `compliance_history[]` remains append-only and is consumed from the passport on resume. No change to Schema 12.
- **Sprint contract (v3.6.2):** reviewer sprint contracts load from the passport on resume (Phase 1 paper-content-blind stage remains valid across the reset boundary because the contract + paper metadata are carried in the passport).
- **Socratic reading probe (v3.5.1):** reading probe fires at most once per session. Across a reset boundary, the probe counter resets — the next session may fire its own probe. This is by design: each session is its own Socratic unit.
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
### `caveats`
Non-empty array, items non-empty strings, minItems 1. The schema physically prevents an
empty caveats field. A benchmark report with no caveats either has no known limitations
(implausible for any real-world evaluation) or the author didn't think about limitations
(which disqualifies the report more than any specific limitation would).
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
without a stated `human_baseline.sample_size`. n=1 is worse than n=10, but it is
infinitely better than unstated.
- **Author-conducted with no caveat disclosure.** Using `author-conducted` is allowed.
Not noting it in `caveats` as a limitation is an editorial failure, not a schema
failure — but reviewers will notice.
Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.
Not noting it in `caveats` as a limitation is an editorial failure, not a schema
failure — but reviewers will notice.
- **Self-scoring with no warning acknowledgment.** `self-scored` produces a warning on
every validator run. If you publish a report with `self-scored` and don't address it
in caveats, the omission is visible in the JSON to anyone who checks.
The manifest includes many generic trigger phrases such as "full pipeline," "publish paper," and "research to paper," which are broad enough to match ordinary academic-assistance requests outside a tightly scoped invocation context. Overbroad activation can cause this orchestration skill to run unexpectedly, chaining multiple delegated stages and data flows when a user may have intended a narrower task.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
## Hermes Execution
The orchestrator (main Hermes agent):
1. **Stage 1**: Clarify user intent → select modes → create Project Plan
2. **Stage 2-6**: Sequential delegate_task calls, each loading the relevant skill
3. **Stage 7-9**: Lightweight parallel or sequential based on dependencies
4. **Stage 10**: Write final package to disk
The manifest description limits the skill to verifying references, citations, and data for factual accuracy. However, the documented behavior includes Phase D originality verification, including paragraph-level plagiarism screening and self-plagiarism checks, which is a distinct function beyond factual verification of citations and data.
The skill hardcodes verification sources and workflows that are biased toward English-language and globally indexed systems without user opt-in or locale-sensitive fallback. In an academic pipeline that may process multilingual or regional scholarship, this can systematically misclassify legitimate non-English or locally indexed references as NOT_FOUND or lower-confidence, causing false integrity failures and disproportionate exclusion of certain sources.
The manifest describes a 10-stage end-to-end pipeline: planning → deep-research → academic-paper → academic-paper-reviewer → revision → polish → ethics → disclosure → format → deliver. This file instead defines coordination around research, writing, integrity, review, revision, finalize, and completion, with no concrete planning, polish, ethics, disclosure, or deliver stages matching the manifest sequence.
The routing triggers use broad natural-language keywords like 'research', 'review', 'draft', and 'format' without strong exclusion logic. In an agent environment, this can cause accidental invocation or misrouting of user requests into a complex pipeline, potentially leading to unintended processing of user files and autonomous stage transitions.
The file states the orchestrator should only coordinate, but later instructs it to mutate draft content in place and participate in document-production behavior. This creates a dangerous authority mismatch: operators and downstream systems may trust the orchestrator as non-substantive while it is actually empowered to alter user materials, which can lead to unexpected content modification and reduced oversight.
The orchestrator instructs in-place draft mutation during cite-time provenance finalization but does not require explicit user-facing notice or consent before modifying user-authored content. Silent or unexpected edits to manuscripts can corrupt user work, introduce compliance text the user did not approve, or change publication-ready artifacts without clear accountability.
This markdown skill defines the agent's role and behavior but does not state when it should be invoked, what exact trigger phrases apply, or any exclusion conditions. For markdown files, missing specificity on trigger scope can cause unintended invocation by an orchestrator or user because the boundary between applicable and non-applicable requests is unclear.
The trigger list for entering external review mode includes broad natural-language examples plus an open-ended "etc.", which makes the activation boundary ambiguous. In an orchestrated multi-stage agent, ambiguous triggers can cause unintended mode switching, leading the system to parse arbitrary user text as reviewer feedback and perform the wrong workflow, which is a real prompt-scope/control issue even if not directly exploitable for code execution.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
Orchestrator obligations on resume:
- Locate the target `kind: boundary` entry by matching `hash`. Hard error if no match, or if a later `kind: resume` entry already carries `consumes_hash == <hash>` (double-resume is forbidden).
- Do NOT ask the user to re-summarize prior stages; the passport is authoritative. Load artifacts by reference (paths or IDs recorded in the entry).
- Honor the `verification_status` field. If `STALE` or `UNVERIFIED`, display a warning and prompt the user to re-verify before continuing. If `VERIFIED`, proceed without prompting.
- If `pending_decision` is set on the ledger entry, re-prompt the user for that decision BEFORE invoking any downstream stage. Display `pending_decision.question` and each option's `value`. After the user picks, look up the matching entry in `options[]` by `value`, then use that entry's `next_stage` and `next_mode` to determine actual routing. Record the chosen `value` as `chosen_branch` on the new `resume` entry. `next` on the boundary entry is advisory and is superseded by the matched option's `next_stage`. A user-supplied `stage=<n>` override on the resume command does NOT satisfy `pending_decision` — the decision prompt always fires when `pending_decision` is present. CLI `stage=`/`mode=` overrides still win over option routing if the user supplies them after the decision prompt.
- Emit a `### Resume Acknowledged` section at the start of the new session with: hash, source session `session_marker` + `generated_at`, recovered stage, and next-stage plan.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
Orchestrator obligations on resume:
- Locate the target `kind: boundary` entry by matching `hash`. Hard error if no match, or if a later `kind: resume` entry already carries `consumes_hash == <hash>` (double-resume is forbidden).
- Do NOT ask the user to re-summarize prior stages; the passport is authoritative. Load artifacts by reference (paths or IDs recorded in the entry).
- Honor the `verification_status` field. If `STALE` or `UNVERIFIED`, display a warning and prompt the user to re-verify before continuing. If `VERIFIED`, proceed without prompting.
- If `pending_decision` is set on the ledger entry, re-prompt the user for that decision BEFORE invoking any downstream stage. Display `pending_decision.question` and each option's `value`. After the user picks, look up the matching entry in `options[]` by `value`, then use that entry's `next_stage` and `next_mode` to determine actual routing. Record the chosen `value` as `chosen_branch` on the new `resume` entry. `next` on the boundary entry is advisory and is superseded by the matched option's `next_stage`. A user-supplied `stage=<n>` override on the resume command does NOT satisfy `pending_decision` — the decision prompt always fires when `pending_decision` is present. CLI `stage=`/`mode=` overrides still win over option routing if the user supplies them after the decision prompt.
- Emit a `### Resume Acknowledged` section at the start of the new session with: hash, source session `session_marker` + `generated_at`, recovered stage, and next-stage plan.
Forcing English-language confirmation without user choice can cause users to misunderstand gating decisions, branch selection, or integrity-check acknowledgments. In a pipeline with mandatory checkpoints and resume semantics, language mismatch increases the chance of erroneous consent or incorrect transitions, which is a workflow safety issue even if it is not a direct code-execution flaw.
The pause transition is triggered by broad natural-language phrases like "stop here," which can easily appear in normal academic discussion or quoted content. In an orchestrator that controls multi-stage execution and resume behavior, this creates a prompt/command ambiguity issue that can unintentionally alter pipeline state, disrupt execution, or cause incorrect checkpoint handling.
The manifest describes a 10-stage pipeline ending with deliver, but this file documents an added 'Stage 6' whose purpose is generating process-history records and collaboration evaluations. That is a substantive behavior expansion beyond the manifest description, not merely an implementation detail, because it adds a distinct output artifact and analysis function over user/AI behavior.
The protocol directs the agent to document the complete human-AI collaboration history for sharing/reporting, but does not require consent review, sensitive-data screening, or minimization. This can turn transient chat content into a durable disclosure artifact containing personal, confidential, or regulated information.
The protocol instructs the system to compile comprehensive user-provided content, including full session history and verbatim quotes, for later sharing/export. This creates a built-in data exfiltration channel from private conversation context into portable documents, even when those details are not necessary for the publication workflow.
No suspicious patterns detected.