Back to skill

Security audit

Academic Paper

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a coherent academic-writing workflow, but it includes an under-disclosed optional path that can use shell commands and API keys to send manuscript or reference material to third-party AI providers.

Install only if you are comfortable with a multi-agent academic drafting workflow processing your manuscript. Do not enable ARS_CROSS_MODEL or store provider API keys in shell profile files unless you understand that manuscript/reference content may be sent to OpenAI or Google and may incur API costs. Review generated drafts for hidden citation comments and automatic citation edits before sharing or submitting.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
shared/cross_model_verification.md:49
Finding

Optional Cross-Model Protocol Uses Undeclared Shell and Network Access with Insecure Credential Handling

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (61)

Hidden Instructions

High
Category
Prompt Injection
Confidence
97% confidence
Finding

This instructs the agent to emit hidden HTML comment payloads alongside visible citations. Hidden-channel outputs are dangerous because downstream systems, reviewers, or users may not see them, yet other components can parse and act on them; this creates an opportunity for covert metadata transmission, policy bypass, or prompt/data smuggling across agent boundaries.

Content

Scanner excerpt · agents/draft_writer_agent.md (reported line 478)May include surrounding context.

md
When emitting any citation in the draft body, write the citation in two layers:

1. **Visible layer**: standard author-year form (e.g. `Smith (2024)` or `(Smith, 2024)`).
2. **Hidden layer**: immediately after the visible form, append an HTML comment of the shape `<!--ref:slug-->`, where `slug` is the `citation_key` already present in the corpus context provided in this prompt.

Examples: `Smith (2024) <!--ref:smith2024-->` or `(Smith, 2024)<!--ref:smith2024-->`.

Hidden Instructions

High
Category
Prompt Injection
Confidence
98% confidence
Finding

This expands the hidden-channel mechanism by requiring both a hidden reference slug and a hidden anchor marker after every citation, creating a richer covert instruction/data plane embedded in natural-language output. In a multi-agent pipeline, that materially increases the risk of invisible cross-stage signaling, parser manipulation, and metadata injection, especially because the hidden markers are mandated at scale throughout the draft.

Content

Scanner excerpt · agents/draft_writer_agent.md (reported line 497)May include surrounding context.

Every visible citation in the draft body MUST be followed by BOTH a slug marker AND an anchor marker:

text
<visible> <!--ref:slug--><!--anchor:<kind>:<value>-->

Anchor kinds (closed enum):

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · shared/benchmark_report_pattern.md (reported line 101)May include surrounding context.

md
### `caveats`

Non-empty array, items non-empty strings, minItems 1. The schema physically prevents an
empty caveats field. A benchmark report with no caveats either has no known limitations
(implausible for any real-world evaluation) or the author didn't think about limitations
(which disqualifies the report more than any specific limitation would).

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · shared/benchmark_report_pattern.md (reported line 154)May include surrounding context.

md
without a stated `human_baseline.sample_size`. n=1 is worse than n=10, but it is
  infinitely better than unstated.

- **Author-conducted with no caveat disclosure.** Using `author-conducted` is allowed.
  Not noting it in `caveats` as a limitation is an editorial failure, not a schema
  failure — but reviewers will notice.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · shared/benchmark_report_pattern.md (reported line 158)May include surrounding context.

md
Not noting it in `caveats` as a limitation is an editorial failure, not a schema
  failure — but reviewers will notice.

- **Self-scoring with no warning acknowledgment.** `self-scored` produces a warning on
  every validator run. If you publish a report with `self-scored` and don't address it
  in caveats, the omission is visible in the JSON to anyone who checks.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill advertises multiple broad natural-language triggers such as 'write paper', 'academic paper', and 'guide my paper', including bilingual variants. These phrases are common user requests and can cause the skill to activate in situations where the user did not explicitly intend to invoke this multi-agent workflow, increasing the chance of unintended delegation and file processing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill hard-codes Traditional Chinese as the non-English abstract language and does not offer user, journal, or locale-based selection. In multilingual or cross-region contexts, this can cause incorrect outputs, exclusion of users who need another language variant, and accidental non-compliance with submission requirements.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The agent is instructed to 'auto-correct detected errors' directly in the draft, despite a description centered on verification and flagging. Silent content mutation creates integrity risk: the agent may alter user-authored material, citations, or claims without review, potentially introducing factual errors, attribution mistakes, or unwanted changes that the user does not notice.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Automatic modification of the draft without a clear warning undermines user control over their document and can cause unnoticed integrity loss. In a paper-writing pipeline, even small automatic citation edits can cascade into broken references, altered meaning, or accidental fabrication if the model guesses missing metadata.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest description limits the agent to verifying citations against journal format requirements and flagging non-compliant entries. These lines add broader scholarly-integrity functions—self-plagiarism detection and checking references against Retraction Watch—that go beyond citation formatting/compliance and are not reflected in the stated role.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

Checking a retraction database is a literature-validity and publication-integrity task, not an obvious requirement of formatting citations or cross-checking reference consistency. For a citation compliance agent, this is a separate evaluative capability outside the manifest’s stated purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The instruction to auto-correct 'silently' is risky because it removes transparency and prevents the user from validating what changed. In an academic context, silent edits to citations or references can create compliance, authorship, and accuracy problems that are difficult to detect after the fact.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

A protocol that tells the agent to fix errors directly in the draft, combined with only post-hoc logging, is unsafe because it prioritizes mutation over consent. This is especially dangerous here because the same file also authorizes source-related corrections, making it plausible that the agent could alter references or citation mappings incorrectly and overwrite the user’s original content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The instruction says Chinese references must appear before English references "per Taiwan academic convention," which enforces a specific regional convention. The file does not clearly frame this as optional, user-selected, or limited to a documented region-specific workflow, so it risks a language/locale policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file includes prescriptive natural-language instructions for a specific locale/language register: 'Additional rules for Chinese academic register.' Under the policy, forcing a specific language or locale without explicit user choice is a violation unless the constraint is documented as optional or justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The markdown instructs that for zh-TW content the formatter should use XeLaTeX with CJK font support, and elsewhere hard-codes Traditional Chinese font choices and language labels. This reflects a locale-specific constraint in the skill instructions without stating that the user must explicitly opt into that locale behavior at the moment of formatting.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

Automatically detecting the user's language to set defaults before explicit consent introduces a privacy and policy concern because it infers a user attribute from content without opt-in. In this academic-writing context the impact is limited, but it can still lead to unwanted profiling, incorrect locale assumptions, or user surprise if language-based defaults are silently applied.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The plan-mode trigger phrases are broad enough that normal user requests like 'help me plan my paper' may route the workflow into a reduced 3-question intake and bypass the fuller 11-step interview. In this skill, that can cause incorrect mode selection, incomplete parameter collection, and unintended handoff behavior, which is a real security/quality issue because downstream agents may act on underspecified or misclassified user intent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description says the skill 'Designs the literature search strategy and manages source selection for the paper' but does not specify concrete invocation phrases, boundaries, or exclusion conditions. In a manifest file, this broad natural-language description could match many ordinary research-related requests and create unintended activation overlap with other writing or research skills.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The text says the agent 'is activated in Phase 1' but does not define what concrete condition places a request into Phase 1 or how that differs from adjacent agents' responsibilities. This leaves the activation condition ambiguous and may cause uncertain or unintended invocation in multi-agent workflows.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The instruction 'Default to English primary + attempt Chinese keyword search' imposes language behavior without explicit user choice. The policy allows locale or language constraints only when the user is offered a choice or the constraint is clearly justified and documented.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger list includes broad natural-language phrases such as "help me with my revision" and "revision roadmap," which can plausibly appear in ordinary conversation outside the intended skill-selection context. This raises the risk of unintended invocation, causing the agent to process sensitive reviewer comments or drafts when the user did not explicitly intend to enter this mode.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Line L179 specifies 'Abstract fixed at 250 words (EN) or 400 characters (zh-TW)', which embeds a language/locale-specific behavior in the instructions. The file does not indicate that the user may choose the language or that the locale restriction is justified for a region-specific use case.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This markdown file is governed by SQP-3, which applies to all file types and includes language or locale policy violations. The title, description, and example user input/output all prescribe Chinese/Traditional Chinese as the required language, but the file does not present this as an opt-in choice or explain a region-specific necessity.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

This markdown example presents all mentor prompts and workflow instructions exclusively in English, despite the topic being specific to Taiwan and involving potentially multilingual users. The file provides no opt-in or alternative language guidance, which can amount to a language/locale policy violation when a skill implicitly forces one language.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.