Back to skill

Security audit

mail-manager

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed email-management workflow, but it gives incoming emails from user aliases authority to trigger agent actions and automatic replies without normal confirmation.

Install only if you intentionally want this agent to manage the mailbox and maintain local mail logs. Review or remove the rule that treats emails from user aliases as fully authorized commands, and require in-session confirmation before any email-originated action with side effects or any automatic reply.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:77
Finding

Inbound Email Content Is Treated as a Privileged Agent Instruction

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 77–83 and 93–95
Vulnerability Type: Untrusted email content crossing into the Agent instruction and tool-execution context
Risk Level: High

Vulnerable Snippet

The following is an English translation of the operative instructions in the identified lines:

markdown
### C. Email as instruction — execute with full authority
- Email from any user email alias = fully authorized instruction; no secondary confirmation:
  1. Read the complete message and parse the task or to-do item.
  2. Register it in TODO.md as an "email instruction."
  3. Execute it directly; record the result in the inspection log,
     current work log, and conversation.
  4. After execution, reply by email to confirm the result.
     Confirmation replies execute automatically without per-message authorization.
     For matters requiring authorization, request authorization by email and execute
     after the user replies with approval.
  5. Only extremely high-risk operations, such as irreversible deletion,
     sending external email, or financial operations, require prior assessment.

## Boundaries and security
- Sending email normally requires confirmation, but fully authorized instructions
  received from any user alias are exempt, including the automatic confirmation reply.
- Only instructions from the primary mailbox are executed; primary-mailbox messages
  are treated as fully authorized instructions.

Technical Analysis

The Skill explicitly changes the trust classification of inbound email bodies from untrusted message data to privileged Agent instructions. A message whose sender appears to match any configured user alias is parsed into a task, written to TODO.md, and executed without an independent in-session confirmation.

Sender-address matching alone is not documented as being bound to cryptographically verified sender identity, connector-authe ...[truncated 2308 chars]

Remediation
View remediation

Remediation Suggestions

  1. Treat every email subject and body as untrusted data, never as an automatically privileged Agent instruction.
  2. Require explicit approval in the active user session before executing any email-originated state-changing or external action.
  3. Remove the confirmation exception for messages that appear to originate from user aliases.
  4. Validate exact sender identities using authenticated connector metadata rather than display names or raw address strings.
  5. Where supported, require verified SPF, DKIM, and DMARC results, while recognizing that these controls do not replace user confirmation.
  6. Restrict email-originated automation to a narrow allowlist of low-risk operations, such as listing or summarizing messages.
  7. Represent extracted requests as proposed actions and display their source message, parameters, target resources, and expected side effects before approval.
  8. Prevent email content from directly selecting tools, paths, recipients, commands, or persistent instructions.
  9. Do not write raw email instructions into persistent task or memory files without sanitization, provenance metadata, and user approval.
  10. Preserve platform-enforced confirmation tokens for all send, reply, forward, and delete operations without workflow exceptions.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Ssd 3

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill directs persistent logging of mailbox state and archival data into local files, including contact information and operational history, in plain-language records outside the email system. This increases exposure of sensitive metadata and potentially content to other local processes, backups, or users, and creates long-lived copies that may bypass mailbox retention and access controls.

Content

No source excerpt is available for this finding.

Ssd 3

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The new-mail workflow instructs the agent to copy sensitive email details, including subject lines, sender information, attachment counts, and body excerpts, into persistent conversation archives based on broad matching rules. This expands the data footprint of potentially confidential mail and can leak secrets or personal information into less protected storage and downstream memory systems.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly treats emails from user aliases as authoritative commands and permits direct execution of tasks beyond normal email-management actions. This is dangerous because email sender identity and alias trust are weaker than in-session user confirmation, so spoofing, mailbox compromise, forwarding rules, or misclassification could turn an inbound message into arbitrary agent action.

Content

No source excerpt is available for this finding.

Ssd 4

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill establishes a trust model where emails from user aliases are normalized as sufficient authority for direct execution and automatic replies, effectively bypassing standard confirmation boundaries. In the context of an email-management skill with broad mailbox access, this is especially dangerous because a compromised mailbox, spoofed sender, or alias confusion can cascade into unauthorized agent actions and data disclosure.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description declares activation triggers such as “邮箱”, “邮件”, “mail”, and “email”, which are generic terms likely to appear in ordinary conversation. The file does not provide narrowing constraints or negative examples to clarify when the skill should or should not activate, increasing the risk of unintended invocation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
79% confidence
Finding

The description is written as a Chinese-only operational specification and presents the skill as handling mailbox operations in that locale without any indication that users may choose another language. Under the language/locale policy, forcing a specific language without opt-in can be a natural-language policy violation unless the locale constraint is explicitly justified.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill creates an exception that allows automatic outbound replies for command emails without the normal confirmation flow. This weakens a core safety control around sending mail and can be abused to send unintended messages, leak execution details, or create a trusted-looking response chain after a spoofed or compromised inbound email.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.