Back to skill

Security audit

WA Relay

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent WhatsApp relay/firewall skill, but it can relay and retain private message content, so users should review privacy expectations before installing.

Install only if you want the agent to read, transcribe, forward, and locally log third-party WhatsApp messages for the owner. Configure the owner number carefully, require confirmation before outbound replies, and periodically review or delete wa-relay-log.md if messages may contain sensitive information.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (10)

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · SKILL.md (reported line 26)May include surrounding context.

md
### Outbound Flow (Owner → Third Party)

1. The owner instructs the agent: "Tell [contact] that..." or "Reply to [contact] with..."
2. The agent uses the `message` tool to send the message to the third party
3. The agent confirms delivery to the owner

### Message Log

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · SKILL.md (reported line 102)May include surrounding context.

md
Natural language commands the agent should recognize:

- "Reply to Martín: [message]" → Send message to Martín
- "Tell Banana that..." → Send message to Banana
- "What did Martín say?" → Check wa-relay log
- "Show me recent messages" → Summarize recent third-party messages

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

The command 'Forward that to Martín' is riskier because 'that' can ambiguously refer to prior context, memory, or sensitive content not intended for external sharing. Without strict scoping and confirmation, this creates a plausible path for accidental exfiltration of internal conversation history, owner context, or other private data to a third party.

Content

Scanner excerpt · SKILL.md (reported line 103)May include surrounding context.

md
Natural language commands the agent should recognize:

- "Reply to Martín: [message]" → Send message to Martín
- "Tell Banana that..." → Send message to Banana
- "What did Martín say?" → Check wa-relay log
- "Show me recent messages" → Summarize recent third-party messages
- "Forward that to Martín" → Forward last relevant content to Martín

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill instructs the agent to transcribe audio, describe media, and forward that content to the owner, but the description does not prominently warn that incoming media may be transformed and relayed. That omission can violate user expectations and increase exposure of sensitive voice or media content received from third parties.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill requires relaying full third-party message contents, including audio transcripts and media descriptions, to the owner. Even if intended as a relay, this broad collection and transmission of full content increases the chance of exposing sensitive personal data, secrets, or regulated information without minimization.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Maintaining a running conversation log in persistent memory stores a durable record of third-party communications and replies, including identifiers and timestamps. Persistent storage materially increases privacy risk, blast radius after compromise, and the chance of secondary misuse by other skills or future prompts.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill persistently logs third-party WhatsApp messages, names, numbers, and replies in a local memory file, but the description does not clearly warn about this retention behavior. This creates a privacy and data-governance risk because sensitive personal communications may be stored longer than expected and later surfaced to the owner or other agent functions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The example explicitly instructs the agent to forward and log private WhatsApp messages from third parties to the owner, but provides no notice, consent model, retention policy, or data-handling constraints. In a relay/firewall skill, this context makes the issue more significant because the core function is systematic interception, summarization, and storage of other people's communications, which can create privacy, compliance, and trust risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The audio example directs the agent to transcribe a voice message and share the transcription with the owner without any warning that audio will be converted to text and disclosed. This is dangerous because transcription increases sensitivity and searchability of content, potentially exposing information the sender expected to remain in ephemeral audio form; in this skill, that behavior is part of the designed workflow, so the risk is amplified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language examples and expected outputs are consistently in Spanish, with no indication that other languages are supported or that Spanish is optional. This can be a language-policy issue when a skill constrains outputs to a specific language without documented user choice or justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.