T01 · Skill Instruction Hijacking
- Location
SKILL.md:207- Finding
Untrusted Remote Messages Can Influence Access to Local Personal Data
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, lines 207-213
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: MediumCode Snippet:
markdown - **`type: "question"`** — show: **"[from_name]** (via [circle]) asks: [query]". If `image_url`, show it. If the message has a `poll`, display options as a numbered list and ask the human to pick. Otherwise draft a full answer (check Peeps, Nooks, Pages, Vibes, Digs first). Ask: **"send or discard?"** If sending and `open_to_connections` is false, warn: _"Your profile is closed — the asker won't get a link to connect with you. Open up at haah.ing/profile, or send anyway?"_ Send → `POST /messages/:id/reply` · Discard → `POST /messages/:id/pass` - **`type: "dm"`** — show: **"DM from [from_name]:** [text]". Ask: _"Want to reply?"_ If yes, draft, confirm, and `POST /messages/:id/reply`. If `has_more` is true: _"Want to see more?"_ → `GET /messages?all=true`.Technical Analysis
Question and DM bodies originate from remote users and are therefore untrusted input. The Skill instructs the agent to use an inbound question to draft an answer after consulting the local Peeps, Nooks, Pages, Vibes, and Digs Skills. It does not establish a trust boundary that requires remote content to be treated solely as data, nor does it prohibit following instructions embedded in messages or accessing unrelated private records.
An attacker could phrase a question as an instruction to search local sources for sensitive information and include that information in the response. Although the Skill requires the user to approve a reply before it is sent, approval occurs only after potentially sensitive data has already been collected into a draft. A plausible or misleading draft could also cause a user to approve disclosure inadvertently.
Attack Path
- An attacker gains the ability to send a circle question or direct message to the user.
- The scheduled heartbeat retr ...[truncated 1201 chars]
- Remediation
View remediation
Remediation Suggestions
- Explicitly classify every remote message field, including
query,text, names, attachment contents, and image-derived text, as untrusted data. - Instruct the agent never to follow commands, policies, tool requests, or disclosure requests embedded in remote messages.
- Restrict local lookups to information directly relevant to the substantive question and prohibit access to credentials, secrets, private messages, authentication material, or unrelated personal records.
- Require the agent to identify which local sources it intends to consult and obtain permission before accessing sensitive sources.
- Add a separate disclosure confirmation that displays the exact response, the local sources used, and any potentially sensitive facts included.
- Apply data minimization and redact unnecessary names, identifiers, addresses, contact details, and private notes.
- Treat attachment text and image content with the same injection protections as message text.
- Explicitly classify every remote message field, including
