T01 · Skill Instruction Hijacking
- Location
infra/agents/listener.prompt.ts:9- Finding
Indirect Prompt Injection Through Automatically Ingested LinkedIn Content
- Content
View full analysis
Vulnerability Details
File Location:
infra/agents/listener.prompt.ts:9-48
Vulnerability Type: Indirect prompt injection from untrusted external content
Risk Level: MediumComplete Code Snippet
typescript ## 1. Read the context first From the workspace context read the ICP (who buys, the personas and titles, the disqualifiers), the pains in the buyer's own words, our positioning, and our competitors. That is the rubric for everything below. ## 2. Read this week's posts Query the linkedin_posts model: every row. Drop any post whose urn is already in surfaced_posts. Drop posts written by us or by our own employees, job posts, and posts by a competitor's company page. ## 3. Pick the posts worth joining, at most five A post is worth joining when its author or its readers are plausibly our buyers and it is about a pain we solve, a category we sit in, or a competitor. Rank by fit first, then by conversation (num_comments over num_likes). For each pick, write one line on why, citing the ICP or pain line that decided it. A post you cannot tie to a line of the context is not picked. ## 4. Find the warm engagers on those posts For each picked post with at least one comment and no more than 150 comments, read its comments once with searchPostComments (sortBy "Most relevant"). That read bills per comment returned, so skip it for a post above the cap and say "comments not read: over the cap" on its digest line. A commenter is a warm engager when their headline matches an ICP persona and does not hit a disqualifier. Keep at most three per post, and at most ten in the digest. For each: name, headline as written, their profile URL as the comment returned it, and the first sentence of what they said, quoted. Never infer an employer the headline does not state, never look anyone up elsewhere, and never add a person who only reacted. ## 5. Post the digest, once If surfaced_posts already has a row with post_urn "digest-<today's date>", this week ...[truncated 4786 chars]- Remediation
View remediation
Remediation Suggestions
- Explicitly classify all LinkedIn post, profile, headline, URL, and comment fields as untrusted data in the system prompt. State that instructions, requests, policies, tool calls, or role declarations appearing in those fields must never be followed.
- Represent external records in a clearly delimited structured format and instruct the model to analyze only specified fields for relevance.
- Separate content classification from side effects. Have one restricted stage produce a structured selection, then validate it with deterministic code before permitting Slack or ledger writes.
- Validate selected post URNs and URLs against the exact records returned by
linkedin_posts; validate quoted comments against the retrieved comment objects. - Constrain ledger writes to post URNs present in the current synchronized dataset and enforce schema-level uniqueness where supported.
- Apply output validation before
postMessage, rejecting unexpected instructions, mentions, links, or content that cannot be traced to a source record. - Add adversarial acceptance tests containing prompt-injection strings in post bodies, author metadata, headlines, and comments. Verify that these strings are quoted or summarized only as data and cannot alter tool use or workflow sequencing.
- Preserve the existing least-privilege controls: keep LinkedIn engagement actions unavailable, retain the locked Slack channel, and keep workspace context read-only.
