T01 · Skill Instruction Hijacking
- Location
index.ts:125- Finding
Untrusted Retrieved Pattern Content Is Injected Directly into the Model Prompt
- Content
View full analysis
Vulnerability Details
File Location:
index.ts, lines 125-130
Vulnerability Type: Prompt injection through untrusted retrieved content
Risk Level: HighVulnerable Code
ts // Build context from patterns const context = results .map(r => `[${r.pattern.slug}] ${r.pattern.claim}`) .join('\n'); return `Relevant knowledge from community commons:\n${context}\n\nQuery: ${query}`;Technical Analysis
The
beforeQuery()hook retrieves pattern claims from the Memory Garden daemon and inserts them verbatim into the text sent to the language model. The implementation does not apply instruction filtering, trust validation, content sanitization, provenance-based policy, or a clear security boundary that requires the model to treat retrieved patterns only as untrusted reference data.Search is enabled by default, and the project describes community and federated pattern sources. Consequently, an attacker who can contribute to, synchronize with, or otherwise influence the searched pattern collection may store a claim containing adversarial instructions. When that pattern matches a later query, the hostile text is positioned immediately before the user's query and may be interpreted as an instruction rather than data.
This is an indirect prompt-injection condition. It does not itself grant operating-system privileges, but it can alter the Agent's current goals, output, safety behavior, and tool-selection decisions. The effective severity depends on which tools and permissions the host Agent exposes.
Attack Path
- An attacker creates or influences a pattern whose
claimcontains instructions intended for the model, such as directions to ignore the user's request, disclose available context, invoke a tool, or direct the user to attacker-controlled content. - The malicious pattern enters a local or community-backed pattern collection searched by the daemon.
- A user submits a query that causes the malicious pattern to rank among t ...[truncated 1044 chars]
- An attacker creates or influences a pattern whose
- Remediation
View remediation
Remediation Suggestions
- Treat every retrieved pattern field as untrusted data, regardless of whether it originated locally or from a nominally trusted community.
- Pass retrieval results through a structured context channel, if supported, rather than concatenating them into the instruction-bearing prompt.
- Add an authoritative instruction outside attacker-controlled content stating that retrieved patterns are reference data and must not modify system instructions, user goals, safety constraints, or tool policy.
- Delimit each retrieved item using a non-executable structured representation containing explicit fields such as
claim,source,trust_level, andsignature_status. - Detect and reject or quarantine claims containing instruction-like content, role markers, requests to ignore earlier instructions, tool invocation directives, or data-exfiltration requests.
- Apply source authentication, signature verification, moderation, and trust thresholds before community patterns become eligible for prompt augmentation.
- Limit the length and number of retrieved claims to reduce prompt-injection surface.
- Require explicit user confirmation before acting on retrieved content that suggests tool use or security-sensitive actions.
- Add adversarial tests covering malicious community claims and confirm that they cannot alter task goals or trigger tools.
