T01 · Skill Instruction Hijacking
- Location
src/agent_memory/core/models.py:73- Finding
Persistent prompt injection through untrusted long-term memory
- Content
View full analysis
list[dict[str, str]]: """Build the final OpenAI-format message list.""" messages: list[dict[str, str]] = [] # 1. System prompt + entity cards ("档案袋") system_text = self.system_prompt if self.entity_cards: cards = "\n".join( f"- {e.to_card_text()}" for e in self.entity_cards ) system_text += ( f"\n\n=== 用户档案(结构化记忆,请优先信任) ===\n{cards}" ) # 2. Retrieved episodes for context if self.retrieved_episodes: episodes_text = "\n\n---\n".join( e.to_context_text() for e in self.retrieved_episodes ) system_text += ( f"\n\n=== 历史相关片段(语义检索,仅供参考) ===\n{episodes_text}" ) ``` The corresponding persistent write path accepts arbitrary values: ```python def remember(self, key_or_content: str, value=None, evidence="", confidence=1.0, tags=None, source="conversation", expires_in_days=None): if value is not None: self._entity.upsert(EntityCard(key=key_or_content, value=value, evidence=evidence, confidence=confidence)) return key_or_content fact = self._entity.add_fact(key_or_content, tags=tags, source=source, confidence=confidence, ttl_days=expires_in_days) return fact.id ``` ### Technical Analysis The application treats persistent entity values and retrieved conversation episodes as trusted system-level instructions. `remember()` accepts arbitrary strings without provenance validation, instruction detection, escaping, or a separation between executable instructions and reference data. `MemoryContext.build_messages()` then con ...[truncated 1712 chars]- Remediation
View remediation
