T02 ยท Agent Memory Poisoning
- Location
scripts/shrink.py:622- Finding
Unvalidated Vision-Model Output Is Persisted in Agent Session History
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill does what it claims, but it sends session images plus nearby conversation to Anthropic and permanently rewrites agent history with unvalidated model-generated text.
Review this skill before installing if your OpenClaw sessions may contain secrets, customer data, regulated records, or untrusted screenshots. Prefer explicit ANTHROPIC_API_KEY selection, use dry-run, limit scope with a single agent or session, avoid --all-sessions unless needed, and treat --redact as output redaction rather than proof that sensitive source data was never sent to Anthropic.
scripts/shrink.py:622Unvalidated Vision-Model Output Is Persisted in Agent Session History
scripts/shrink.py:367Redaction Is Applied Only After Sensitive Data Has Been Sent to Anthropic
The skill context makes this more dangerous because it operates on session history, where screenshots often contain account data, internal dashboards, error traces, and secrets, and the surrounding messages may explain or reproduce those secrets in text. Sending contextual excerpts to improve descriptions is functionally a context-leak pathway, not just image summarization.
def extract_surrounding_context(lines, line_idx, content, content_idx, context_depth=5):
"""Extract conversation context around an image for context-aware description."""
context_parts = []
# 1. Get text from the SAME message
The prompt explicitly instructs the model to extract and preserve all readable text, numbers, names, IDs, URLs, and other data visible in the image, which maximizes retention of sensitive content by default. In a session-shrinking tool, that means the compressed output may still preserve secrets and PII while also transmitting them to a third party during processing.
This call site activates the context-leak behavior during normal processing of each image, making the issue operational rather than theoretical. Because the tool is designed to batch-process session files, a single run can repeatedly exfiltrate neighboring conversation content for many images at scale.
src = dedup_cache.get(f"{fp}_src", "earlier")
print(f" โป๏ธ Duplicate detected โ reusing description from {src}")
else:
# Extract conversation context
context = extract_surrounding_context(lines, line_idx, content, content_idx, context_depth)
if log and verbose and context:
context_preview = context.split("\n")[0][:100]
Without declared permissions the skill's intent is opaque and cannot be validated.
The trigger phrases are broad enough to match ordinary discussion about shrinking context or reducing image bloat, which can cause the skill to activate in situations where the user did not intend a destructive session rewrite. Because the skill reads session history, sends images and conversation context to an external API, writes modified JSONL files, and may prompt for a gateway restart, accidental invocation has meaningful privacy and integrity consequences.
Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.
VISION_MODEL_OAUTH = "claude-haiku-4-5" # OAuth tokens (Max/Max Pro) โ only Haiku works direct
VISION_MODEL_APIKEY = "claude-sonnet-4-6" # API keys โ full model access
MAX_DESCRIPTION_TOKENS = {"standard": 500, "full": 800}
ANTHROPIC_API_URL = "https://api.anthropic.com/v1/messages"
ANTHROPIC_VERSION = "2023-06-01"
MODEL_PRICING = {
The context extraction routine collects prior user messages, same-message text, and the next assistant response, then packages that material into the prompt sent with the image. This broadens data exposure well beyond the image itself and can leak unrelated confidential session content, credentials, or personal information to the external model provider.
The script sends base64 image contents plus extracted surrounding conversation context to Anthropic's external API, but it does so as the core workflow rather than after an explicit opt-in disclosure. Because session images and nearby messages may contain secrets, personal data, or proprietary content, users can unknowingly exfiltrate far more data than expected.
This request performs the actual external transmission of sensitive payload data to the Anthropic API. While network access is expected for a vision summarization feature, it becomes dangerous here because the payload includes raw image data and potentially sensitive conversational context without strong safeguards or default minimization.
}
try:
resp = requests.post(ANTHROPIC_API_URL, headers=headers, json=payload, timeout=30)
# Check if we should failover
if resp.status_code in FAILOVER_CODES and key_idx < len(api_keys) - 1:
The script permanently rewrites session files with model-generated descriptions, causing sensitive information extracted from images to persist in plain text in local history. That can increase future exposure because data that was previously buried in base64 blobs or images becomes searchable, easier to copy, and may remain even when redaction was not enabled.
if log and verbose:
print(f"\n๐พ Backup saved: {os.path.basename(backup_path)}")
# Write modified file
with open(session_file, "w") as f:
f.writelines(lines)
The script automatically harvests Anthropic credentials from both environment variables and local auth profile files, including agent-scoped credentials, without a strong disclosure boundary. This expands the trust surface and can cause the tool to silently use credentials the user did not intend to expose for this operation.
No suspicious patterns detected.