T05 · Unauthorized Access and Privilege Escalation
- Location
scripts/smartocr_from_session.py:40- Finding
Automatic Session Selection Can Upload Images from an Unrelated Conversation
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This OCR skill has a coherent purpose, but it needs Review because it can read OpenClaw session history and upload selected images to a configurable OCR endpoint without strong scoping or confirmation.
Review before installing. Use this only for documents you are comfortable sending to the configured SmartOCR service, prefer the default HTTPS endpoint or a trusted HTTPS host, avoid the session helper unless you can verify the exact session and image being processed, and use a limited API key. The skill should add stronger session scoping, endpoint validation, and explicit data-transfer confirmation before it is treated as low-risk.
scripts/smartocr_from_session.py:40Automatic Session Selection Can Upload Images from an Unrelated Conversation
scripts/smartocr.py:39Configurable Plaintext Endpoint Can Expose API Credentials and Sensitive Documents
scripts/smartocr.py:2Unpinned Runtime Dependency Creates a Supply-Chain Exposure
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
Returns:
包含 ocr_type 和 content 的字典
"""
resp = requests.post(
f"{api_url.rstrip('/')}/api/ocr",
headers={"X-API-Key": api_key},
json={"image": image},
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
def ocr(image_data, api_url, api_key, timeout=60):
"""调用 SmartOCR API。"""
resp = requests.post(
f"{api_url.rstrip('/')}/api/ocr",
headers={"X-API-Key": api_key},
json={"image": image_data},
The documented purpose says the skill processes image URLs and local files, but the skill also states it can scan ~/.openclaw/agents/{agent}/sessions/ and extract recent uploaded images from session JSONL files. This is dangerous because it expands behavior into undeclared access of historical conversation data, potentially collecting sensitive user images and transmitting them to a third-party OCR service without sufficiently explicit disclosure or per-use authorization.
The skill documents and enables access to environment variables, local files, and remote network endpoints, but it does not declare any explicit tool scope or permissions boundary. This is dangerous because the agent can read local image files, access session-related paths, and send data to an external OCR API without a clearly declared least-privilege contract, making review, containment, and user consent weaker.
This CLI uploads either a user-supplied image URL or the full contents of a local image file to a remote OCR service, and those images may contain sensitive PII such as vehicle registration documents, receipts, or invoices. The script description and runtime behavior do not present a clear, explicit privacy/data-transfer warning before transmission, which increases the risk of users unknowingly exfiltrating sensitive documents to a third party.
The script behavior materially differs from the skill description: instead of only handling explicit user-provided image URLs or local files, it searches OpenClaw session history and extracts recent images automatically. That can cause unintended access to prior conversation attachments and send sensitive documents to a third-party OCR service without clear user intent.
Reading OpenClaw session files from disk gives the skill access to historical user content beyond the minimum needed for OCR. In this context, session data may include highly sensitive vehicle documents, invoices, or other attachments, so broad disk access increases privacy and data-exfiltration risk.
The code sends extracted session images to an external OCR API, but there is no explicit user-facing warning or consent flow explaining that conversation attachments will leave the local environment. Because the skill targets vehicle licenses and receipts/invoices, the transmitted images are likely to contain personal, financial, or regulated data, making this especially sensitive.
The natural-language description, help text, and runtime messages are all fixed in Chinese, with no indication that users can select another language. Under the stated policy, forcing a specific language without user opt-in is a locale/language policy concern.
The natural-language descriptions, usage text, and error/help strings in this script are presented in Chinese only. There is no indication that the tool is region-specific or that users can opt into another language, which may violate language/locale policy expectations.
No suspicious patterns detected.