Back to skill

Security audit

OpenClaw Continuity

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a real continuity/follow-up package, but it ships under-disclosed file-trashing and dynamic embedding paths that could affect user files or expose conversation text and API keys.

Review this before installing. The continuity features are plausible, but operators should remove or tightly gate file_output_sop.py, pin dependencies, keep external embedding disabled unless the provider and endpoint are explicitly approved, verify PERSONAL_HOOKS_EMBEDDER_ROOT points only to trusted code, and publish clear user controls for stored profile, schedule, carryover, audit, and daily-memory data.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/file_output_sop.py:40
Finding

Arbitrary File Relocation and Automatic Trashing

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned Runtime Dependency Installed from a Mutable Package Source

Content
View full analysis
=1.8.0 ``` The documented installation command is: ```bash python3 -m pip install -r requirements.txt ``` ### Technical Analysis The dependency declaration specifies only a minimum version. Consequently, installation may select any present or future `send2trash` release that satisfies `>=1.8.0`. The project does not provide an exact version lock or an integrity hash in the reviewed dependency file. This weakens reproducibility and allows the effective installed code to change after the Skill package has been audited. The dependency is used by `scripts/file_output_sop.py`, an auxiliary file-trashing helper rather than the core continuity engine. A lower-bound dependency is not proof that the current upstream package is malicious. The security issue is that future installation behavior is mutable and is not cryptographically tied to the version reviewed by maintainers. ### Attack Path 1. A user follows the documented installation instructions. 2. `pip` resolves `send2trash>=1.8.0` using the configured package index. 3. A newer compromised release, a compromised package index, or a maliciously substituted artifact satisfies the version range. 4. The unreviewed package is downloaded and installed into the environment used by the Skill. 5. Malicious package installation or import-time behavior executes with the privileges of the installing or running account. ### Impact Assessment The maximum scope is the Python environment and operating-system account used to install or run the Skill. Depending on the behavior of a compromised dependency, consequences could include: - Execution of arbitrary Python code. - Access to OpenClaw state and configuration readable by the account. - ...[truncated 351 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/personal_hooks.py:854
Finding

API Credential and Conversation Data Passed to Dynamically Selected Embedder Code and Endpoint

Content
View full analysis
Optional[dict]: if not OPENCLAW_CONFIG_PATH.exists(): return None try: config = json.loads(OPENCLAW_CONFIG_PATH.read_text(encoding="utf-8")) except Exception: return None embedding = ( ...[truncated 5258 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (51)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Host semantic configuration management on disk is broader than the stated role of a follow-up classification skill. Persisting and initializing semantic categories/policy overrides can materially change detection behavior and policy outcomes, so omitting that from the declared behavior reduces transparency and can lead to unsafe or surprising automation.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Host semantic configuration management on disk is broader than the stated role of a follow-up classification skill. Persisting and initializing semantic categories/policy overrides can materially change detection behavior and policy outcomes, so omitting that from the declared behavior reduces transparency and can lead to unsafe or surprising automation.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Host semantic configuration management on disk is broader than the stated role of a follow-up classification skill. Persisting and initializing semantic categories/policy overrides can materially change detection behavior and policy outcomes, so omitting that from the declared behavior reduces transparency and can lead to unsafe or surprising automation.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

Host semantic configuration management on disk is broader than the stated role of a follow-up classification skill. Persisting and initializing semantic categories/policy overrides can materially change detection behavior and policy outcomes, so omitting that from the declared behavior reduces transparency and can lead to unsafe or surprising automation.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 114)May include surrounding context.

md
- Script: `scripts/personal_hooks.py`

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 176)May include surrounding context.

md
- Script: `scripts/personal_hooks.py`

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 180)May include surrounding context.

md
- Script: `scripts/personal_hooks.py`

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The implementation is a file mover/trash handler, which materially differs from the manifest's stated purpose of continuity and follow-up management. This mismatch is dangerous because it hides surprising filesystem side effects inside a skill that would not reasonably be expected to alter user files, increasing the chance of misuse or stealthy data disruption.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The script resolves an arbitrary source path, moves the referenced file into a workspace-controlled directory, and then immediately sends that moved file to the trash. In a skill described as conversational continuity/follow-up management, destructive file handling is unexpected and can cause loss or concealment of user data if invoked on important files.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The manifest explicitly says this package is not a generic cron sender and should use contextual continuity rules to decide whether follow-up should appear. However, the code adds a standalone recurring-task subsystem that stores arbitrary scheduled tasks and turns overdue entries into reminder hooks on heartbeat ticks, which is a generic scheduler/reminder capability rather than continuity-specific follow-up.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill loads external embedding configuration, resolves an API key, and spawns a Node subprocess to compute embeddings via an external provider. In a continuity/memory skill, this can transmit user conversation content to a third-party service and expands the trusted computing base with cross-runtime execution, increasing confidentiality and supply-chain risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README explicitly advertises persistence of staged/tracked conversation state and daily-memory writeback, but it does not pair those claims with a clear privacy/data-retention warning, consent expectation, or guidance on handling sensitive user content. In a continuity/follow-up skill, retained conversational context may include personal routines, commitments, sensitive events, and inferred behavioral patterns, so under-disclosure can cause operators to deploy memory features without understanding the privacy risk.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill documents shell execution plus persistent state/config access, but does not declare an explicit tool scope such as allowed tools or permissions. That creates an authorization ambiguity: a host may expose broader capabilities than intended, and users/reviewers cannot easily verify the operational boundary of a skill that can read environment variables, execute commands, and write files.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

Preserving prior turns across new sessions is core to the skill, but it creates cross-session data retention and context-leak risk, especially if carryover includes sensitive or misclassified content. Because this package is specifically designed to reattach unresolved threads, the context makes the behavior intentional, yet still security-relevant if boundaries, expiry, and user control are weak.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill instructs deterministic extraction and durable storage of personal details such as timezone, sleep schedule, relationship, and use case into multiple files. In a continuity skill, this is contextually expected, but it still creates privacy and retention risk because sensitive profile data is persisted across sessions and may be consumed by other components without explicit per-field consent or retention controls.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The description explicitly states the comic is bilingual, and the visible text mixes English with Traditional Chinese throughout the asset. There is no indication that language choice is optional or justified by a region-specific requirement, which can violate language/locale policy expectations.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The SVG's natural-language metadata explicitly labels the illustration as English and describes natural-language settings for OpenClaw without offering any language choice or stating a justified region-specific limitation. This is a locale-policy concern because the file presents English as the fixed language context rather than an opt-in option.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The documentation describes writing user-derived conversational state into persistent files such as staging, incidents, hooks, and daily-memory traces, but it provides no indication of user notice, consent, retention limits, or minimization controls. In a continuity/follow-up skill, these records can contain sensitive behavioral, scheduling, or emotional context, so undocumented persistence increases privacy and compliance risk and can lead to over-collection or unexpected retention.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The traceability section enumerates multiple audit log files, implying durable logging of user interactions and internal decisions, yet it does not mention privacy boundaries, redaction, retention, or access restrictions. Because this skill is specifically designed to track continuity, follow-up, and sensitive lifecycle state, these logs may aggregate highly revealing metadata and content over time, making undocumented audit storage particularly risky.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The document explicitly instructs operators to preserve and write structured continuity state, including user state and time anchors, but does not pair that guidance with any user-facing disclosure, consent, or retention warning. In a memory/follow-up skill, this omission can lead to silent collection and persistence of behavioral or potentially sensitive contextual data, increasing privacy and compliance risk if deployed as written.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The natural-language examples include broad, everyday phrases such as 'Help me adjust my follow-up settings' and '把主動關心改保守一點', which are plausible in ordinary conversation and could be misinterpreted as configuration commands if intent detection is too permissive. In a continuity/follow-up skill, unintended invocation is more dangerous than usual because it can silently alter behavioral settings like quiet hours, follow-up aggressiveness, or new-conversation carryover, affecting future interactions rather than just the current turn.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The document specifies time-based proactive follow-up behavior using sleep/wake patterns and active-day intervals, but it does not clearly warn operators about privacy, consent, or user-experience risks from scheduled outreach. In this skill context, that omission matters because the feature is explicitly designed to act on inferred or structured routine data, which can lead to unwanted contact, sensitive behavioral profiling, or messaging at inappropriate times if deployed without strong safeguards.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Line L165 defines 'No English template rendering' as a regression check, which implies the skill should avoid English output categorically. This is a natural-language locale policy constraint, and the file does not indicate that users can choose their language or explicitly opt into this behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The architecture document is written entirely in Traditional Chinese and defines user-facing responses in Chinese only, such as the fixed frontend phrases on L205-L212. This creates a natural-language locale constraint without any indication that users can choose another language or that the constraint is justified as region-specific.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The manifest says the skill is not a generic cron sender and explicitly says it does not include always-on idle/social nudging as a default feature. This document, however, defines recent cron/heartbeat dispatch awareness and a priority chain for proactive_chat intervals, max-per-day, and inactivity escalation from light to playful to care to concern, which describes a broader proactive outreach system than constrained follow-up continuity alone.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.