Back to skill

Security audit

Test Continuity

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a coherent conversation-continuity tool, but it includes under-disclosed external embedding and broad file-moving/trashing utilities that should be reviewed before installation.

Install only if you are comfortable with a stateful skill that records selected conversation context, writes daily memory traces, and may create follow-up hooks. Before enabling it, disable or tightly restrict remote embedding fallback, remove or sandbox the file_output_sop.py helper, pin dependencies, and make sure users can inspect, edit, and delete stored continuity data.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/personal_hooks.py:7705
Finding

Automatic Transmission of Raw User Dialogue to an Unrestricted Embedding Endpoint

Content
View full analysis
Optional[dict]: if not OPENCLAW_CONFIG_PATH.exists(): return None try: config = json.loads(OPENCLAW_CONFIG_PATH.read_text(encoding="utf-8")) except Exception: return None embedding = ( config.get("plugins", {}) .get("entries", {}) .get("memory-lancedb-pro", {}) .get("config", {}) .get("embedding", {}) ) if not isinstance(embedding, dict): return None api_key = embedding.get("apiKey") if isinstance(api_key, list): api_key = next((str(item).strip() for item in api_key if str(item).strip()), "") api_key = resolve_env_vars(str(api_key or "")).strip() model = str(embedding.get("model") or "").strip() base_url = resolve_env_vars(str(embedding.get("baseURL") or "")).strip() if not api_key or not model or not base_url: return None return { "api_key": api_key, "model": model, "base_url": base_url.rstrip("/"), "dimensions": embedding.get("dimensions"), "task_query": embedding.get("taskQuery"), "task_passage": embedding.get("taskPassage"), "normalized": embedding.get("normalized"), } ``` ```python def jina_embed_inputs(texts: list[str], config: dict, task: Optional[str]) -> list[list[float]]: if not texts: return [] mode = "query" if task and task == config.get("task_query") else "passage" payload = { "texts": texts, "mode": mode, "plugin_root": EMBEDDER_ROOT or None, "config": { "api_key": config["api_key"], "model": config["model"], "base_url": config["base_url"], "dimensions": config.get("dimensions"), ...[truncated 4603 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/file_output_sop.py:40
Finding

Bundled Utility Can Move and Trash Arbitrary Accessible Files

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned and Unverified Python Runtime Dependency

Content
View full analysis
=1.8.0 ``` The installation guide instructs users to resolve and install the dependency directly: ```bash python3 -m pip install -r requirements.txt ``` ### Technical Analysis The lower-bound-only version constraint allows pip to install any future `send2trash` release compatible with `>=1.8.0`. The project does not provide a lock file, an exact version, or package hashes. As a result, the code reviewed during this audit is not sufficient to determine the code that future installations will obtain. Resolution may vary by installation date, selected package index, dependency mirror, and resolver state. The dependency is used by the unrelated destructive file-output utility rather than the primary continuity engine, increasing the exposed supply-chain surface without a clear requirement from the Skill's documented purpose. ### Attack Path 1. A user follows the documented installation command. 2. Pip contacts the configured package index or mirror. 3. The resolver selects the newest package version satisfying `send2trash>=1.8.0`. 4. If a future release, package index, or configured mirror is compromised, the installer downloads unreviewed package content. 5. Package installation or later import executes with the privileges of the installing user or runtime account. 6. Malicious dependency behavior can access data and resources available to that account. This finding does not establish that the current `send2trash` package is malicious. The vulnerability is the absence of reproducible version and integrity controls. ### Impact Assessment Successful supply-chain compromise could execute code with the privileges used for installation or runtime. Depending on deployment practice, this ...[truncated 422 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (50)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill text repeatedly claims it is not a generic cron sender and does not provide always-on nudging, yet the documented settings include proactive care, intervals, re-engagement retries, heartbeat, and timed follow-up controls that amount to outbound scheduled behavior. That mismatch can mislead operators about how much autonomous messaging or behavioral policy the skill can perform, causing it to be deployed in contexts with insufficient safeguards.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill text repeatedly claims it is not a generic cron sender and does not provide always-on nudging, yet the documented settings include proactive care, intervals, re-engagement retries, heartbeat, and timed follow-up controls that amount to outbound scheduled behavior. That mismatch can mislead operators about how much autonomous messaging or behavioral policy the skill can perform, causing it to be deployed in contexts with insufficient safeguards.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill text repeatedly claims it is not a generic cron sender and does not provide always-on nudging, yet the documented settings include proactive care, intervals, re-engagement retries, heartbeat, and timed follow-up controls that amount to outbound scheduled behavior. That mismatch can mislead operators about how much autonomous messaging or behavioral policy the skill can perform, causing it to be deployed in contexts with insufficient safeguards.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill text repeatedly claims it is not a generic cron sender and does not provide always-on nudging, yet the documented settings include proactive care, intervals, re-engagement retries, heartbeat, and timed follow-up controls that amount to outbound scheduled behavior. That mismatch can mislead operators about how much autonomous messaging or behavioral policy the skill can perform, causing it to be deployed in contexts with insufficient safeguards.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 114)May include surrounding context.

md
- Script: `scripts/personal_hooks.py`

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 176)May include surrounding context.

md
- Script: `scripts/personal_hooks.py`

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 180)May include surrounding context.

md
- Script: `scripts/personal_hooks.py`

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest frames this skill as continuity-driven follow-up with explicit lifecycle rules and explicitly says it is not a generic cron sender. However, the code adds a standalone recurring task subsystem with daily/weekly/every-N-hours frequencies, overdue checks, completion handling, and reminder hook creation, which is a generic scheduler capability independent of conversational continuity.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The frontstage normalization layer rewrites Simplified Chinese phrases and characters into Traditional Chinese unconditionally. This enforces a specific locale/script policy on user-facing output rather than offering a user choice or documenting a justified locale constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README advertises daily-memory writeback from structured continuity state but does not clearly warn users that conversation-derived information may be persisted beyond the current chat. This can undermine user expectations, especially for sensitive events, delegated tasks, and watchful states, leading to privacy harm, over-retention of personal data, and accidental storage of content users thought was ephemeral. The context increases risk because the skill's core purpose is to classify, retain, and reconnect personal conversational context over time.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The README encourages natural-language settings changes using broad phrases such as adjusting follow-up behavior in ordinary conversational wording. In a continuity/follow-up skill, that creates a real risk that normal user dialogue is misinterpreted as a persistent configuration change, causing silent behavioral drift, unwanted follow-ups, or suppression of safety-relevant behavior. The context makes this more dangerous because the skill is explicitly designed to persist state and alter future agent conduct from conversational inputs.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill declares required environment variables and instructs use of shell-driven Python entry points that can read and modify local state, but it does not declare any explicit tool scope or permission boundary. That creates an authorization gap: hosts or reviewers cannot easily constrain filesystem, environment, and shell access to the minimum needed, increasing risk of overbroad execution in a stateful skill.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The setup instruction says to delete SKILL_ONBOARD.md once complete, but it provides no warning, confirmation, or qualification about preserving auditability or avoiding accidental deletion of an unexpected file. Destructive file actions without confirmation are risky in agent-driven environments because they can remove user or operator data based on ambiguous workspace state.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill explicitly states that deterministic onboarding writes user-provided facts into settings.json, profile.json, and USER.md, but it does not pair this with a clear user-facing consent and persistence warning. Silent or insufficiently disclosed persistent writes are dangerous because they can store sensitive relationship, schedule, timezone, and usage-profile data that later influences agent behavior across sessions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The checklist item 'No English template rendering (locale resolution correct)' suggests the skill treats English output as an error state, which can indicate a forced language/locale policy. Under the policy rules, natural-language constraints that enforce a specific locale without user choice should be flagged.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The natural-language triggers are broad enough that ordinary conversational phrases can be interpreted as durable events, including sensitive ones. In this skill's context, that is more dangerous than in generic chat because detections feed persistent event storage, future follow-up behavior, and runtime context injection, amplifying the effect of a mistaken parse.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The architecture permits silent automatic creation of events for sensitive categories like health and emotional state, yet does not require user-facing notice, consent, or a minimization policy. In a continuity/memory skill, this creates a real privacy risk because intimate personal data may be persisted and later surfaced or acted on without the user clearly understanding that storage occurred.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The design documents a status=suggested flow, but the declared schema only allows active | parked | due | completed | cancelled | superseded. This mismatch can cause suggestion-state events to be mis-stored, dropped, or coerced into another status, breaking the confirmation gate and potentially turning tentative tracking into active tracking without explicit user confirmation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The CRUD phrases are ambiguous and the documented matching rule relies on simple active-event selection or keyword overlap, which can cause the wrong event to be completed, cancelled, or parked. Because this skill manages longitudinal follow-up state, a misfire can suppress important reminders or incorrectly preserve sensitive tracking the user intended to stop.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

This markdown states that "language neutrality at the public-package level" still needs explicit release proof, indicating the skill may be published or considered ready without confirmed compliance with language/locale neutrality expectations. Because language/locale policy applies across all file types, this is a natural-language policy concern in the documented release status.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Line L07 states that frontstage text should be kept in natural Traditional Chinese as a blanket rule. This is a natural-language locale constraint and the file does not mention any user preference, opt-in, or justified region-specific limitation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The examples, prompts, signal phrases, and user-facing interaction strings are predominantly fixed in Chinese, and confirmation logic is built around Chinese phrases such as "好" and "不用". There is no visible opt-in or language selection mechanism, so the skill effectively enforces a specific language/locale policy through its natural-language behavior.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The runtime context and carryover builders explicitly instruct the model to reuse tracked user conversation details across turns and sessions, including causes and follow-up context. In a skill centered on continuity this is expected behavior, but it still creates a privacy risk because sensitive personal details may be surfaced later in contexts where the user did not expect them or where other subsystems can access the prompts.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The context prompt instructs that suggested events need explicit user confirmation and must not be auto-confirmed. However, process_event_entry confirms the most recent suggested event whenever the next user message merely matches broad tokens like 好, 對, 嗯, or ok and is under 20 characters, without checking that the reply is actually about the suggestion.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
85% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · scripts/event_store.py (reported line 419)May include surrounding context.

python
lines.append("")
            lines.append("### Pending event suggestions")
            lines.append("These events were detected but need user confirmation. Ask briefly: 「這件事要我幫你記著追嗎?」")
            lines.append("If user says yes/好/對, confirm it. If no/不用/算了, drop it. Do NOT auto-confirm.")
            for evt in suggested:
                title = evt.get("title", "untitled")
                cause = evt.get("cause_summary", "")

Static analysis

No suspicious patterns detected.