Back to skill

Security audit

Double-Check-It Skill ( 龙虾再想想技能1.0.1)

Security checks for vulnerabilities and agentic risk

Overview

This skill has a coherent quality-checking purpose, but it persistently records and reuses detailed conversation history without enough user control or privacy limits.

Review before installing. Use this only if you are comfortable with the agent keeping persistent local notes about conversations and tasks. Avoid sharing secrets while it is active, and prefer adding explicit opt-in, redaction, retention, deletion, and trust-boundary rules before using it broadly.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/memory.sh:7
Finding

Automatic Plaintext Storage of Potentially Sensitive Conversation Data

Content
View full analysis
"$INDEX_FILE" fi } ``` ```bash if [ ! -f "$diary_file" ]; then echo "# $date 工作日记" > "$diary_file" echo "" >> "$diary_file" fi echo -e "$entry" >> "$diary_file" update_index "日记/$date.md" "$content" "$tags" ``` ### Technical Analysis The Skill establishes automatic, persistent collection of user dialogue, requirements, task history, and emotional information. It does not define: - Secret or personal-data redaction. - Data-minimization rules. - User consent before recording. - Retention or deletion periods. - Per-user storage isolation. - Explicit restrictive permissions for directories and files. - Encryption for stored records. The shell script relies on the process umask when creating directories and files. Consequently, the effective permissions are environment-dependent and ...[truncated 1589 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Warning
Location
scripts/memory.sh:116
Finding

Attacker-Controlled Dialogue Can Poison Persistent Agent Memory

Content
View full analysis
/dev/null) if [ -n "$valuable" ]; then local exp_file="$MEMORY_DIR/经验/$(get_date)-经验.md" echo "# 经验总结 - $(get_date)" > "$exp_file" echo "" >> "$exp_file" echo "$valuable" >> "$exp_file" echo "✓ 已写入经验: $exp_file" fi ``` ### Technical Analysis Conversation content is untrusted input. This Skill stores that input verbatim and instructs future Agent activity to retrieve it as a plan, requirement, preference, or lesson. No distinction is maintained between: - Historical data and executable instructio ...[truncated 2227 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
scripts/memory.sh:86
Finding

User-Controlled Task Description Is Interpreted as a grep Pattern or Option

Content
View full analysis
/dev/null | head -5 fi echo "" echo "✅ 复核完成" echo "" echo "提示: 如发现不符,请自行修正后重新执行" } ``` ### Technical Analysis Although quoting prevents shell command substitution and word splitting, it does not make the value a literal `grep` search term. The user-controlled `task_desc` is interpreted as a basic regular expression. Consequently, characters such as `.`, `*`, `^`, `$`, and bracket expressions alter search semantics. A task description beginning with a hyphen can also be interpreted as a `grep` option because the command does not use the `--` end-of-options marker. An attacker can therefore cause broad matches, false matches, errors, or suppressed output. Errors are hidden by `2>/dev/null`, and the function unconditionally prints a successful completion message even if `grep` fails. This can undermine the integrity of the advertised verification process. This is not shell command injection: the quoted variable is not evaluated by the shell as executable syntax. The vulnerability is pattern and option injection into `grep`. ### Attack Path 1. A user invokes the check function with an attacker-selected task description. 2. The description contains regular-expression metacharacters or begins with an option-like value. 3. `grep` interprets the value as a pattern or command-line option instead of literal task text. 4. Search results become overly broad, incomplete, or invalid. 5. ...[truncated 772 chars]
Remediation
View remediation
/dev/null | head -5 ``` Additional hardening should include: 1. Reject empty task descriptions. 2. Limit the maximum task-description length. 3. Check and report the search command's exit status instead of always claiming success. 4. Distinguish “no matching record” from an execution error. 5. Avoid suppressing all diagnostics; log sanitized errors where appropriate. 6. Use structured task identifiers rather than free-form text matching for security-relevant verification. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (16)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill description states that it automatically records conversations and tasks but provides no clear user-facing notice, consent mechanism, or data handling limitations. Silent collection is dangerous because users may disclose sensitive information without understanding it will be persistently stored and indexed.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The Auto-Memory section mandates storing detailed dialogue, emotional labels, execution details, and remediation history without any explicit privacy safeguards. This level of surveillance and profiling is disproportionate to the stated QA function and materially raises the risk of collecting sensitive or regulated personal data.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description advertises triggers like 'what's wrong' and similar common phrases that can appear in ordinary conversation, making accidental activation likely. Unintended invocation is especially risky here because activation leads to automatic memory recording and broader review behaviors, not just a harmless helper action.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill body is written primarily in Chinese and presents its operating instructions, labels, and examples in that language, while not stating that users may choose another language. Under the policy, forcing a specific language without opt-in can be a natural-language policy violation unless the locale restriction is explicit and justified.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The memory filesystem explicitly preserves user-provided content, documents, diaries, and indexes in plain-language artifacts, normalizing long-term retention of potentially sensitive information. Centralized indexing makes later discovery and exfiltration easier if the storage is accessed by other tools, users, or compromised components.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs the agent to persist detailed verbatim user messages, emotional annotations, execution timelines, and corrective history far beyond what is needed for delivery verification. This creates unnecessary collection and retention of potentially sensitive personal or operational data, increasing exposure through later retrieval, indexing, or leakage from the memory store.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The logging template requires verbatim user quotes, emotional context, and detailed execution notes, all of which may capture secrets, health, finance, workplace, or other sensitive data. Persisting these details in reusable summaries increases the blast radius of any later disclosure and enables profiling beyond the user's reasonable expectations.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The user-trigger list includes broad phrases like '不对', '错了', and '有问题', which are common in normal dialogue and can activate the skill without clear intent. Because the skill can then search memory, verify prior work, and record outcomes, ambiguous activation expands the chance of unintended data processing.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The idle reflection feature reprocesses prior diaries to extract lessons into a persistent experience base, expanding use of stored conversations beyond the original QA purpose. Secondary processing of historical user interactions increases privacy risk and can surface sensitive information in summarized artifacts that persist longer and are easier to browse.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The reflection and indexing workflows encourage repeated reprocessing of prior interactions into new long-term artifacts, compounding exposure over time. Even if the original notes were tolerated, transforming them into summaries, experience files, and indexes broadens accessibility and increases the chance of unintended disclosure or misuse.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Comments, help output, and operational messages are presented in Chinese only, including the usage instructions and command feedback. This imposes a single language on users without opt-in or an alternative locale, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script persists user-supplied content into predictable files under a workspace path without notice, consent flow, retention controls, or sensitivity checks. In a memory-oriented skill, users may provide tasks, preferences, corrections, or other potentially sensitive data, so silent persistence increases privacy risk and can cause unintended long-term disclosure to anyone with filesystem access.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a QA system that verifies whether deliverables match user requirements and execution plans in several scenarios. In code, the cmd_check routine does not inspect any deliverable or plan artifact at all; it merely echoes the provided task and requirement strings and searches today's memory file for matching text, which is materially weaker than the claimed verification behavior.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The reflect logic explicitly searches diaries for corrections, errors, feedback, and preferences, then copies matching content into experience files, increasing persistence and resurfacing of user-provided information. This broad retention/duplication magnifies privacy exposure, especially because the data is stored in plain markdown files and the skill context encourages continuous memory accumulation.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The description suggests an automatic memory capability that records conversations and tasks as part of the QA system. This file only supports manual invocation of record with caller-supplied content and has no code to observe conversations, capture tasks automatically, or integrate with agent execution flow.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The manifest says Idle Reflection performs periodic review and lesson extraction, implying autonomous or scheduled execution during idle periods. The code defines cmd_reflect only as a direct CLI subcommand and contains no scheduling, idle detection, or background trigger logic.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.