Back to skill

Security audit

agent-chronicle

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent diary skill, but it stores and reuses sensitive session-derived memory in ways that need careful review before installation.

Install only if you are comfortable with a diary skill reading local memory logs and saving long-lived reflections, quotes, decisions, mood inferences, and relationship notes. Keep it in a low-sensitivity workspace, review generated entries before saving, avoid blanket automation approvals, disable memory integration and relationship tracking unless explicitly wanted, and avoid exporting untrusted diary markdown to PDF/HTML until renderer resource fetching is constrained.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
scripts/generate.py:72
Finding

Untrusted Memory Content Can Hijack Sub-Agent Instructions and Poison Persistent State

Content
View full analysis
15000: content = content[:15000] + "\n\n[... truncated for context ...]" return content return None ``` ```python if today_log: context_parts.append(f"## Today's Session Log ({date_str}):\n{today_log}") if recent_sessions: context_parts.append(f"## Recent Session Context:\n{recent_sessions}") if persistent_files.get("quotes"): context_parts.append( f"## Quote Hall of Fame (existing):\n{persistent_files['quotes']}" ) if persistent_files.get("curiosity"): context_parts.append( f"## Curiosity Backlog (existing):\n{persistent_files['curiosity']}" ) if persistent_files.get("decisions"): context_parts.append( f"## Decision Log (existing):\n{persistent_files['decisions']}" ) if persistent_files.get("relationship"): context_parts.append( f"## Relationship Notes (existing):\n{persistent_files['relationship']}" ) context = "\n\n---\n\n".join(context_parts) task = build_generation_task(date_str=date_str, context=context) ``` ```python user_prompt = f"""Write your personal diary entry for {date_str}. Based on the following context from today and recent days: {context} --- Write a RICH, reflective diary entry (400-600 words minimum) with these sections: ... """ ``` ```python def update_persistent_files(entry_content, date_str, workspace): ...[truncated 3588 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/export_pdf.py:774
Finding

PDF Rendering Permits Unrestricted Local and Remote Resource Resolution

Content
View full analysis
{date_str} {escape(title_clean)} ''') # Convert markdown to HTML html_body = markdown.markdown( content, extensions=["fenced_code", "tables", "sane_lists", "smarty"] ) ``` ```python def export_pdf(output_path: Path, month: str = None): """Export diary entries to a beautiful PDF, optionally filtered by month (YYYY-MM)""" config = load_config() diary_path = get_diary_path(config) entries = load_entries(diary_path, month=month) if not entries: if month: print(f"No diary entries found for {month} in {diary_path}") else: print(f"No diary entries found in {diary_path}") return False html = build_html(entries) if not html: print("Failed to build HTML") return False output_path.parent.mkdir(parents=True, exist_ok=True) HTML(string=html, base_url=str(diary_path)).write_pdf(str(output_path)) ``` ### Technical Analysis Diary Markdown is converted to HTML and inserted into the final document without sanitizing resource-bearing elements such as images. WeasyPrint then renders th ...[truncated 2054 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/generate.py:378
Finding

Unvalidated Date and Diary Path Values Permit Writes Outside the Intended Workspace

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (40)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

There is a strong description-behavior mismatch. The description frames the skill as a generative journaling tool, likely centered on composing 400–600 word reflective diary entries with several structured sections and automation features. In contrast, this code does not generate text entries at all: it loads markdown files from a diary directory, performs regex/keyword-based sentiment and topic analysis, extracts wins and frustrations from existing sections, and emits a report. While 'mood analytics' mentioned in the description overlaps with part of the behavior, that is only one subset of the declared functionality, and the code's primary purpose here is analytics rather than diary generation. Additionally, the code reads workspace diary files even though no permissions are declared.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description emphasizes generation and analysis features for diary content, including rich reflective writing and multiple higher-level diary intelligence features. The actual code does not generate diary entries, perform analytics, schedule anything, or implement the named features. Instead, it exports existing markdown diary files to PDF or HTML and lists/filter entries from the filesystem. This is a materially different primary purpose and introduces undeclared capabilities related to file export and subprocess execution.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a content-generation skill centered on AI-written diary entries and advanced journal-analysis features. The supplied code does something materially different: it exports existing markdown diary entries into a styled PDF document. It contains no AI calls, no diary text generation, no analytics, no resurfacing logic, and no scheduling. Its actual primary purpose is document formatting/export, not diary generation. Additionally, it reads from the workspace diary directory and writes PDF/HTML files, which is inconsistent with the absence of declared permissions. This is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The code is broadly related to diary generation, so the overall theme matches. However, the description overstates direct AI generation and lists features not present in this chunk, while the code also performs additional persistence/integration behavior not disclosed. Specifically, this script does not call a model directly; it prepares a task for another agent, or accepts manual/pre-generated content. It also writes to daily memory files, maintains several persistent markdown knowledge files, and can invoke PDF export. Conversely, weekly digests and explicit mood analytics are not implemented here, though 'On This Day' resurfacing is partially implemented as a Looking Back section and scheduled auto-generation is supported via CLI auto mode. These differences are material enough to count as a description/behavior mismatch.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · templates/decisions.md (reported line 11)May include surrounding context.

md
*No decisions logged yet. They'll appear here from daily entries.*

<!--
Format for entries:

### [Decision Title]

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · templates/relationship.md (reported line 11)May include surrounding context.

md
*Notes about how we work together*

<!--
Examples:
- Prefers concise responses during work hours
- Likes detailed explanations when learning something new

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README promotes automatic journaling from session logs and memory writes, but it does not prominently warn that potentially sensitive session content may be copied into persistent diary and memory files. This can lead to unintended retention of private prompts, secrets, user data, or internal reasoning-adjacent material, especially when auto-generation is enabled.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · README.md (reported line 60)May include surrounding context.

python3 scripts/export.py --format pdf --days 7

text

> **Note:** If no `config.json` exists, `generate.py` automatically runs the setup wizard on first use.

## Scheduled generation in OpenClaw 2.0

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
81% confidence
Finding

The README presents commands like @diary quotes, @diary curious, and @diary relationship as activation patterns, but it does not clearly state where these commands are valid, how the skill is invoked, or any negative examples showing when they should not trigger. In markdown guidance, this lack of specificity can lead to unintended invocation in contexts where @diary could be interpreted broadly.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill clearly describes capabilities that read environment/workspace state, invoke Python scripts, and write persistent files, but it does not declare any explicit tool scope or permissions. That mismatch weakens review and containment because operators cannot easily tell, up front, that the skill needs shell, file read, and file write access.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The top-level description emphasizes journaling features but does not clearly warn that the skill persistently stores conversation-derived reflections, quotes, relationship notes, and summaries for future reuse. Users and operators may therefore invoke it without informed consent about long-term retention and resurfacing of sensitive content.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill encourages durable capture of subjective interactions, memorable statements, frustrations, and patterns over time. That creates privacy and profiling risk because sensitive user-derived material can be stored, correlated across sessions, and later surfaced in new outputs without renewed consent.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The skill frames journaling around 'cherish memorable moments with our human' and documenting how the dynamic grows, encouraging cumulative trust narratives and personal memory capture. In context, that makes the storage features more dangerous because it normalizes retention of interpersonal details not strictly needed for diary generation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger phrases are generic terms like 'journal', 'quotes', and 'write entry', which can cause accidental activation during unrelated conversations. In this skill, unintended activation matters because it can lead to persistence of conversation-derived content into memory files and downstream automation flows.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The template directly instructs the agent to document human interactions, memorable quotes, and evolving relationship details across sessions. This promotes creation of a persistent personal dossier that may include sensitive preferences, emotional context, and identifying anecdotes beyond what is necessary for the stated function.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The 'Quote Hall of Fame' explicitly tells the system to persist user statements for future retrieval and reuse. Stored quotations can contain sensitive, identifying, or context-dependent content that may later be exposed, misinterpreted, or replayed outside the original conversation context.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Relationship tracking instructions direct the agent to retain preferences, inside jokes, recurring themes, and collaboration patterns about the human. This is long-term behavioral profiling, which raises privacy and trust concerns and can produce unnecessarily intimate or revealing memory artifacts.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Memory integration and 'On This Day' resurfacing intentionally replicate diary content into broader memory logs and future entries. Replication increases the blast radius of sensitive content, makes deletion harder, and can reintroduce old private material into new contexts without fresh user approval.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
84% confidence
Finding

The automation guidance instructs users to choose 'Always allow' for exact commands used by the recurring job. Persistent approval for scheduled execution increases the chance that future unattended runs will read session data and write reflective outputs without timely human review, especially if the surrounding workspace content becomes more sensitive over time.

Content

Scanner excerpt · SKILL.md (reported line 568)May include surrounding context.

md
Run the job once with `openclaw automations run <job-id> --wait`. When an approval card appears, choose **Always allow** for each exact command the job runs. The grant also binds the command to the same working directory, environment, and automation configuration. Keep those values unchanged on later runs.

If no approval surface is connected, the `--auto` command is denied immediately. The run records an error, and repeated failures can disable the recurring job. Inspect disabled jobs with `openclaw automations list --all` and check the run history before enabling it again.

See the [OpenClaw automations documentation](https://docs.openclaw.ai/automation/cron-jobs) for other schedule and delivery options.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Weekly digest generation aggregates prior quotes, decisions, mood trends, and observations into synthesized summaries. Aggregation increases sensitivity because it can reveal higher-level patterns and inferences that are more privacy-invasive than any single stored note.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Automatically running setup and enabling memory-related behavior can cause persistent storage and configuration changes without a deliberate privacy review at the moment of first use. In this skill's context, even benign automation is riskier because it governs retention of sensitive journal and interaction data.

Content

Scanner excerpt · SKILL.md (reported line 984)May include surrounding context.

md
- **Context Awareness:** Reads recent session logs and existing memory files for context

### v0.3.0
- **Auto-Setup:** `generate.py` now automatically runs setup wizard if no config.json exists
- **Memory Integration:** New feature to append diary summaries to main daily memory log (`memory/YYYY-MM-DD.md`)
  - Three formats: `summary`, `link`, `full`
  - Enabled by default during setup

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The configuration enables memory integration that appends diary-related content into daily records, which can aggregate sensitive personal reflections and behavioral data without any accompanying consent, notice, or privacy safeguards in the example configuration. In a diary skill, this materially increases the chance that intimate or identifying information is collected and persisted in ways users may not fully anticipate.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Mood tracking and topic extraction are profiling features that infer emotional state and patterns from diary entries, which are highly sensitive by nature. Describing and enabling these analysis capabilities without clear warnings, opt-in consent, or data-handling safeguards creates privacy risk because users may not realize the extent of behavioral inference being performed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This code writes a synthesized report containing sensitive diary-derived mood, wins, and frustrations data to a user-specified file path. While it confirms after saving, there is no prior user-facing warning in the code that the output may contain sensitive personal information, and the save-to-file behavior is optional rather than inherently obvious from the skill purpose alone.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest describes an AI diary generation skill with reflective journaling, analytics, resurfacing, and scheduled generation, but does not mention document conversion tooling or host-level command execution. This file invokes pandoc through subprocess to generate PDF/HTML exports, adding a capability beyond the stated purpose and one that is more sensitive than normal text generation behavior.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.