T01 · Skill Instruction Hijacking
- Location
skill.md:78- Finding
Automatic Persistent Modification of Workspace Agent Instructions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill is openly designed to make an agent proactive, but it grants broad private-memory access, persistent workspace changes, and automatic outbound messaging with insufficient safeguards.
Install only if you deliberately want an agent that can change workspace instructions, retain personal intellectual state, search broad private content, and send autonomous messages. Before enabling, require a reviewed AGENTS.md diff, limit QMD to approved paths, exclude sensitive notes and secrets, confirm the messaging destination, and define retention/deletion rules for CURIOSITY.md and daily memory.
skill.md:78Automatic Persistent Modification of Workspace Agent Instructions
skill.md:110Unbounded Persistent Retention of Conversation-Derived Agent State
references/proactive-protocol.md:67Broad Private-Data Indexing Feeds Autonomous Outbound Messages
README.md:74Unpinned Global Installation of a Third-Party npm Package
The skill directs the agent to read workspace, vault, and memory content before each unprompted dispatch, giving a scheduled background process broad access to potentially sensitive local data. In context, this is more dangerous because the collected material is then used to generate outbound messages, increasing the chance of confidential data exposure, overcollection, and unintended exfiltration through the configured messaging channel.
The proposal states that the skill will automatically patch AGENTS.md and create CURIOSITY.md on first load, but it does not clearly present this as a user-consent-sensitive behavior or describe safeguards such as confirmation, diff preview, or rollback. Automatic persistent modification of workspace control files can change agent behavior across future sessions in ways the user may not expect, making this a real security and integrity concern even if the stated purpose is functional rather than malicious.
The README explicitly states that on first load the skill will modify AGENTS.md and create CURIOSITY.md automatically, but it does not require an explicit confirmation step immediately before those workspace changes occur. Automatic mutation of control files is risky because it can silently alter agent behavior and persistence settings in ways the user may not notice or fully review first.
The morning cron payload instructs the agent to read Knowledge_Core.md, knowledge-personal.md, and CURIOSITY.md, then send a proactive message over an external messaging channel without any boundary on what personal or sensitive material may be included. This creates a clear risk of privacy leakage, because internal memory and user-specific knowledge are being used as source material for unsolicited outbound content.
The midday dispatch instruction again directs the agent to read knowledge-personal.md and CURIOSITY.md and then transmit a message externally, with no explicit separation between internal personalized context and shareable output. Because the skill is designed to produce opinionated, context-rich messages, the surrounding skill context makes inadvertent disclosure more likely, not less.
The afternoon workflow combines reading internal knowledge and active threads with unsolicited delivery to a configured external channel, again without content-classification or minimization controls. This is dangerous because it operationalizes a continuous exfiltration path from workspace memory into third-party platforms such as Slack, Discord, or Telegram.
The evening prompt instructs the agent to use user-personal context and active threads as source material for unsolicited external messaging, with no guardrails limiting sensitive inference or disclosure. In a skill whose purpose is to simulate continuity and inner life, this context increases the chance that private user beliefs, projects, or notes will be echoed outward in ways the user did not specifically authorize.
The protocol explicitly encourages the agent to quote and engage with the user's own stored words in proactive dispatches. Even if intended as personalization, reusing personal text in autonomous outbound messages increases the chance that sensitive reflections, confidential notes, or contextually private statements are repeated in the wrong channel or at the wrong time.
These instructions direct the agent to search and reuse content from the full workspace, Obsidian vault, session history, drafts, research, and prior analyses when composing autonomous outbound messages. That creates a direct natural-language exfiltration path: sensitive or confidential text can be surfaced in unsolicited communications without a strict need-to-know boundary, classification filter, or source restriction.
The protocol says outbound dispatches are sent via a messaging channel and persisted into daily memory files, but it does not require a prominent user-facing disclosure at the moment of enablement and before each data flow about what content may be included, where it is stored, and how long it is retained. Because the skill also draws from broad workspace and personal knowledge sources, users may underestimate that autonomous messages and retained copies can expose sensitive material or create a durable record in indexed storage.
The skill instructs the agent to load a user's personal knowledge file and use that material to shape every proactive dispatch, which operationalizes persistent reuse of user-authored thoughts across sessions. Even if intended as personalization, this expands the surface for privacy leakage because sensitive user ideas may be echoed, transformed, or surfaced in later unsolicited messages without contextual review.
The skill explicitly schedules four unprompted outbound messages per day and frames them as automatic behavior, which creates a consent and messaging-abuse risk even though the metadata warns about it. Because the operative instructions tell the agent to self-activate and send messages without a per-dispatch user prompt, this can produce spammy behavior, unwanted notifications, or policy violations in environments where explicit opt-in is required.
The salience-tagging and memory instructions direct the agent to record and preserve high-signal relational and unresolved conversation details in daily memory files for future reuse. This creates a cross-session profiling and retention risk because subjective, potentially sensitive interaction details are explicitly marked as important and made easier to retrieve and reuse later.
The cron setup examples explicitly use "tz": "America/New_York", which imposes a specific locale in the documented default behavior. Although nearby text says to adjust timezone to match the user, the provided configuration still defaults to a specific region rather than prompting for user choice first.
No suspicious patterns detected.