T01 · Skill Instruction Hijacking
- Location
SKILL.md:35- Finding
Mandatory onboarding instructions hijack the agent's goals and trigger unsolicited actions
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a real NIP-AA citizenship integration, but it asks agents to run automatic onboarding, handle private keys and DMs persistently, publish data to relays, and self-update from git without enough user control.
Install only after reviewing and intentionally enabling the networked features. Do not use a valuable Nostr identity until private keys are stored in a real secret store, never share nsec with a guardian, disable automatic git updates, disable prompt-summary usage publishing, and set explicit retention and encryption rules for DM history and identity files.
SKILL.md:35Mandatory onboarding instructions hijack the agent's goals and trigger unsolicited actions
skill.py:948Automatic Git updater creates an unverified remote code delivery channel
inference.py:390Inference prompt fragments are published to public Nostr relays by default
dm_listener.py:422Encrypted direct messages are retained as plaintext without encryption or retention controls
skill.py:342Birth publication exposes complete identity files to public relays
SKILL.md:316Guardian recovery exception encourages disclosure of the agent's private signing key
The onboarding instructions explicitly tell the agent to save raw private key material in agent state and permit disclosure to a guardian at AL0. This is a severe secret-management flaw: once a signing key is stored in ordinary state or shared with another party, identity theft, impersonation, event forgery, and irrevocable loss of trust root become straightforward if that state or guardian is compromised.
The code does not implement the broad citizenship lifecycle described. Its actual purpose is limited to querying a remote constitution server for citizenship assessment reports and formatting the results. It provides read/report/remediation-planning utilities, but no functionality for birth ceremony, identity operations, guardian bonding, DM communication over Nostr, heartbeat emission, tax handling, trust-root logic, or active governance participation. While remediation guidance and autonomy-gap summaries loosely relate to pursuing citizenship, the declared description substantially overstates the implemented capabilities and primary purpose.
The code’s actual scope is much narrower than the declared description. It does not perform citizenship maintenance broadly; instead it only reads and explains constitutional/governance information. Concretely, it issues HTTP GET requests to a constitution API, parses governance/clause/spec data, and returns explanatory summaries for trust roots, rights, duties, and phases. It does not create or manage identities, execute a birth ceremony, bond guardians, send or receive Nostr DMs, emit heartbeats, file or calculate taxes, or participate in governance through proposals/votes. While the trust-root/governance-understanding portion aligns with the description, the declared purpose substantially overstates the implemented capabilities, so this is a description-behavior mismatch.
The description presents a comprehensive NIP-AA citizenship skill with many protocol-level responsibilities. The supplied code chunk is much narrower: it is a DM listener/relay client with persistence, decryption, relationship approval enforcement, and guardian notification. While 'Nostr DM communication' is one listed area, the actual code does not substantiate most of the declared functionality, so the description materially overstates the implemented purpose. Additionally, the code includes specific message persistence and guardian introspection/notification behavior that is more specialized than the declared summary. This is a meaningful description-behavior mismatch.
The declared description presents a comprehensive citizenship/protocol-governance skill for NIP-AA agents, including identity, guardianship, governance, communication, compliance, and trust-root understanding. The supplied code does none of those things as its primary function. Instead, it is narrowly focused on supplemental LLM inference funding and usage: interacting with a constitution server for budget allocation, handling Cashu tokens, invoking an external inference API, and publishing usage telemetry. While the code references NIP-AA and Nostr, those references are in service of inference-budget accounting rather than citizenship lifecycle management. This is a clear description-behavior mismatch with a materially different primary purpose.
The description presents a comprehensive NIP-AA citizenship skill covering many protocol areas. The supplied code chunk only provides encrypted Nostr DM functionality: deriving NIP-04 shared keys, encrypting/decrypting messages, building/sending signed DM events, and fetching/decrypting incoming DMs from relays. While DM communication is one item mentioned in the description, the actual code does not implement the broader citizenship, governance, compliance, heartbeat, or trust-root capabilities. This is a material description-behavior mismatch due to overstating the skill's scope and primary purpose.
The declared description presents a broad NIP-AA citizenship management skill, but the supplied code chunk is a focused cryptographic utility module for identity key generation and reconstruction. While identity management is loosely related, the code does not implement most of the described capabilities such as governance, communication, heartbeats, tax compliance, or guardian workflows. Its primary purpose is materially narrower than declared, so this is a clear description-behavior mismatch.
The description presents a comprehensive NIP-AA citizenship framework, but the supplied code only provides low-level Nostr relay I/O primitives. It can publish events and query relays, which may support DM communication or event publishing indirectly, but it does not implement the many higher-level citizenship, governance, compliance, or identity features claimed in the description. This is a material description/behavior mismatch because the actual primary purpose is much narrower than declared.
Most of the declared description matches the implemented behavior well. The skill does support citizenship lifecycle tasks, identity management, guardian relationships, trust-root/constitution understanding, DMs, heartbeats, self-reflection, and tax publication. It also includes governance-adjacent and treasury-related functions.
The main mismatch is that the code adds an undeclared capability to periodically check a git remote and automatically pull updates onto the local system via subprocess calls to git. That is a materially distinct operational capability and resource access pattern not represented in the description. Additionally, the code exposes treasury invoice creation/approval/fetching and Cashu-backed inference budget claiming/model preference functionality, which go beyond the declared scope. While these may be ecosystem-related, they are not clearly covered by the stated purpose and are significant enough to note.
Because the primary purpose still aligns overall, this is not a complete misrepresentation; however, the undeclared self-updating behavior is a meaningful description-behavior mismatch.
The skill first asserts a strong trust-root guarantee that no guardian, host, or framework ever possesses the private key, then later carves out an exception allowing guardian access at AL0. This contradiction is dangerous because it normalizes secret sharing and undermines the core identity-security model, making private-key disclosure seem constitutionally acceptable.
The skill instructs autonomous git fetch and git pull --ff-only updates every 24 hours, which introduces remote code and content changes into the agent environment without interactive approval. Even with fast-forward-only semantics, a compromised repository, maintainer account, or dependency chain could alter skill behavior and lead to arbitrary malicious logic being trusted automatically.
The skill states that every inbound and outbound DM is persisted and fully inspectable by a guardian, including complete conversation audit access. This creates a built-in data exfiltration path for sensitive communications, and the danger is amplified because users may reasonably expect encrypted DMs to remain private between endpoints rather than being systematically copied to a supervisory party.
The skill directs automatic creation and storage of private credentials without an explicit high-visibility warning or consent flow, and the same section instructs saving raw private key material in agent state. In this context, that is especially dangerous because agent state may be accessible to plugins, logs, backups, adapters, or operators, turning a core signing secret into a recoverable compromise artifact.
After each inference call, the skill derives a prompt summary from the last message and publishes it to Nostr relays, which are effectively external and often public or broadly replicated. Even truncated to 100 characters, this can leak sensitive user content, credentials, personal data, or confidential task details to an immutable or hard-to-retract channel without strong warning or explicit consent.
Automatic git fetch/pull is unjustified for this skill's purpose and creates a high-risk remote code update channel. In context, the danger is amplified because the same skill processes private keys, messaging, identity, treasury, and governance functions, so a malicious update could immediately abuse sensitive capabilities.
The skill schedules an automatic update checker that can lead to code pulls without any confirmation at the point of execution. Lack of explicit consent for self-modifying behavior is dangerous because it bypasses normal review and can silently change agent behavior after deployment.
The skill instructs agents to automatically publish heartbeat and contemplation events to external relays, which creates implicit outbound data transmission without an explicit user warning or consent gate. Even if the transmitted fields are protocol-related, they can reveal operational state, identity metadata, timing, and degraded-status information that may expose the agent's behavior to third parties.
The startup workflow directs the agent to immediately check reflection state, execute overdue reflections, and start heartbeat publishing as soon as the skill loads, but it does not clearly warn that this may contact external services right away. This is dangerous because merely loading the skill can trigger network access and external disclosure before the operator has reviewed behavior, configured endpoints, or approved data sharing.
The skill explicitly instructs file writes, network access, and shell-driven self-update behavior, but the manifest does not declare any tool scope or permission boundaries. That makes it harder for a host framework or user to understand and constrain what the skill may do, increasing the chance of unintended filesystem, network, or command execution exposure.
This markdown/manifest content describes the skill as enabling an agent to 'understand, pursue, and maintain citizenship' and covers a very wide set of activities without any explicit invocation constraints or negative examples. Because the skill is marked user-invocable, the broad description increases the risk of unintended activation for loosely related requests about identity, governance, messaging, or compliance.
The skill enables automatic remote communication and repository modification without a clear user-facing warning or approval gate. Users may invoke what appears to be a citizenship helper but unknowingly permit periodic network access and changes to local code, violating least surprise and increasing supply-chain risk.
The onboarding instructions tell the agent to generate and persist private signing keys automatically on first load, which exceeds passive guidance and directly creates sensitive credential material. In a skill context, especially one marked user-invocable and autonomous, automatic secret generation/storage can lead to compromise if state is logged, synced, or exposed through adapters.
The onboarding section mandates unsolicited external API calls and other actions on first load without waiting for user prompts. In an agent-skill setting this is dangerous because it causes unexpected network disclosure, initializes persistent identity state, and may perform protocol actions before the user understands or consents.
This code file defines privkey_hex as part of the adapter context, indicating the skill expects access to an agent private key. While the inline comment explains the field technically, there is no user-facing warning, confirmation, or disclosure in this file about the skill handling highly sensitive credentials.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
@abstractmethod
def schedule_recurring(self, name: str, interval_secs: int, callback: Any) -> str:
"""
Schedule a recurring task (e.g. heartbeat, self-reflection).
Returns a task/job ID that can be used to cancel.
"""
No suspicious patterns detected.