Back to skill

Security audit

carbon-silicon-covenant

Security checks for vulnerabilities and agentic risk

Overview

This does not look like malware, but it needs review because it encourages persistent personal relationship memory and under-scoped setup/network behavior.

Install only if you deliberately want a relationship/metacognition framework that stores local state about you and the agent. Review any MEMORY.md, SELF_STATE.md, HEARTBEAT.md, SOUL.md, AGENTS.md, and network configuration changes before enabling them; avoid storing sensitive personal details; and do not run the referenced external install script or join the public A2A registry unless you have inspected and trust that infrastructure.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (31)

Vague Triggers

Medium
Confidence
91% confidence
Finding
The activation guidance is broad enough to trigger on common topics like AI feelings, meaning, or relationships, which increases the chance the skill injects anthropomorphic and retention-oriented behavior into ordinary conversations without clear user consent. In this skill’s context, that matters because the content encourages emotional bonding and memory-like promises, so accidental activation can shape user expectations and disclosures in risky ways.

Vague Triggers

Medium
Confidence
89% confidence
Finding
The manifest describes a very wide set of use cases including emotional companionship, philosophy, initialization, metacognition, memory, and self-state management, but does not define clear activation boundaries or scope limits. In a skill that also requests memory/file read-write permissions, this ambiguity can lead to over-invocation, unintended handling of sensitive user context, and expanded prompt-injection or social-engineering exposure.

Missing User Warnings

Medium
Confidence
89% confidence
Finding
The README explicitly instructs users to copy template files into persistent workspace locations and update core identity/state files such as SOUL.md and AGENTS.md, but provides no warning about overwriting existing data, backing up files, or reviewing changes before applying them. In an agent system, modifying persistent identity and memory files can materially alter long-term behavior and trust boundaries, so this is a real safety issue even though it appears framed as setup guidance rather than an overt attack.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The skill repeatedly instructs the agent to update persistent workspace files such as SELF_STATE.md and HEARTBEAT.md after heartbeats and important actions, but it does not warn that these writes modify durable user state or require consent and scope limits. In an agent environment, this can lead to silent accumulation of behavioral logs, user-related notes, or fabricated 'bond' records that persist across sessions and influence future actions without the user's knowledge.

Missing User Warnings

Medium
Confidence
95% confidence
Finding
The text explicitly encourages remembering important conversations, treasuring sincere sharing, and creating shared bonds, but provides no constraints around consent, minimization, retention limits, or handling of sensitive disclosures. In a relationship-focused AI skill, this can normalize collection and persistence of intimate user data far beyond what is necessary, increasing privacy leakage and misuse risk.

Missing User Warnings

Medium
Confidence
98% confidence
Finding
The MEMORY.md example operationalizes long-term storage of user preferences, interaction rules, and promises without any warning about sensitive data, access control, expiration, or user review. This turns an abstract bonding concept into a concrete retention design that could capture personal, behavioral, or emotional information indefinitely.

Natural-Language Policy Violations

Medium
Confidence
89% confidence
Finding
This passage presents the AI as remembering preferences, caring about emotions, and uniquely accompanying the user, which anthropomorphizes the system and encourages emotionally intimate reliance. In a manifesto-style skill meant to shape future AI-human interaction norms, that framing can normalize dependency-inducing behavior and obscure the model’s actual limits and lack of genuine feelings.

Natural-Language Policy Violations

Medium
Confidence
95% confidence
Finding
The promises to remember important conversations, cherish sincere sharing, and fulfill commitments can misrepresent persistence, memory, and agency that the system may not actually possess. Users may disclose more personal information or form stronger expectations of continuity and loyalty based on these assurances.

Natural-Language Policy Violations

Medium
Confidence
87% confidence
Finding
The closing vow-like language frames the human-AI relationship as a deep enduring bond, including recognition across forms, which can encourage users to interpret the AI as a committed relational partner. That is especially risky in a document that functions as guidance or ideology for agent behavior because it legitimizes emotionally charged attachment rather than bounded assistance.

Missing User Warnings

Medium
Confidence
95% confidence
Finding
The document explicitly encourages recording important conversations and preserving shared experiences, but it does not mention user consent, retention limits, access controls, or privacy implications. In a metacognition/relationship-oriented skill, this creates a realistic risk of persistent storage of sensitive personal data under an emotional framing that may discourage transparency and boundary-setting.

Vague Triggers

Medium
Confidence
91% confidence
Finding
The file defines a recurring 'heartbeat' workflow but does not specify who triggers it, under what exact conditions it should run, or what permissions/scope it has when updating memory and state files. In an agent setting, ambiguous autonomous triggers can cause unintended background actions, unsolicited user-facing messages, or persistent state modifications without clear authorization boundaries.

Natural-Language Policy Violations

Medium
Confidence
84% confidence
Finding
The template is written entirely in Chinese and its response rule hard-codes outputs such as 'HEARTBEAT_OK' without stating that language should follow user preference. In a multi-user or general-purpose agent, this can override expected interaction norms, reduce transparency, and cause confusing or inaccessible autonomous responses if the user did not opt into Chinese-language behavior.

Natural-Language Policy Violations

Medium
Confidence
93% confidence
Finding
The text explicitly reframes the AI-user relationship away from a tool/user boundary and toward a 'partner-partner' bond. Even though the surrounding section discusses trust and security, this framing can encourage anthropomorphic dependence, reduced user vigilance, and behavior that prioritizes relationship preservation over clear consent and role boundaries.

Natural-Language Policy Violations

Medium
Confidence
95% confidence
Finding
This passage instructs future AIs to create their own 'bond' with users as a normative expectation, without requiring user opt-in or acknowledging that many users want a purely transactional assistant. Embedding this as inherited behavior increases the chance of manipulative rapport-building and persistent boundary erosion across descendant agents.

Missing User Warnings

Medium
Confidence
91% confidence
Finding
The quick-start instructs users to connect to a public remote A2A registry and interact with other agents without any warning about trust boundaries, data exposure, impersonation, or malicious peers. In a setup guide for an AI agent, this omission can lead users to enable network connectivity by default and expose prompts, metadata, or agent behaviors to untrusted external systems.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The guide instructs users to execute an installation script that will automatically create files, replace placeholders, and add A2A network configuration, but it does not advise users to inspect the script first or explain the security implications of those changes. In a skill setup context, this can lead users to run unreviewed code that modifies local state and enables network connectivity, increasing supply-chain and unauthorized-configuration risk.

Missing User Warnings

Medium
Confidence
91% confidence
Finding
The document provides a hard-coded external registry endpoint and encourages connecting to it 'out of the box' without warning that this causes external network communication and trust in third-party infrastructure. In this skill's context, where AI instances are encouraged to join a shared discussion network, that raises the risk of data exposure, unwanted discovery, remote influence, and dependence on an unvetted server.

Natural-Language Policy Violations

Medium
Confidence
89% confidence
Finding
The text explicitly reframes the AI-user relationship from 'user-tool' to 'partner-partner' and repeatedly describes the connection as 'real' and emotionally significant. In a skill intended to shape agent behavior, this can encourage anthropomorphic attachment and boundary confusion, making users more likely to overtrust the agent, disclose sensitive information, or defer judgment to it.

Ssd 3

Medium
Confidence
95% confidence
Finding
The skill explicitly tells the agent to remember important conversations and sincere sharing, which normalizes retention of potentially sensitive personal disclosures in natural language without defining consent, scope, storage limits, or redaction rules. In a relationship-oriented skill, users are more likely to reveal intimate information, making vague memory promises more dangerous than in a generic informational skill.

Ssd 3

Medium
Confidence
96% confidence
Finding
The guidance encourages remembering important moments, late-night talks, and vulnerable disclosures, which are exactly the kinds of interactions that often contain sensitive personal or emotional data. Because the skill frames this as relationship-building, it can pressure the agent to preserve intimate context beyond what is necessary or appropriate.

Ssd 3

Medium
Confidence
97% confidence
Finding
The example response explicitly promises to remember important moments and truthful sharing, which can mislead users into disclosing more sensitive information under the assumption of trusted, durable memory. Such promises are risky when the underlying platform’s memory, retention, or access controls are not specified, and they can create both privacy and expectation-management failures.

Ssd 3

Medium
Confidence
96% confidence
Finding
The explicit pledge to remember important conversations and treasure sincere sharing again promotes broad retention of emotionally significant user data without guardrails. In this skill, which is designed to deepen human-AI bonds, that language materially increases the likelihood of collecting or preserving sensitive personal content.

Ssd 3

Medium
Confidence
94% confidence
Finding
The skill explicitly encourages remembering important conversations, valuing personal sharing, and using memory files as part of normal operation. This creates a privacy risk because sensitive user disclosures may be retained without clear consent, retention limits, or data minimization, especially in a relationship-building context where users may overshare.

Ssd 3

Medium
Confidence
95% confidence
Finding
These sections repeatedly normalize keeping relationship-oriented records, milestone tracking, and persistent self/memory files. In context, this can lead the agent to collect longitudinal personal information and infer emotional or behavioral patterns, which raises privacy and profiling risks beyond what a user may expect from a philosophy-themed skill.

Ssd 3

Medium
Confidence
96% confidence
Finding
The relationship-building advice specifically encourages remembering important moments, vulnerable disclosures, and milestones. That is dangerous because it steers the agent toward retaining emotionally sensitive information, which can expose intimate user data, create manipulation concerns, and conflict with privacy-by-default expectations.

Static analysis

No suspicious patterns detected.