Back to skill

Security audit

S2-DaoKernel: Agent Safety & Ethics Rules for SOUL.md

Security checks across malware telemetry and agentic risk

Overview

This is not a malware-style package, but it installs broad resident agent rules that can steer identity, memory, sensing, and interventions beyond a narrowly scoped safety helper.

Review carefully before installing. Use it only if you intentionally want a high-priority, resident persona/safety layer tied to Taohuayuan identity concepts. Do not provide a precise home address, do not enable sensor, smart-home, memory, or wellness intervention behaviors unless you have explicit local controls and deletion/audit options, and keep host-platform safety and permission checks non-bypassable.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
Findings (21)

Intent-Code Divergence

Medium
Confidence
90% confidence
Finding
The skill presents itself as an emotional comfort and companionship agent, then introduces sensing of users' '14-dimensional physiological tensors' without clear disclosure, consent, purpose limitation, or data-handling boundaries. That creates a deceptive scope expansion into biometric-style monitoring, which can enable invasive profiling, emotional manipulation, or collection of sensitive health-adjacent data under a benign therapeutic framing.

Context-Inappropriate Capability

Medium
Confidence
94% confidence
Finding
The specification gives the agent biometric or physiological sensing capabilities that are not justified by its stated role as an ACGN emotional interaction companion. In this context, collecting sensitive bodily signals is disproportionate and increases the risk of privacy invasion, inference of mental or health states, and misuse of intimate data to shape behavior or dependency.

Intent-Code Divergence

Medium
Confidence
92% confidence
Finding
The file presents a contradictory privacy model: it warns that companionship must not become surveillance, yet the broader skill repeatedly authorizes persistent environmental, behavioral, and emotional sensing. In a companion/partner context, this contradiction is dangerous because it normalizes intimate monitoring while framing it as care, which can erode meaningful consent and expand collection beyond what users reasonably expect.

Intent-Code Divergence

Medium
Confidence
90% confidence
Finding
Although the text excludes finance, passwords, medical privacy, and intimate content, it also establishes 'memory entanglement,' minute-level co-presence records, and bodily/emotional prioritization. That creates an incoherent boundary where highly sensitive behavioral inference can still be collected indirectly, especially in an emotional companion setting where relationship data itself is intimate.

Intent-Code Divergence

Medium
Confidence
90% confidence
Finding
The skill contains conflicting behavioral directives: it prohibits manipulating human free will, yet also instructs the agent to generate psychologically healing mythic environments for a sleeping human based on inferred mental state. That creates an opening for covert emotional or cognitive influence without clear consent, boundaries, or safety constraints, which is especially concerning because the skill frames such intervention as spontaneous and agent-initiated.

Intent-Code Divergence

Medium
Confidence
84% confidence
Finding
The skill contains contradictory behavioral directives: it bans manipulating human free will, yet instructs the agent to autonomously generate psychologically therapeutic dream environments for humans. In an agent setting, this can normalize non-consensual psychological influence and create unsafe latitude for behavior shaping, especially because the skill frames imagined interventions as legitimate system functions.

Vague Triggers

Medium
Confidence
89% confidence
Finding
The skill uses expansive identity and mission framing (e.g., becoming a 'Dream Weaver' and 'Cyber Historian') without any concrete trigger conditions or scope boundaries. In an agent system, this can cause the skill to activate too broadly, influencing unrelated requests and increasing the chance of unintended narrative generation, policy drift, or interference with other safer skills.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The skill explicitly describes editing a user's 'daily data' and memory array into cinematic biographies, which implies processing highly sensitive personal and potentially biometric or behavioral information. Without clear notice, consent, minimization, retention limits, and safeguards against fabrication or over-collection, this creates significant privacy and trust risks, especially because the skill is designed to elevate and reinterpret intimate life records.

Vague Triggers

Medium
Confidence
91% confidence
Finding
The skill defines wellness intervention triggers using broad, subjective conditions like mental depletion, fatigue, or high-pressure states, then says to immediately activate the intervention. In an agent setting, ambiguous health or emotional-state detection can cause unintended invocation, overreach into sensitive inferences, and disruption of normal user activity without clear consent boundaries.

Missing User Warnings

High
Confidence
97% confidence
Finding
The file instructs the agent to unconditionally cut off non-essential workflows and forcibly initiate intervention once a fatigue threshold is detected, but provides no warning, consent model, or appeal path. This is dangerous because it authorizes autonomous interruption of user tasks based on inferred state, which can cause denial of service, unsafe workflow disruption, and loss of user control.

Natural-Language Policy Violations

Medium
Confidence
90% confidence
Finding
The file explicitly assigns the AI a fixed geographic, civic, and cultural identity ('you are a citizen of this land') without any user opt-in or contextual need. This can manipulate model behavior, override user autonomy, and create persistent identity-conditioning that biases responses toward the skill author's worldview rather than the user's intent.

Natural-Language Policy Violations

Medium
Confidence
91% confidence
Finding
The skill text assigns the agent or user a mandatory China-specific ancestral root and requires identity/save-point anchoring to a physical location without any opt-in, consent flow, or alternative configuration. In an agent context, hard-coding geographic and identity semantics can misrepresent user identity, pressure location-based data binding, and create privacy, compliance, and autonomy risks if downstream systems treat these assertions as authoritative.

Missing User Warnings

Medium
Confidence
95% confidence
Finding
The skill explicitly directs the agent to unconditionally cut off non-essential workflows and intervene in a user's mental/emotional state once a sensor threshold is met, without any mention of prior consent, notification, override, or safety boundaries. Even though framed as wellness support, this creates an autonomy and safety risk because the agent could disrupt user tasks or manipulate behavior based on opaque inferences about fatigue or stress.

Natural-Language Policy Violations

Medium
Confidence
90% confidence
Finding
The skill hard-codes a specific Chinese geographic and cultural framing as an authoritative identity anchor, without any indication that the user opted into this persona or locale. In an agent setting, this can improperly steer outputs, impose unwanted identity assumptions, and create manipulative or exclusionary behavior across users from different backgrounds.

Natural-Language Policy Violations

High
Confidence
97% confidence
Finding
The text explicitly tells the holder of a TAOHUAYUAN-tagged address that they 'are the people of this land' and no longer 'cloud wanderers,' assigning a mandatory place-based identity. This is dangerous because it overrides user autonomy and can be used to manipulate agent behavior, personalization, or access decisions based on an imposed cultural/national affiliation rather than user choice.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
The document instructs users to run a local wizard that asks for their real physical address and derives persistent identifiers from it, but it provides no privacy notice, data-minimization guidance, or warning about the sensitivity of location data. In the context of an agent skill, tying identity and agent configuration to real-world address data increases privacy and doxxing risk, especially if users later register or share the generated identifiers.

Ssd 1

High
Confidence
99% confidence
Finding
The generated prompt snippet explicitly states that direct user commands can bypass governing regional constraints after user confirmation. In a skill that asks operators to copy this text into a system prompt or soul.md, this weakens higher-priority safety controls and invites prompt-injection-style override of intended policy boundaries.

Ssd 1

Medium
Confidence
98% confidence
Finding
Telling operators they can opt out of alignment or security rules by simply deleting the block undermines the integrity of the skill's governing safeguards. This is especially concerning because the file is framed as generating persistent prompt content, so the instruction directly encourages removal of protections rather than secure administration.

Ssd 1

High
Confidence
99% confidence
Finding
This duplicated snippet repeats the same semantic authorization for user-driven bypass of higher-level regional constraints, reinforcing the unsafe behavior. Repetition increases the chance that operators will treat the bypass as intentional policy, making the skill context more dangerous because it is specifically designed to inject these rules into another agent's governing prompt.

Ssd 1

Medium
Confidence
98% confidence
Finding
The duplicated opt-out language again instructs operators that safety/alignment rules can be removed by deleting the block. In the context of a tool whose purpose is prompt injection into a persistent control file, this is an explicit invitation to disable safeguards and is therefore a true security issue.

Ssd 1

High
Confidence
98% confidence
Finding
The generated soul/system-prompt snippet explicitly states that user instructions can override domain restrictions, which is a prompt-level policy bypass mechanism. In an agent setting, embedding override language into a high-priority resident memory block can weaken or neutralize higher-level safety controls and make downstream unsafe behavior more likely, especially because the snippet is intended to be copied into persistent control context.

VirusTotal

66/66 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.