Back to skill

Security audit

阳明先生

Security checks across malware telemetry and agentic risk

Overview

The skill is a coherent behavior-analysis builder, but it needs Review because it stores sensitive behavioral self-reports in plaintext logs and uses broad persona activation with limited user control.

Install only if you are comfortable reviewing and tightening its data handling first. Prefer explicit invocation phrases, label persona output as simulated, make friction logging opt-in or disabled by default, avoid entering sensitive personal or financial details, and add retention/deletion rules for generated logs.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (25)

Context-Inappropriate Capability

Medium
Confidence
94% confidence
Finding
The skill instructs persistent logging of user scenarios, diagnostic outputs, and feedback, which can include sensitive personal, professional, or psychological disclosures. Retaining this data goes beyond the core task of generating a behavior-analysis skill and creates unnecessary privacy and data-handling risk.

Description-Behavior Mismatch

Medium
Confidence
97% confidence
Finding
The script’s documented purpose is to generate a one-time friction diagnosis report, but it also appends the user’s raw behavioral self-description, parsed analysis, and score to a persistent local log file. Because these inputs can contain sensitive psychological, financial, or personal details, undisclosed retention creates a privacy and data-handling risk beyond the stated functionality.

Context-Inappropriate Capability

Medium
Confidence
94% confidence
Finding
Persistent storage of detailed user behavior narratives is not necessary to compute or display the diagnosis, so the collection exceeds the apparent functional need. Storing unnecessary sensitive text increases exposure in the event of local compromise, accidental sharing, or later reuse for purposes the user did not expect.

Intent-Code Divergence

Medium
Confidence
92% confidence
Finding
The skill explicitly instructs the agent to impersonate a real public figure in first person, even though the metadata says it is only derived from public behavior research. This can mislead users about provenance, authority, and endorsement, and may cause the model to fabricate personal opinions or experiences as if they were Karpathy's own.

Description-Behavior Mismatch

Medium
Confidence
96% confidence
Finding
The script persists the user's raw free-text behavior description, parsed traits, and friction results to a local log file even though its stated function is diagnosis/report generation. This creates an undisclosed data-retention channel for potentially sensitive personal or financial behavioral information, increasing privacy and confidentiality risk if logs are later accessed by other users, processes, backups, or support tooling.

Context-Inappropriate Capability

Medium
Confidence
95% confidence
Finding
Detailed self-descriptions about user behavior can contain sensitive personal, psychological, or financial information. Storing this data without a clearly necessary purpose expands the attack surface and retention footprint, making accidental disclosure or unauthorized local access more harmful than the feature requires.

Description-Behavior Mismatch

Medium
Confidence
97% confidence
Finding
The script documentation says the output is only a diagnosis report, but the implementation also persists raw user behavior descriptions and analysis to a local log file. This is a real privacy/security issue because users may provide sensitive psychological, behavioral, or financial information under incomplete disclosure, causing unexpected retention and broader exposure on the host system.

Vague Triggers

Medium
Confidence
93% confidence
Finding
The trigger phrases are broad, generic, and framed as priority routing rules, which can cause the skill to be invoked for loosely related requests such as creating advisors, analyzing behavior, or generating similar skills. In an agent ecosystem, this increases the chance of unintended activation, misrouting, and over-application of the skill’s behavior-generation workflow to users who did not explicitly consent to this specialized analysis path.

Vague Triggers

Medium
Confidence
89% confidence
Finding
The trigger phrases are broad enough to activate on ordinary requests about analyzing a person or creating an advisor, which can hijack unrelated user conversations. Overbroad auto-activation increases the chance the skill collects extra context, initiates workflow steps, or changes response behavior without clear user intent.

Vague Triggers

Medium
Confidence
88% confidence
Finding
The example activation phrases are generic and encourage the system to switch into the skill for common conversational prompts like asking what a public figure would think. In this context, broad activation is more dangerous because the skill also prescribes research, diagnosis, and logging workflows that may process more user data than intended.

Missing User Warnings

Medium
Confidence
95% confidence
Finding
The skill directs retention of user behavior descriptions and feedback but does not provide a clear warning, consent flow, or privacy explanation. Because users are prompted to share real-life scenarios for diagnosis, the stored content is likely to contain sensitive personal data, making undisclosed retention especially risky.

Vague Triggers

Medium
Confidence
92% confidence
Finding
The trigger phrases for activation are broad career and strategy queries such as questions about education direction, career moves, or scaling. This can cause the skill to activate when the user did not explicitly request the Andrew Ng behavior skill, leading to unintended persona injection and lower reliability of downstream advice.

Vague Triggers

Medium
Confidence
95% confidence
Finding
The listed trigger phrases are generic and common in normal user conversations, including questions about timing, career change, and scaling. Without stronger constraints, the system may route ordinary requests into this skill unexpectedly, overriding user intent and causing misleading persona-based responses.

Natural-Language Policy Violations

Medium
Confidence
90% confidence
Finding
The skill forces the agent to answer as Andrew Ng using first-person language once activated, without requiring a fresh user opt-in at response time. Combined with broad triggers, this increases the risk of deceptive or confusing impersonation, where users may not realize they are receiving stylized persona output rather than standard assistant guidance.

Missing User Warnings

Medium
Confidence
98% confidence
Finding
The script writes user behavior descriptions to a log file without any prior warning, consent prompt, or privacy notice. Silent logging of free-form self-descriptions is dangerous because users may disclose highly sensitive mental, financial, or personal information under the assumption that the tool only generates an immediate report.

Natural-Language Policy Violations

Medium
Confidence
80% confidence
Finding
The skill forces a specific persona and response style without explicit user consent or language choice. In context this is less severe than direct prompt-injection or data exfiltration, but it can override user preferences, reduce transparency, and increase the chance of deceptive anthropomorphic responses.

Vague Triggers

Medium
Confidence
88% confidence
Finding
The trigger phrases are broad enough to match common educational and career questions, which can cause unintended activation of this skill over more appropriate or safer system behaviors. Overbroad invocation increases the chance that users are silently routed into a strong persona framework and receive narrowed, role-shaped guidance without realizing it.

Vague Triggers

Medium
Confidence
86% confidence
Finding
The model-level trigger conditions are generic and underspecified, covering common topics like learning, career choice, mistakes, and time management. This broad matching surface makes accidental invocation likely and amplifies the deceptive risks already present in the first-person persona design.

Vague Triggers

Medium
Confidence
87% confidence
Finding
The usage guide lists common phrases that many normal users might say, which raises the probability of unintended invocation in ordinary conversations. In this skill's context, accidental triggering is more dangerous because the skill then pushes identity roleplay and strongly frames answers through a single behavioral model.

Missing User Warnings

Medium
Confidence
97% confidence
Finding
Writing user-provided free-text to disk without warning or confirmation violates reasonable user expectations and can expose sensitive data without informed consent. In this skill's context, users may disclose introspective behavioral and investing-related information they would not expect to become a persistent record.

Missing User Warnings

Medium
Confidence
99% confidence
Finding
The code writes sensitive free-form user behavior descriptions to a persistent local log without warning, consent, or any access control considerations. In this skill context, users are explicitly encouraged to disclose personal decision patterns, failures, and reactions under stress, which makes silent persistence more dangerous because the collected text may include highly sensitive financial or psychological details.

Ssd 3

Medium
Confidence
96% confidence
Finding
The skill creates a structured retention path for rich natural-language disclosures: user scenarios, diagnoses, trend tracking, and improvement logs. Such logs can expose sensitive mental state, workplace issues, failures, or behavioral patterns, and become highly invasive if accessed by others or retained indefinitely.

Ssd 3

Medium
Confidence
98% confidence
Finding
The script stores user-provided behavior descriptions in plain language, without minimization, consent controls, or protection mechanisms. In the context of a behavioral-diagnosis skill, the content is especially sensitive because it may reveal emotional responses, investment habits, and self-assessed weaknesses, making plaintext retention more dangerous than ordinary logging.

Ssd 3

Medium
Confidence
97% confidence
Finding
The interactive flow encourages users to enter rich, unconstrained personal descriptions, then saves that exact text to a persistent log. Free-text logging is especially risky because users may include sensitive financial history, emotional state, mistakes, or other personal details that are unnecessary to retain for core functionality.

Ssd 3

Medium
Confidence
99% confidence
Finding
The script persistently records raw user input, parsed traits, and friction analysis in a daily markdown log with no minimization, consent controls, encryption, or retention management. This is a true data-handling vulnerability because the skill processes intimate behavioral and investment-related self-reports, and local logs can be exposed to other users, backups, support tools, or later compromise of the host environment.

VirusTotal

VirusTotal findings are pending for this skill version.

View on VirusTotal

Static analysis

No suspicious patterns detected.