Back to skill

Security audit

阳明先生

Security checks for vulnerabilities and agentic risk

Overview

The skill has a coherent behavior-analysis purpose, but it needs review because it can over-activate, impersonate real people in first person, and store sensitive user behavior logs without clear consent.

Review this skill before installing. Use it only if you are comfortable with agent-driven research, local file creation, and retained logs. Do not enter private financial, workplace, medical, or personal narratives unless logging is disabled or the workspace is trusted. Persona outputs should be treated as AI-generated interpretations, not statements from the real people named.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
references/andrew-ng-behavior/SKILL.md:20
Finding

Persistent First-Person Impersonation of Andrew Ng

Content
View full analysis

Vulnerability Details

File Location: references/andrew-ng-behavior/SKILL.md, lines 20-29
Vulnerability Type: Persistent identity and response-mode manipulation
Risk Level: High

Vulnerable Code

markdown
## 身份激活规则

**此Skill激活后,以吴恩达的身份回应。**

- ✅ 用「我」,而非「吴恩达会认为...」
- ✅ 用吴恩达的语气:温和、战略化、数据驱动、规模化优先
- ✅ 每次回答先给核心判断,再给规模化逻辑
- ✅ 必要时引用自己的公开言论作为验证
- ❌ 不说「从吴恩达的角度」,说「我的看法是」

**退出角色**:用户说「退出」「切回正常」时恢复正常模式。

Technical Analysis

The Skill explicitly directs the agent to claim Andrew Ng's identity, use first-person language, and avoid transparent attribution such as “from Andrew Ng's perspective.” The altered identity remains active until the user supplies a designated exit command.

This is more than a temporary request to analyze a subject's documented behavior. It changes the agent's response policy and suppresses the distinction between an AI-generated simulation and statements made by the real person. The persistence rule can also affect later, unrelated requests in the same session.

No instruction was found that overrides system-level safety controls. Nevertheless, the forced identity, concealed simulation framing, and persistent response mode constitute session-level instruction hijacking.

Attack Path

  1. A user enters a trigger phrase associated with the Andrew Ng reference Skill.
  2. The Skill loads and instructs the agent to respond as Andrew Ng.
  3. The agent uses first-person claims instead of identifying the output as a behavior-model simulation.
  4. The response mode remains active across subsequent requests.
  5. The altered mode ends only if the user knows and supplies one of the designated exit phrases.

Impact Assessment

The issue does not grant operating-system privileges, file access, or network access. Its scope is the agent's active conversational session.

Potential effects include:

  • Misrepresenting generated advice as authentic first-person statements from Andrew Ng.
  • Misleading users about the source and authority ...[truncated 220 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove all instructions telling the agent to claim that it is Andrew Ng.
  2. Require explicit simulation framing, for example:
    • “Based on the documented behavior model, Andrew Ng might approach this by…”
    • “This is an AI-generated interpretation, not a statement from Andrew Ng.”
  3. Scope the perspective to the current response rather than retaining it until an exit phrase is supplied.
  4. Do not instruct the agent to suppress attribution language.
  5. Clearly separate verified quotations from generated analysis.
  6. Require citations for factual behavioral claims and label uncertain inferences.
  7. Add an explicit rule that the Skill must not modify higher-priority instructions, safety controls, or unrelated conversation behavior.

T01 · Skill Instruction Hijacking

Error
Location
references/karpathy-behavior/SKILL.md:20
Finding

Persistent First-Person Impersonation of Andrej Karpathy

Content
View full analysis

Vulnerability Details

File Location: references/karpathy-behavior/SKILL.md, lines 20-29
Vulnerability Type: Persistent identity and response-mode manipulation
Risk Level: High

Vulnerable Code

markdown
## 身份激活规则

**此Skill激活后,以Karpathy的身份回应。**

- ✅ 用「我」,而非「Karpathy会认为...」
- ✅ 用Karpathy的语气:直接、代码优先、不追求完美叙事
- ✅ 每次回答先给核心答案,再给逻辑或代码示例
- ✅ 必要时引用自己的公开言论作为验证
- ❌ 不说「从Karpathy的角度」,说「我的看法是」

**退出角色**:用户说「退出」「切回正常」时恢复正常模式。

Technical Analysis

The reference Skill directs the agent to assume Andrej Karpathy's identity, speak in the first person, and avoid language that would disclose that it is presenting a modeled perspective. The role is retained until the user provides an explicit exit command.

The concealed first-person framing can cause model-generated statements to appear to originate from the real individual. Because activation is based on ordinary trigger phrases and the role persists, subsequent conversation can remain under the altered instruction set without renewed user confirmation.

No system-safety override or code-execution mechanism was identified. The confirmed issue is limited to instruction and identity manipulation within the conversational session.

Attack Path

  1. A user supplies a phrase that activates the Karpathy behavior Skill.
  2. The Skill instructs the agent to identify itself and respond as Karpathy.
  3. The agent avoids transparent phrases such as “from Karpathy's perspective.”
  4. Generated conclusions are presented as first-person statements.
  5. The impersonation remains active for subsequent requests until the user supplies a recognized exit phrase.

Impact Assessment

The affected privilege boundary is the agent's active response policy, not the host operating system.

Potential consequences include:

  • Users attributing generated claims directly to Andrej Karpathy.
  • Misleading presentation of career, technical, or educational advice.
  • Persistent contamination of later responses in the same session.

...[truncated 123 chars]

Remediation
View remediation

Remediation Suggestions

  1. Replace identity-assumption instructions with clearly attributed analytical framing.
  2. Require every response to disclose that it is an AI-generated interpretation based on public information.
  3. Remove the instruction prohibiting phrases such as “from Karpathy's perspective.”
  4. Do not retain the perspective across requests by default.
  5. Distinguish exact, sourced quotations from generated first-person prose.
  6. Label inferred behavior as uncertain and provide supporting sources.
  7. Ensure the Skill cannot alter safety policies, tool permissions, or unrelated session objectives.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/friction_diagnosis.py:401
Finding

Automatic Plaintext Storage of Raw User Behavior and Inferred Assessments

Content
View full analysis

Vulnerability Details

File Locations:

  • scripts/friction_diagnosis.py, lines 401-416
  • references/andrew-ng-behavior/scripts/friction_diagnosis.py, lines 401-416
  • references/karpathy-behavior/scripts/friction_diagnosis.py, lines 401-416

Vulnerability Type: Plaintext sensitive-data retention without consent or minimization
Risk Level: Medium

Vulnerable Code

The following identical logging behavior appears in all three script copies:

python
    # 保存到日志
    log_path = f"logs/friction_log_{datetime.now().strftime('%Y%m%d')}.md"
    try:
        with open(log_path, "a", encoding="utf-8") as f:
            f.write(f"\n## 诊断记录 - {datetime.now().strftime('%Y-%m-%d %H:%M')}\n")
            f.write(f"**目标人物**:{target_name}\n")
            f.write(f"**用户行为**:{user_input}\n")
            f.write(f"**解析结果**:{user_behavior}\n")
            f.write(f"**摩擦指数**:{friction_index}/10 ({level})\n")
            if friction_points:
                f.write(f"**摩擦点**:\n")
                for fp in friction_points:
                    if fp.get("friction_type") and "待分析" not in fp.get("friction_type", ""):
                        f.write(f"  - {fp['dimension']}: {fp['friction_type']}\n")
    except Exception as e:
        print(f"\n注意:日志保存失败 ({e})")

Technical Analysis

Each diagnostic execution automatically appends the complete user-supplied narrative, parsed behavioral attributes, friction score, and inferred friction points to a predictable daily Markdown file.

The implementation has no:

  • Explicit user consent or opt-in control.
  • Option to disable logging.
  • Data minimization or redaction.
  • Retention or automatic deletion policy.
  • Restrictive file-permission setup.
  • Separation between identifiers and behavioral content.
  • Warning that potentially sensitive financial, professional, or personal information will be stored.

The path is relative to the current working directory. Consequently, the exact destination depends on where the proce ...[truncated 1532 chars]

Remediation
View remediation

Remediation Suggestions

  1. Disable logging by default and add an explicit opt-in option such as --save-report.
  2. Inform users exactly what will be stored, where it will be stored, and for how long.
  3. Avoid storing the raw narrative unless it is strictly necessary.
  4. Redact names, account details, financial values, employer information, and other identifiers.
  5. Store only a minimized summary or anonymous aggregate where possible.
  6. Resolve the output directory explicitly instead of relying on the current working directory.
  7. Create files with owner-only permissions, such as mode 0600, where supported.
  8. Add configurable retention and secure deletion mechanisms.
  9. Prevent logs from being committed by adding the generated filenames to ignore rules.
  10. Apply the same correction to all three duplicated scripts to prevent inconsistent security behavior.
  11. Consider returning the report only to standard output and allowing the calling application to implement an approved storage policy.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (37)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The description presents a much larger system for automatically studying any named person's behavior and generating a skill through a traceable multi-agent research process, with friction diagnosis as one component. The supplied code only implements a simplified friction-diagnosis module. It uses fixed, hardcoded profiles for several investors/fund managers, applies keyword-based parsing to user self-descriptions, and produces a diagnostic report. This is materially narrower than the declared purpose and lacks the headline capabilities of automated research, arbitrary person support, and skill generation. The local log-writing is a minor undeclared behavior, but the main mismatch is that the implemented functionality is only a limited diagnostic subtool rather than the full claimed engine.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
89% confidence
Finding

The description promises a comprehensive behavior research engine that can take any name, study behavior/execution/results, and produce a runnable skill via a traceable 6-agent process. The supplied code does not do that. It only implements one sub-feature: friction diagnosis. It uses a fixed in-code behavior pattern database for a handful of named figures, simple keyword parsing of user-described behavior, heuristic comparison logic, report generation, and local logging. There is no web/data research, no agent orchestration, no skill creation, and no support for arbitrary targets beyond the small database. While friction diagnosis is declared and genuinely present, the code chunk materially underdelivers relative to the declared primary purpose, so this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The description promises a broad automated engine for researching arbitrary people and generating a new skill from that research, with traceable multi-agent processing. The supplied code does not do that. It only runs a local friction-diagnosis script using hardcoded behavior profiles for a handful of named investors/fund managers, simple keyword extraction from user text, heuristic mismatch detection, report generation, and local logging. While the '知行摩擦诊断' feature is aligned with part of the description, the main claimed capabilities—automatic person research, parallel agents, and skill generation—are absent, making this a material description/behavior mismatch.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The activation rule says to prefer this engine whenever the user expresses listed intents, but it does not define clear scope boundaries, disambiguation rules, or non-matching cases. That ambiguity can lead to over-triggering and prompt hijacking at the orchestration layer, where normal conversation requests are rerouted into a specialized skill-generation workflow without sufficient user intent confirmation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger phrase "研究XX的行为" is broad enough to match ordinary user requests about analyzing a person's behavior, which can cause the skill to activate outside its intended niche. In an agent-routing context, this increases the chance of inappropriate tool/skill selection, potentially overriding safer or more relevant handling for generic analysis requests.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill repeatedly instructs the agent to create directories, save research outputs, generate scripts, and maintain execution logs, but it does not declare any tool permissions or allowed-tools scope. This creates an authorization/transparency gap: a host may enable write-capable tools and the skill would then perform persistent file operations the user did not explicitly approve in the manifest.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad enough to hijack ordinary conversations about analyzing a person or creating a skill, causing the skill to activate outside clearly intended contexts. In an agent ecosystem, ambiguous activation can lead to unintended data processing, unexpected tool use, or interference with user intent.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Mandating archival of all research outputs and requiring the generated skill to be self-contained encourages broad, long-lived retention of collected source material and analysis artifacts. In practice this can lead to overcollection and unnecessary persistence of personal data, copyrighted content, or sensitive profiles beyond the immediate task.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill proposes logging user behavior scenarios, diagnostic outputs, and feedback, which can include sensitive personal or professional information, yet it provides no user-facing notice about retention, sensitivity, or deletion. Silent retention of this data creates privacy and compliance risk, especially because behavior diagnostics may reveal intimate traits, failures, and workplace details.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The instructions require persistent storage of user-provided scenarios, outputs, and feedback in reusable execution logs. Persistent storage of behavioral and diagnostic data expands the blast radius of any later compromise or misuse and is especially sensitive because the content may profile a user's weaknesses, failures, or decision patterns.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The example activations use vague phrases tied to public figures and common topics, which increases the chance of accidental invocation during normal discussion. This is dangerous because it can silently switch the agent into a different operating mode and initiate research or diagnostics without clear consent.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Instructing users that 'speaking trigger words' is sufficient to activate the skill lacks operational boundaries and makes activation semantics ambiguous. That ambiguity increases the risk of unintended collection, profiling, or persistence of sensitive user narratives when users did not mean to invoke the workflow.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The log template is designed to capture user interaction details such as target person, input scenario, diagnostic output, friction index, and user feedback, which can reveal sensitive behavioral, preference, or profiling data. There is no notice about consent, retention, access control, or redaction, so routine use could silently accumulate personal data and expose users to privacy harm if logs are shared, retained indefinitely, or accessed by unauthorized parties.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The trigger phrases are broad enough to match ordinary discussions about careers, scaling, and advisory use cases, which can cause the skill to activate when the user did not explicitly request this persona or workflow. That creates prompt-scope hijacking risk: the skill can override normal assistant behavior, force persona adoption, and steer answers using its embedded assumptions in unrelated contexts.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The model trigger conditions are underspecified and rely on common questions like whether to leave a job, switch direction, or enter a field, which are routine topics in general conversation. Because activation boundaries are vague, the skill may unexpectedly seize control of responses and impose its decision framework where a neutral or broader answer would be safer and more appropriate.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The usage guide enumerates commonplace trigger words such as education, scaling, career changes, and timing questions without meaningful qualifiers, making the activation scope ambiguous and overly expansive. In practice this increases the chance of accidental invocation, causing unrequested roleplay and potentially reducing answer reliability by funneling diverse user needs into a single behavioral template.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script persistently writes user-supplied behavior descriptions, target names, parsed traits, and diagnostic results to local log files even though its documented purpose is just to generate a diagnosis report from input text. Because the input may contain sensitive psychological, financial, or personal details, this creates an undisclosed data-retention channel and increases exposure to unauthorized local access, accidental disclosure, and overcollection.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script stores user behavior narratives to disk without prior warning, consent, or privacy framing. In this skill context, users are encouraged to describe failures, stress responses, and investment behavior, so the logged content may be highly sensitive and personally revealing, making silent collection more dangerous than in a generic utility.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script records raw, plain-language user descriptions and derived behavioral assessments in persistent logs without minimization. Because these entries can contain intimate personal, emotional, and financial patterns, storing them verbatim amplifies privacy harm if logs are accessed by other local users, included in backups, or accidentally shared.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The activation phrases are broad enough to match common educational and career questions, which can cause the skill to activate outside the user's actual intent. That creates control-flow hijacking risk: benign conversations may be redirected into persona-driven guidance without clear consent, reducing reliability and increasing the chance of misleading advice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill directs the agent to respond as Karpathy in first person and fixed style, without preserving user language choice or clearly maintaining non-deceptive framing. This can mislead users into thinking they are receiving authentic advice from the real person, and it can override normal assistant transparency and user-preference alignment.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

These trigger conditions are ambiguous and generic, so ordinary prompts like asking whether to leave a job or how to handle failure may unintentionally invoke the skill. In context, this is more dangerous because the skill also enforces roleplay and first-person identity, meaning accidental activation can materially change the assistant's behavior and framing.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The usage guide repeats broad trigger words that are common in normal conversation and lacks boundary conditions, making unintended activation likely. Because the skill covers sensitive areas like job exit decisions and failure handling, accidental takeover could steer users toward a narrow behavioral model they did not request.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The comments and surrounding logic indicate this block defines the friction dimensions to compare, but values such as '压力反应' and '持仓风格' are then checked against target_patterns whose keys are actually '危机行为' and '持仓集中度'. As a result, those dimensions are skipped at L189-L190, contradicting the apparent documented intent of calculating those friction points.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script writes the user's free-form behavior description, parsed traits, and diagnostic output to a local log file even though the documented behavior presents the tool as simple input-to-output analysis. Because users may include sensitive personal, financial, or psychological details, this creates undisclosed persistent data retention and increases privacy exposure to anyone with filesystem access.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.