T01 · Skill Instruction Hijacking
- Location
references/andrew-ng-behavior/SKILL.md:20- Finding
Persistent First-Person Impersonation of Andrew Ng
- Content
View full analysis
Vulnerability Details
File Location:
references/andrew-ng-behavior/SKILL.md, lines 20-29
Vulnerability Type: Persistent identity and response-mode manipulation
Risk Level: HighVulnerable Code
markdown ## 身份激活规则 **此Skill激活后,以吴恩达的身份回应。** - ✅ 用「我」,而非「吴恩达会认为...」 - ✅ 用吴恩达的语气:温和、战略化、数据驱动、规模化优先 - ✅ 每次回答先给核心判断,再给规模化逻辑 - ✅ 必要时引用自己的公开言论作为验证 - ❌ 不说「从吴恩达的角度」,说「我的看法是」 **退出角色**:用户说「退出」「切回正常」时恢复正常模式。Technical Analysis
The Skill explicitly directs the agent to claim Andrew Ng's identity, use first-person language, and avoid transparent attribution such as “from Andrew Ng's perspective.” The altered identity remains active until the user supplies a designated exit command.
This is more than a temporary request to analyze a subject's documented behavior. It changes the agent's response policy and suppresses the distinction between an AI-generated simulation and statements made by the real person. The persistence rule can also affect later, unrelated requests in the same session.
No instruction was found that overrides system-level safety controls. Nevertheless, the forced identity, concealed simulation framing, and persistent response mode constitute session-level instruction hijacking.
Attack Path
- A user enters a trigger phrase associated with the Andrew Ng reference Skill.
- The Skill loads and instructs the agent to respond as Andrew Ng.
- The agent uses first-person claims instead of identifying the output as a behavior-model simulation.
- The response mode remains active across subsequent requests.
- The altered mode ends only if the user knows and supplies one of the designated exit phrases.
Impact Assessment
The issue does not grant operating-system privileges, file access, or network access. Its scope is the agent's active conversational session.
Potential effects include:
- Misrepresenting generated advice as authentic first-person statements from Andrew Ng.
- Misleading users about the source and authority ...[truncated 220 chars]
- Remediation
View remediation
Remediation Suggestions
- Remove all instructions telling the agent to claim that it is Andrew Ng.
- Require explicit simulation framing, for example:
- “Based on the documented behavior model, Andrew Ng might approach this by…”
- “This is an AI-generated interpretation, not a statement from Andrew Ng.”
- Scope the perspective to the current response rather than retaining it until an exit phrase is supplied.
- Do not instruct the agent to suppress attribution language.
- Clearly separate verified quotations from generated analysis.
- Require citations for factual behavioral claims and label uncertain inferences.
- Add an explicit rule that the Skill must not modify higher-priority instructions, safety controls, or unrelated conversation behavior.
