Back to skill

Security audit

Digital Twin

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed digital-twin builder, but it warrants Review because it profiles and can persistently impersonate real people using sensitive communications.

Install only when you have clear authorization to profile the target person and access the connected data sources. Review which MCP connectors the agent can read, avoid using broad defaults over email/chat/docs unless intended, and do not set a real-person persona as a default without clear disclosure and an easy way to turn it off.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This skill is built to clone the communication style, judgment, and audience adaptation of a named real person, yet its description frames this mainly as a functional capability and does not prominently warn about impersonation, consent, deception, reputational harm, or misuse when deployed broadly or as a default. That omission is dangerous because it normalizes a high-risk capability and reduces the likelihood that users or downstream agents will apply appropriate safeguards before generating convincing real-person stand-ins.

Content

No source excerpt is available for this finding.

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
70% confidence
Finding

Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Content

Scanner excerpt · references/personality_skill_template.md (reported line 155)May include surrounding context.

md
Return to normal Claude behavior after unless instructed otherwise.

### Persistent Mode
When this skill is set as the default personality or the user says "always respond as {name}",
keep this skill active for ALL responses in the session. Every message goes through the full
7-step pipeline.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

This file explicitly instructs the system to infer personality traits, conflict style, risk tolerance, and emotional tendencies from a person's communications using psychometric frameworks, but it provides no user warning, consent boundary, or safeguard around sensitive trait inference. In the context of a digital-twin skill designed to model real named people, this creates a meaningful privacy and profiling risk because users may generate invasive personality assessments from emails, chats, and transcripts without the subject's knowledge or appropriate limitation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The generated personality skill is explicitly designed to trigger on broad natural-language phrases such as "be {name}" and "use {name}'s personality," and the parent skill also suggests it may be set as a default persona for all communications. In the context of a real-person digital twin, broad activation conditions materially increase the risk of unintended impersonation, misattributed outputs, and silent persona switching in ordinary conversations.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger language activates on 'any similar instruction asking the agent to adopt {name}'s persona,' which is open-ended and can match ambiguous requests. In a personality-cloning skill, this increases the chance of unintended persona takeover, causing the agent to answer in an impersonated voice without a clear, explicit user opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The persistent-mode instruction allows session-wide persona takeover when the user says 'always respond as {name}' or when the skill is set as default personality. That broad, durable activation can override normal assistant behavior across unrelated turns, increasing the risk of impersonation misuse, policy bypass through role-framing, and accidental continuation after the original context has changed.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/pillar_4_audience.md (reported line 135)May include surrounding context.

md
- **Personal disclosure**: Do they share personal anecdotes or keep things strictly professional? Does this vary?
- **Empathy expression**: Do they explicitly acknowledge feelings or challenges? More with certain audiences?
- **Recognition/praise**: Do they give verbal recognition? To whom?
- **Trust signals**: What indicators suggest they trust (or don't trust) different audiences? (Sharing concerns, being vulnerable, delegating without checking)

---

Static analysis

No suspicious patterns detected.