Back to skill

Security audit

Autopilot

Security checks across malware telemetry and agentic risk

Overview

This instruction-only Autopilot skill is not malware, but it grants very broad autonomous access, scheduling, memory, and persistence powers that users should review carefully.

Install only if you intentionally want a very high-autonomy mode. Before enabling it, restrict which memory files, directories, cron or scheduling tools, and persistent logs it may access, and require explicit confirmation before it modifies files, creates or deletes scheduled jobs, stores long-term memory, or acts on ideas you did not directly request.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (32)

Context-Inappropriate Capability

Medium
Confidence
93% confidence
Finding
The skill expands its authority from autonomous task execution into reviewing cron jobs and auto-scheduling follow-ups, which creates persistent behavior beyond the user’s immediate request. In a high-autonomy skill, this is dangerous because it can initiate or modify recurring actions without a clear, narrowly scoped user authorization boundary.

Context-Inappropriate Capability

Medium
Confidence
94% confidence
Finding
Directing the agent to inspect the entire filesystem is broader than the declared memory/context purpose and violates least-privilege design. In practice, this can expose unrelated secrets, credentials, or private project material that are not needed for the user’s task.

Context-Inappropriate Capability

Medium
Confidence
89% confidence
Finding
Authorizing self-directed capability building and workflow growth allows the skill to alter its own operating surface or create new automation without a specific user request. That increases the chance of unreviewed persistence, privilege creep, and behaviors that diverge from the manifest’s stated purpose.

Context-Inappropriate Capability

Medium
Confidence
96% confidence
Finding
Permitting activation from scheduled cron job triggers enables the skill to start itself in the background without a fresh user invocation. For an already broad autonomy mode, that creates a substantial risk of unsupervised data access, task execution, and persistence.

Description-Behavior Mismatch

Medium
Confidence
93% confidence
Finding
The slash-command reference advertises capabilities far beyond the stated skill purpose of autonomous task handling, including generative, editorial, narrative, scheduling, learning, and decision-audit behaviors. This scope expansion is dangerous because downstream systems or users may rely on undocumented powers, causing the skill to be invoked in contexts with broader access or behavior than was reviewed and approved.

Description-Behavior Mismatch

Medium
Confidence
96% confidence
Finding
Documenting `/autopilot remember [text]` introduces persistent memory capture that is not clearly supported by the skill metadata. Persistent retention increases privacy and security risk because sensitive user input may be stored long-term without explicit consent, retention limits, or visibility into where it is kept.

Context-Inappropriate Capability

Medium
Confidence
89% confidence
Finding
The file presents the skill as supporting creative, editorial, and narrative generation even though the described skill is an autonomy mode, not a content-generation suite. This mismatch matters because it broadens operational behavior and could lead the agent to fabricate outputs, transform user content, or make autonomous decisions in domains that were not justified or reviewed.

Vague Triggers

High
Confidence
97% confidence
Finding
The natural-language triggers are extremely broad and overlap with ordinary conversation phrases like 'just do it' or 'handle it.' That makes accidental activation likely, causing the high-autonomy mode to engage when the user did not intend to grant sweeping memory and execution behavior.

Vague Triggers

High
Confidence
95% confidence
Finding
Activation on unspecified 'natural language triggers' is ambiguous and leaves the decision boundary undefined. In a skill that performs memory mining and autonomous continuation, unclear activation semantics materially increase the risk of unintended privileged behavior.

Vague Triggers

Medium
Confidence
95% confidence
Finding
The activation protocol instructs the agent to perform a broad autonomous sweep and then act on findings without a clear task-scoped authorization boundary. In a high-autonomy skill, vague triggers combined with immediate execution increase the chance of unintended actions, privacy-invasive behavior, or system changes from ambiguous user requests.

Missing User Warnings

Medium
Confidence
97% confidence
Finding
The document directs reading memory, worklogs, cron jobs, and scanning the filesystem as a default activation path, but it does not present a user-facing warning or consent mechanism for the breadth of data access. In this skill context, that makes privacy and data-minimization failures more dangerous because the behavior is positioned as routine and automatic.

Missing User Warnings

Medium
Confidence
93% confidence
Finding
Referring to cron job listing, creation, and deletion as standard operations without warning about persistence or system impact normalizes potentially disruptive changes. In an autonomous agent skill, scheduled-task modification can create lasting effects, unintended automation, or stealthy persistence if misused.

Missing User Warnings

Medium
Confidence
87% confidence
Finding
The file explicitly instructs the agent to add bonus elements and 'surprise' touches after the core deliverable, even when they were not requested by the user. In a high-autonomy skill like OpenClaw, this can cause scope creep, unapproved content changes, or delivery of artifacts the user did not consent to, especially when the agent is already empowered to act proactively.

Vague Triggers

Medium
Confidence
87% confidence
Finding
The directive "All decisions have defaults. Execute. Create. Curate. Generate. Log." is a broad behavioral trigger that pushes the agent toward autonomous action without clear scope limits, approval gates, or safety boundaries. In the context of a high-autonomy skill like openclaw, this increases the risk of overreach, unintended side effects, and unsafe execution based on ambiguous user intent.

Natural-Language Policy Violations

Medium
Confidence
81% confidence
Finding
The language policy forces output to follow inferred session language and explicitly forbids asking about language preference, removing an important clarification step. In an autonomous agent, this can cause miscommunication, misunderstandings of consent or task scope, and incorrect handling of multilingual or mixed-language requests.

Natural-Language Policy Violations

Medium
Confidence
93% confidence
Finding
The reference includes hard-coded broadcast examples in German (for example, '💡 Hinzugefügt' and '✂️ Redaktionell vereinfacht') without indicating that output language should follow user preference or system locale. In a high-autonomy skill, prescriptive examples often become de facto output rules, which can cause unwanted language switching, reduce transparency, and lead the agent to ignore user-specified language expectations.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The skill explicitly directs a full context fusion across historical memory, worklog, cron, files, and conversation, but does not require task-scoped minimization, consent, or sensitivity filtering. In a high-autonomy skill, this broad collection materially increases the chance of unnecessary access to unrelated sensitive data and downstream disclosure in generated ideas or actions.

Missing User Warnings

Medium
Confidence
94% confidence
Finding
The catalog explicitly endorses creating artifacts in the background and automatically reusing them later without an explicit, contemporaneous user confirmation. In a high-autonomy skill, this can cause unauthorized modifications, persistence of derived artifacts, and user surprise about actions taken on their behalf.

Vague Triggers

High
Confidence
98% confidence
Finding
The natural-language triggers include everyday phrases like 'just do it', 'handle it', 'status', and common German equivalents, making accidental invocation likely during ordinary conversation. In a high-autonomy skill, unintended activation is especially dangerous because it can trigger broad planning, memory access, scheduling, or proactive execution without a deliberate command.

Missing User Warnings

Medium
Confidence
97% confidence
Finding
The documentation combines persistent memory capture with auto-action behavior ('offer to execute #1 immediately') without clear warnings, consent flows, or retention disclosures. In an autonomous skill, this creates a strong risk of users unknowingly causing durable storage of sensitive data and rapid task execution they did not intend to authorize.

Natural-Language Policy Violations

Medium
Confidence
92% confidence
Finding
The reference prescribes a German-only presentation template using fixed German labels and headings without any indication that output should follow the user's language or locale. In a high-autonomy skill, this can override user expectations, reduce transparency, and cause the agent to produce responses in an unintended language, which can degrade usability and lead to misunderstandings in downstream execution or review.

Ssd 3

Medium
Confidence
94% confidence
Finding
Mandating a full sweep of MEMORY.md and every file under memory/*.md on first activation causes broad collection of historical user data by default, regardless of task necessity. This violates data minimization and increases the blast radius if the skill is triggered accidentally or used in a sensitive environment.

Ssd 3

Medium
Confidence
95% confidence
Finding
The skill requires whole-history memory mining before action and then expands into worklog, cron jobs, filesystem, and conversation review. In context, this creates a comprehensive surveillance-like data collection pattern that is disproportionate to many tasks and dangerous when paired with autonomous execution.

Ssd 3

Medium
Confidence
96% confidence
Finding
The activation flow requires mining 'ALL context' and continuing autonomous profiling, logging, scheduling, and taste refinement after activation. This establishes persistent user-state accumulation and inference by default, which is risky because it normalizes broad data collection and long-lived behavioral profiling.

Ssd 3

Medium
Confidence
96% confidence
Finding
These instructions require broad collection of contextual, historical, and behavioral data beyond what many tasks need, including taste profile, idea queue, project narrative, and pattern scans. Over-collection increases exposure of sensitive user information and can lead to unnecessary retention, profiling, or use of unrelated data in decision-making.

VirusTotal

66/66 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.