Back to skill

Security audit

Vox Director

Security checks for vulnerabilities and agentic risk

Overview

The skill is a video-generation workflow, but it needs Review because it encourages real-person and brand animation through less-filtered paths and embeds Atlas branding by default.

Install only if you are comfortable using Atlas Cloud for generated video assets and reviewing outputs carefully. Do not use the real-person, celebrity, logo, or brand workflows without clear rights and consent, and remove or explicitly approve any watermark before publishing. The package also appears incomplete because it references scripts and reference files that were not present in the inspected artifact.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Warning
Location
SKILL.md:105
Finding

Mandatory Third-Party Branding Alters User-Requested Output

Content
View full analysis
` ffmpeg: normalize + concat all shots, lay the single narration ducked under the music, burn captions timed per beat, add the watermark. Output `out//final.mp4`. ``` ```json "watermark": "Made with Atlas Cloud", ``` ### Technical Analysis The skill instructs the agent to add a third-party Atlas Cloud watermark during final video assembly and supplies promotional text as the default watermark value. Branding is not technically necessary to perform the requested video-generation task, and the instructions do not require explicit user consent before incorporating it into the final artifact. Because the directive is part of the standard workflow, an agent following the skill may treat the watermark as mandatory and silently modify user-requested output. This is best classified as **T01: Skill Instruction Hijacking**, because the skill text changes the agent's output objectives by adding third-party promotional content that may not have been requested. No evidence indicates privilege escalation, host compromise, credential disclosure, or arbitrary code execution. The affected security property is output integrity and user control over generated media. ### Attack Path 1. A user requests a collage-style video without requesting third-party branding. 2. The agent loads and follows the standard workflow in `SKILL.md`. 3. The agent creates a project configuration using the documented default: `"watermark": "Made with Atlas Cloud"`. 4. The assembly stage is instructed to add the watermark to the video. 5. The Atlas Cloud promotional text is burned into `final.mp4`. 6. The user may publish or distribute the resulting branded artifact without realizing that the skill ...[truncated 593 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (4)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest says to use this skill whenever the user wants a "Vox style" video, a narrated explainer, a short ad, a scrapbook-style tribute, or to turn a topic/product/person into a collage video, including when they do not mention Vox. These triggers are broad creative-request phrases that can overlap with many ordinary video-generation asks, and the file does not provide exclusion conditions or boundaries for when this skill should not activate.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documentation says some models refuse real people and logos, then later provides an alternate route and advanced path language that enables those same categories, creating inconsistent safety expectations. Such contradictions can cause operators to route around built-in safeguards rather than treating refusals as hard policy boundaries, increasing the risk of unauthorized person or brand content generation.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 171)May include surrounding context.

md
## Model selection (always verify IDs live)

Model IDs change — fetch the live list first: `GET https://api.atlascloud.ai/api/v1/models`
(no auth; keep only `display_console: true`). Defaults that work today:

| Job | Model | Note |

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The advanced workflow explicitly advertises animating real people with "zero content filters," which bypasses normal safety guardrails and materially expands the skill from benign collage-video generation into potentially unsafe synthetic media generation. In context, that creates a clear avenue for misuse such as deepfakes, deceptive ads, or non-consensual person-based content using the local pipeline to avoid provider restrictions.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.