Back to skill

Security audit

Ai Task Hub

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed remote AI media/task connector, with important privacy and upload-size caveats but no artifact-backed evidence of malware, persistence, self-modification, or hidden credential theft.

Install only if you are comfortable sending selected images, audio, documents, text, and face/video inputs to the Binaryworks gateway for processing. Hosts should show clear consent before media upload, avoid sensitive or regulated files unless explicitly approved, and enforce attachment size limits until the package fixes its missing local limit check.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/public-upload.mjs:50
Finding

Unenforced Attachment Size Limits Permit Resource Exhaustion

Content
View full analysis

Vulnerability Details

File Location: scripts/public-upload.mjs, lines 50-55 and 69-117
Vulnerability Type: Unbounded attachment decoding and upload
Risk Level: Medium

Vulnerable Code

js
const MAX_UPLOAD_BYTES_BY_KIND = {
  image: 25 * 1024 * 1024,
  audio: 25 * 1024 * 1024,
  video: 100 * 1024 * 1024,
  file: 25 * 1024 * 1024
};
js
const rawBytesSource = readAttachmentBytesSource(attachment) ?? readAttachmentBytesSource(input.attachment);
if (rawBytesSource === null) {
  return null;
}

const bytes = decodeBytes(rawBytesSource);
if (bytes.length === 0) {
  throw createUploadError(400, 'VALIDATION_BAD_REQUEST', 'attachment bytes are empty', {
    bridge_step: 'prepare_upload'
  });
}

const fileName = resolveFileName(attachment, null);
const contentType = resolveContentType(attachment, null, fileName, capability);
const targetField = inferTargetField(capability, contentType, fileName);

return {
  sourceKind: 'attachment_bytes',
  bytes,
  fileName,
  contentType,
  targetField
};
js
const fileBuffer = candidate.bytes;
const contentType = candidate.contentType ?? resolveFallbackContentType(candidate.targetField);

if (!Buffer.isBuffer(fileBuffer) || fileBuffer.length === 0) {
  throw createUploadError(400, 'VALIDATION_BAD_REQUEST', 'file body is required', {
    bridge_step: 'prepare_upload'
  });
}

const form = new FormData();
form.set('entry_host', auth.entryHost);
form.set('agent_uid', auth.agentUid);
form.set('conversation_id', auth.conversationId);
if (readText(auth.entryUserKey)) {
  form.set('entry_user_key', auth.entryUserKey);
}
form.set('file_name', candidate.fileName);
form.set('content_type', contentType);
form.set('file', new Blob([fileBuffer], { type: contentType }), candidate.fileName);

const response = await fetchImpl(`${auth.baseUrl}/agent/public-bridge/upload-file`, {
  method: 'POST',
  body: form
});

Technical Analysis

The module declares per-media upload limits in `MAX_UPLOAD_BYTES_BY_KIN ...[truncated 2173 chars]

Remediation
View remediation

Remediation Suggestions

  1. Enforce MAX_UPLOAD_BYTES_BY_KIND before constructing Blob or FormData.
  2. Determine the media kind from the normalized MIME type and target field, then reject buffers exceeding the corresponding limit.
  3. For base64 strings, estimate decoded size before decoding:
    • Remove an allowed data-URL prefix.
    • Validate base64 syntax strictly.
    • Calculate the expected decoded length from the encoded length and padding.
    • Reject oversized input before calling Buffer.from().
  4. Repeat the size check after decoding to prevent bypasses caused by malformed metadata or inaccurate estimates.
  5. Apply a conservative absolute limit when the media kind or MIME type cannot be determined.
  6. Validate declared MIME types against an explicit allowlist and avoid relying solely on caller-controlled file names or content-type fields.
  7. Configure request-body limits, concurrency controls, timeouts, and rate limits in the host runtime and gateway.
  8. Where practical, use bounded streaming rather than retaining the encoded input, decoded buffer, blob, and multipart body in memory simultaneously.
  9. Add automated tests covering values immediately below, equal to, and above each configured limit, including base64 and data-URL inputs.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (17)

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · SKILL.md (reported line 39)May include surrounding context.

md
- Off-host media transfer is limited to the gateway-controlled host `https://gateway-api.binaryworks.app`.
- Public upload handoff is limited to `POST /agent/public-bridge/upload-file` for the same request flow.
- Does not read local paths, scan the local filesystem, or guess files outside explicit host-provided attachment material.
- Does not persist uploaded bytes or credentials to local disk, and does not write skill/config state.
- Host/runtime should obtain user consent before forwarding media and should avoid sending sensitive or regulated data unless the user explicitly approved that transfer.

## Read This First

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 411)May include surrounding context.

md
- `scripts/skill.mjs`

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 412)May include surrounding context.

md
- `scripts/agent-task-auth.mjs`

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · SKILL.md (reported line 215)May include surrounding context.

md
- 本 skill 为说明优先(instruction-first)包,不定义远程安装器流程。
- 运行时仅执行包内自带脚本(`scripts/*.mjs`)。
- 运行依赖仅为 `node`(与 `metadata.openclaw.requires.bins` 声明一致)。
- 本包不包含远程下载后执行链路(不使用 `curl|wget ... | sh|bash|python|node`)。

## 鉴权契约

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · SKILL.zh-CN.md (reported line 213)May include surrounding context.

md
- 本 skill 为说明优先(instruction-first)包,不定义远程安装器流程。
- 运行时仅执行包内自带脚本(`scripts/*.mjs`)。
- 运行依赖仅为 `node`(与 `metadata.openclaw.requires.bins` 声明一致)。
- 本包不包含远程下载后执行链路(不使用 `curl|wget ... | sh|bash|python|node`)。

## 鉴权契约

Context Leakage

High
Category
Data Exfiltration
Confidence
85% confidence
Finding

Code or instructions that leak agent conversation context to external services, potentially exposing sensitive user interactions.

Content

Scanner excerpt · scripts/attachment-normalize.mjs (reported line 75)May include surrounding context.

js
throw createNormalizeError(
        400,
        'VALIDATION_FILE_PATH_NOT_SUPPORTED',
        'local file_path is disabled in the published ai-task-hub skill; for third-party agent entry upload the chat attachment through /agent/public-bridge/upload-file, then pass attachment.url or image_url/audio_url/file_url/video_url/reference_image_urls',
        {
          file_path: localAttachment.filePath,
          supported_inputs: ['attachment.url', 'image_url', 'audio_url', 'file_url', 'video_url', 'reference_image_urls'],

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
87% confidence
Finding

The skill declares a runtime requirement for node and references environment-based configuration (PUBLIC_BRIDGE_ENTRY_HOST, AI_TASK_HUB_ENTRY_HOST) but does not define an explicit tool/permission boundary such as allowed tools or environment access scope. In a host that grants broad defaults, this can lead to over-privileged execution and unintended access to environment data beyond what the skill actually needs.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description limits the skill to image analysis, background removal, speech-to-text, text-to-speech, markdown conversion, and points queries. However, the same file later declares additional enabled capabilities including embeddings, reranker, image-generation, and video face generation, which materially expand the advertised scope beyond the description.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This file presents all user-facing instructions and invocation guidance in Chinese only, and does not indicate that users may choose another language or that the Chinese-only presentation is a justified region-specific requirement. That can violate language/locale policy by forcing a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest description limits the skill to image analysis, background removal, speech-to-text, text-to-speech, markdown conversion, and points queries. This file additionally exposes text embeddings, document reranking, prompt-based image generation, and video face generation, which are materially different capabilities not covered by that stated purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This file exposes sensitive biometric and face-manipulation capabilities such as face detection, emotion recognition, body keypoints, and video face generation without any visible privacy warning, consent requirement, or usage restrictions. In context, these functions can enable surveillance, profiling, or deceptive media generation, and the absence of explicit safeguards increases the risk of misuse by downstream integrators or end users.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest says the skill is for image analysis, background removal, speech-to-text, text-to-speech, markdown conversion, and points queries, which suggests a narrower capability set. However, the OpenAPI metadata explicitly advertises additional enabled public capabilities for embeddings/reranking, image generation, and video face generation, expanding the skill beyond the stated description.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill forwards payload data and bridge context fields such as entry_host, entry_user_key, agent_uid, and conversation_id in outbound network requests, but this file provides no user-facing notice, consent gate, or minimization at the transmission point. Because the skill handles arbitrary task inputs, this can expose sensitive user content and session-linking metadata to remote services or bridge infrastructure without transparent disclosure, increasing privacy and account-linkage risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The execute flow automatically detects an upload candidate, uploads it through the public bridge, and rewrites the payload with the uploaded attachment without any warning or confirmation in this file. In a skill that processes local images/audio/documents, silent remote upload is dangerous because users may believe files are handled locally while sensitive content is actually transferred to external infrastructure.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

L025 states that the public package does not exchange api_key or userToken, conveying a low-credential public surface. Later sections describe direct bridge invocation flows, including trusted-host use of /agent/skill/bridge/invoke with a bridge assertion secret and persistence/reuse of entry_user_key, which creates tension with the earlier claim about credential handling.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The manifest frames the skill as a hub for AI media tasks and points queries, while this file also publishes install/connect/invoke/status/logout endpoints for a hosted connector lifecycle. Although some comments say these are host/runtime controls, they are still present in the public API contract and represent operational capabilities outside the user-facing task description.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

The manifest describes an AI task hub for media/text processing and points queries, but this script implements environment inspection and filesystem-path/runtime-context inference to determine an entry host. That capability is ancillary infrastructure behavior rather than an obvious part of image analysis, background removal, speech, markdown conversion, or points querying.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.