Back to skill

Security audit

Step3-VL Finetune

Security checks for vulnerabilities and agentic risk

Overview

The skill is a mostly documentation-only fine-tuning guide, but it publishes internal infrastructure details and overstates multimodal support.

Review this before installing or sharing publicly. Treat the hardcoded server, registry, internal IP, and document path as environment-specific sensitive details, replace them with your own placeholders, and do not send prompts or data to the listed HTTP endpoint unless you control the network and have appropriate authentication or tunneling. Also expect the documented training workaround to be language-only, not full multimodal fine-tuning.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Note
Location
SKILL.md:28
Finding

Exposure of Internal Infrastructure and Deployment Metadata

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 28–31
Vulnerability Type: Internal infrastructure information disclosure
Risk Level: Low

Evidence

markdown
- **Server**: `wphu@gpu506.aibee.cn`
- **Container**: `step3vl-finetune`
- **Project directory**: `/app/` (inside the container)
- **Model path**: `/data/algorithm/tracking/Checkpoints/Step3-VL-10B`

Technical Analysis

The publicly distributable Skill documentation embeds an organization-specific username and hostname, container name, internal filesystem layout, and model-storage path. These values are not credentials and do not provide direct access by themselves, but they disclose operational details that are unnecessary for a reusable fine-tuning guide.

Such information can assist reconnaissance by identifying account naming conventions, infrastructure hostnames, container workloads, application locations, and valuable model-storage directories. An attacker could combine this metadata with credentials obtained through phishing, credential reuse, or another vulnerability to make subsequent targeting more precise.

Attack Path

  1. An attacker obtains or reads the Skill package.
  2. The attacker extracts the server account, hostname, container name, and filesystem paths from SKILL.md.
  3. The attacker uses the identifiers for infrastructure discovery, targeted phishing, password spraying against separately exposed authentication services, or more focused post-compromise navigation.
  4. If access is obtained through an independent weakness, the disclosed paths help the attacker locate the application and model assets more quickly.

Impact Assessment

This issue does not independently grant privileges or prove that the disclosed host is externally reachable. Its direct impact is limited to information disclosure. However, it reduces the effort required to identify and target internal systems and could facilitate access to pro ...[truncated 59 chars]

Remediation
View remediation

Remediation Suggestions

  • Replace organization-specific identifiers with neutral placeholders such as user@gpu-server, /app, and /models/model-name.
  • Store environment-specific deployment values in access-controlled operational documentation rather than the distributed Skill.
  • Review the repository history and related documentation for additional infrastructure details.
  • Establish documentation review controls to detect internal hostnames, usernames, private paths, registry addresses, and service endpoints before publication.
  • If the hostname or account name is sensitive, assess whether it has been exposed elsewhere and monitor relevant authentication services for targeted access attempts.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:227
Finding

Plaintext Internal Inference Endpoint and Additional Internal Resource Disclosure

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 227–229
Vulnerability Type: Plaintext service reference and internal resource information disclosure
Risk Level: Medium

Evidence

markdown
- vLLM inference service: `http://172.18.10.103:8600/v1`
- Docker image: `harbor.aibee.cn/auto_car/preference-align:v1.1`
- System design document: `memory/reports/2026-03-25-rlhf-system-design.md`

Technical Analysis

The Skill publishes a private inference-service address, a private container-registry path, and an internal design-document location. The inference endpoint explicitly uses unencrypted HTTP.

If a user follows the guide and sends inference requests across a network where traffic can be observed or modified, HTTP provides no transport confidentiality or server authentication. A network-positioned attacker could potentially inspect prompts and generated responses, alter requests or responses, or impersonate the service. The actual exploitability depends on network reachability and whether another trusted encryption layer exists; neither authentication nor such a layer is documented.

The registry and document references also reveal organization-specific project naming and internal resource locations. They do not provide direct access, but they improve an attacker's knowledge of the deployment and software supply chain.

Attack Path

  1. An attacker reads the Skill and learns the private service address, registry namespace, image name, and internal document path.
  2. The attacker obtains a network position capable of reaching or observing traffic to the documented endpoint, such as through a compromised internal host or network segment.
  3. A user submits prompts or data to the inference endpoint over HTTP.
  4. The attacker captures sensitive request or response content or modifies traffic in transit.
  5. Separately, the disclosed registry and document identifiers can be used to focus phishing ...[truncated 648 chars]
Remediation
View remediation

Remediation Suggestions

  • Remove internal service, registry, and document references from the distributable Skill or replace them with clearly marked placeholders.
  • Expose inference services through authenticated HTTPS with valid certificate verification.
  • If direct TLS termination is not feasible, require a mutually authenticated secure tunnel and document that requirement explicitly.
  • Require authentication and authorization for the inference API; restrict access through network segmentation and allowlisting.
  • Avoid transmitting sensitive prompts, images, or outputs over plaintext connections.
  • Keep private registry coordinates and internal design-document paths in access-controlled deployment documentation.
  • Review inference-service logs and network controls to determine whether the plaintext endpoint has handled sensitive information.
Vulnerability Patterns
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (3)

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
80% confidence
Finding

The manifest description says the skill includes a complete configuration, training, and inference workflow for Step3-VL fine-tuning. Yet the guide later documents a monkey patch that bypasses multimodal features and explicitly lists 'implement full vision + language joint training' as future work, contradicting the claim of a complete multimodal fine-tuning flow.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest and title present this as a multimodal Step3-VL fine-tuning workflow, and the inference examples also imply image-conditioned use. However, the core workaround explicitly says to 'skip the vision encoder' and the patched forward uses only language-model embeddings, which is a materially narrower behavior than multimodal fine-tuning.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This markdown file documents code that creates directories and writes adapter_model.bin and adapter_config.json to an output path, but it does not include any warning that running the workflow will modify the filesystem or may overwrite existing outputs. For markdown files, SQP-2 applies when the skill description omits warnings about behaviors that could affect user data or system state.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.