Back to skill

Security audit

huawei-cloud-openviking-embedding-switch

Security checks for vulnerabilities and agentic risk

Overview

The skill is purpose-aligned, but it can delete OpenViking vector-database data and restart processes without strong confirmation, path validation, or recoverable rollback safeguards.

Review before installing or using. This skill is not showing deception or exfiltration, but it should only be used by an operator who understands that it may delete the OpenViking vector index and restart the service. Prefer adding an explicit confirmation flag, dry-run mode, canonical path checks for SANDBOX_DIR, and a move-to-backup rollback flow before allowing automated use.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/switch-embedding-model.sh:95
Finding

Destructive Vector Database Deletion Lacks Enforced Confirmation and Recoverable Rollback

Content
View full analysis

Vulnerability Details

File Location: scripts/switch-embedding-model.sh, lines 95–126 and 177–181
Vulnerability Type: Destructive operation without an enforced confirmation or data rollback mechanism
Risk Level: High

Vulnerable Code

bash
# ── Step 3: Modify ov.conf ──
echo ""
echo "[3/6] Modifying ov.conf..."

# BUG-5 fix: backup original config for rollback
cp "${SANDBOX_DIR}/process_dir/ov.conf" "${SANDBOX_DIR}/process_dir/ov.conf.bak"

python3 -c "
import json
conf_path = '${SANDBOX_DIR}/process_dir/ov.conf'
with open(conf_path, encoding='utf-8') as f:
    data = json.load(f)
dense = data['embedding']['dense']
dense['provider'] = 'openai'
dense['model'] = '${MODEL_NAME}'
dense['api_key'] = 'not-needed'
dense['api_base'] = 'http://127.0.0.1:${LLAMA_PORT}/v1'
dense['dimension'] = ${TARGET_DIMENSION}
with open(conf_path, 'w', encoding='utf-8') as f:
    json.dump(data, f, indent=2, ensure_ascii=False)
print('  ov.conf updated (backup at ov.conf.bak)')
"

# ── Step 4: Delete incompatible vectordb index ──
echo ""
echo "[4/6] Checking vectordb index compatibility..."
COLLECTION_META="${SANDBOX_DIR}/data/vectordb/context/collection_meta.json"
if [ -f "$COLLECTION_META" ]; then
  CURRENT_DIM=$(python3 -c "import json; print(json.load(open('$COLLECTION_META'))['Dimension'])")
  if [ "$CURRENT_DIM" != "$TARGET_DIMENSION" ]; then
    echo "  Dimension mismatch: $CURRENT_DIM → $TARGET_DIMENSION"
    echo "  Deleting vectordb/context..."
    rm -rf "${SANDBOX_DIR}/data/vectordb/context"
    echo "  Deleted."
  else
    echo "  Dimensions match ($CURRENT_DIM), no deletion needed."
  fi
else
  echo "  No existing collection_meta.json, skipping."
fi

The failure rollback restores only the configuration:

bash
if [ "$HEALTHY" != "True" ]; then
  echo "  ❌ Server not healthy after ${HEALTH_TIMEOUT}s"
  echo "  Rolling back config..."
  cp "${SANDBOX_DIR}/process_dir/ov.conf.bak" "${SANDBOX_DIR}/process_dir/ov.conf"
  echo "  Check log: ${SANDB
...[truncated 2576 chars]
Remediation
View remediation

Remediation Suggestions

  1. Require an explicit option such as --confirm-delete-index before deleting the database context. Fail closed when it is absent.
  2. In interactive operation, display the existing dimension, measured target dimension, exact directory, and deletion consequences before requesting confirmation.
  3. Add a --dry-run mode that reports all planned configuration, process, and filesystem changes.
  4. Replace immediate deletion with an atomic rename, for example:
    bash
    mv "${SANDBOX_DIR}/data/vectordb/context" \
       "${SANDBOX_DIR}/data/vectordb/context.backup.$(date +%s)"
    
  5. Restore the preserved directory on every failure path, including failed execution, failed health checks, unchanged PID, dimension mismatch, and detected startup errors.
  6. Delete the preserved directory only after the new server, PID, collection dimension, and logs have all passed verification.
  7. Validate that SANDBOX_DIR is non-empty, canonical, and constrained to the expected OpenViking sandbox root before any recursive filesystem operation.
  8. Document whether the vector store is reproducible from authoritative source data and require a separate backup when it is not.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (49)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
93% confidence
Finding

The documented workflow includes rm -rf deletion of the vectordb/context directory, which is an inherently destructive operation. Even though the intent is maintenance, a path mix-up, symlink issue, or incorrect working directory could cause irreversible data loss, and the skill provides no hardened safeguards around that deletion.

Content

Scanner excerpt · SKILL.md (reported line 32)May include surrounding context.

├── Detect current config (Read ov.conf embedding.dense section) ├── Validate endpoint (Check llama-server /v1/embeddings) ├── Modify ov.conf (Update provider, model, api_base, dimension) ├── Delete vectordb index (If dimension changed: rm -rf vectordb/context) ├── Restart server (Kill + exec, NOT stop/start) └── Verify (Health + PID + dimension + log check)

text

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · SKILL.md (reported line 74)May include surrounding context.

md
> **Prerequisite check: job-env-manager running**
> ```bash
> curl -s http://127.0.0.1:8090/api/v1/envs/openviking | python3 -c "import sys,json; print(json.load(sys.stdin)['state'])"
> ```

- **job-env-manager** running on `http://127.0.0.1:8090`

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · SKILL.md (reported line 101)May include surrounding context.

Task 2: Validate Target Embedding Endpoint

bash
curl -s http://127.0.0.1:${LLAMA_PORT}/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model":"${MODEL_NAME}","input":"test"}' \
  | python3 -c "import sys,json; d=json.load(sys.stdin); print(len(d['data'][0]['embedding']))"

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

This step explicitly instructs destructive recursive deletion of the vector database when dimensions differ. In context it may be operationally necessary, but it still poses substantial integrity risk because an agent could delete the wrong directory or erase recoverable index data without backup or confirmation.

Content

Scanner excerpt · SKILL.md (reported line 123)May include surrounding context.

md
### Task 4: Delete Incompatible vectordb Index

> **⚠️ Critical:** If dimensions differ, `rm -rf vectordb/context` is required. Otherwise `EmbeddingRebuildRequiredError` on startup.

If dimension is unchanged, skip this step.

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · SKILL.md (reported line 181)May include surrounding context.

Quick verification:

bash
# 1. Server healthy
curl -s http://127.0.0.1:1933/health \
  | python3 -c "import sys,json; assert json.load(sys.stdin)['healthy']; print('OK')"

# 2. Collection dimension matches target

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
96% confidence
Finding

The command embeds a recursively destructive filesystem operation with a variable-expanded path. If ${SANDBOX_DIR} is unset, malformed, or unexpectedly broad, the deletion target can become incorrect, causing loss of unrelated data inside the sandbox or potentially beyond it depending on execution context and path composition.

Content

Scanner excerpt · references/config-reference.md (reported line 74)May include surrounding context.

md
| `vectordb/context/index/default/versions/` | Yes — index version data | Delete (will be recreated) |
| `viking/` metadata | No — user/session metadata is dimension-independent | Preserve |

**Simplest approach:** `rm -rf ${SANDBOX_DIR}/data/vectordb/context` — deletes everything under context, server recreates from scratch.

## start.sh Override Behavior

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · references/config-reference.md (reported line 89)May include surrounding context.

To check what env vars the openviking environment has:

bash
curl -s http://127.0.0.1:8090/api/v1/env-templates/openviking | python3 -m json.tool

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
88% confidence
Finding

The documented flow includes a destructive rm -rf operation against the vector database path as part of a model switch. Even though the target path is specific in the diagram, any implementation derived from this design that accepts or constructs the path from mutable inputs without strict validation could turn a maintenance action into arbitrary file deletion, and even the intended deletion causes irreversible index loss if triggered incorrectly.

Content

Scanner excerpt · references/dataflow-diagram.md (reported line 62)May include surrounding context.

md
Agent->>FS: Read ov.conf
    Agent->>FS: Write modified ov.conf<br/>(embedding.dense → llama endpoint)
    Agent->>FS: rm -rf vectordb/context<br/>(if dimension changed)
    Agent->>JEM: GET /envs/openviking<br/>(find server PID)
    Agent->>FS: kill old server PID
    Agent->>JEM: POST /envs/openviking/exec<br/>(start new server)

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · SKILL.md (reported line 92)May include surrounding context.

Fix:

bash
SANDBOX_DIR=$(curl -s http://127.0.0.1:8090/api/v1/envs/openviking \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['cwd'])")
rm -rf "${SANDBOX_DIR}/data/vectordb/context"
# Then restart the server (Step 5 in SKILL.md)

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · references/troubleshooting.md (reported line 14)May include surrounding context.

Fix:

bash
SANDBOX_DIR=$(curl -s http://127.0.0.1:8090/api/v1/envs/openviking \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['cwd'])")
rm -rf "${SANDBOX_DIR}/data/vectordb/context"
# Then restart the server (Step 5 in SKILL.md)

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · references/troubleshooting.md (reported line 81)May include surrounding context.

Fix:

bash
SANDBOX_DIR=$(curl -s http://127.0.0.1:8090/api/v1/envs/openviking \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['cwd'])")
rm -rf "${SANDBOX_DIR}/data/vectordb/context"
# Then restart the server (Step 5 in SKILL.md)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
96% confidence
Finding

rm -rf is executed on a path built from job-env-manager API output with no canonicalization, prefix enforcement, or non-empty-path checks. If the API response is malformed, compromised, or unexpected, this could delete arbitrary host files reachable by the executing user, making the guidance materially dangerous.

Content

Scanner excerpt · references/troubleshooting.md (reported line 16)May include surrounding context.

bash
SANDBOX_DIR=$(curl -s http://127.0.0.1:8090/api/v1/envs/openviking \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['cwd'])")
rm -rf "${SANDBOX_DIR}/data/vectordb/context"
# Then restart the server (Step 5 in SKILL.md)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

The command removes a PID file using a path derived from local API output without validating the directory, which can target unintended files if the path is wrong or manipulated. The impact is lower than recursive deletion but still dangerous because it removes process-coordination artifacts based on untrusted input.

Content

Scanner excerpt · references/troubleshooting.md (reported line 83)May include surrounding context.

bash
SANDBOX_DIR=$(curl -s http://127.0.0.1:8090/api/v1/envs/openviking \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['cwd'])")
rm -f "${SANDBOX_DIR}/data/.openviking.pid"
find "${SANDBOX_DIR}/data/vectordb" -name "LOCK" -delete

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · references/troubleshooting.md (reported line 115)May include surrounding context.

Fix:

bash
# Check llama env state
curl -s http://127.0.0.1:8090/api/v1/envs/llama | python3 -c "import sys,json; print(json.load(sys.stdin)['state'])"
# Start if needed
curl -s -X POST http://127.0.0.1:8090/api/v1/envs/llama/start
# Verify --embeddings flag

Privileged Container / Container Escape

High
Category
Privilege Escalation
Confidence
90% confidence
Finding

Potential security issue detected. Manual review is recommended.

Content

Scanner excerpt · references/troubleshooting.md (reported line 124)May include surrounding context.

md
---

## Problem 8: nsenter Fails with "Operation not permitted"

**Fix:** Do not use `nsenter`. Use the job-env-manager `exec` API instead.

Privileged Container / Container Escape

High
Category
Privilege Escalation
Confidence
90% confidence
Finding

Potential security issue detected. Manual review is recommended.

Content

Scanner excerpt · references/troubleshooting.md (reported line 126)May include surrounding context.

md
---

## Problem 8: nsenter Fails with "Operation not permitted"

**Fix:** Do not use `nsenter`. Use the job-env-manager `exec` API instead.

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · scripts/switch-embedding-model.sh (reported line 32)May include surrounding context.

sh
# ── Step 1: Detect sandbox directory ──
echo "[1/6] Detecting OpenViking sandbox..."
ENVS_RESP=$(curl -s "${JEM_BASE}/envs/openviking")
OV_STATE=$(echo "$ENVS_RESP" | python3 -c "import sys,json; print(json.load(sys.stdin)['state'])")
SANDBOX_DIR=$(echo "$ENVS_RESP" | python3 -c "import sys,json; print(json.load(sys.stdin)['cwd'])")

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
97% confidence
Finding

A recursive rm -rf is executed on a path built from SANDBOX_DIR, which originates from API data, without canonicalization or strict path validation. If the API response is wrong, compromised, or unexpected, this could delete arbitrary directories and cause severe data loss beyond the intended vectordb index.

Content

Scanner excerpt · scripts/switch-embedding-model.sh (reported line 114)May include surrounding context.

sh
if [ "$CURRENT_DIM" != "$TARGET_DIMENSION" ]; then
    echo "  Dimension mismatch: $CURRENT_DIM → $TARGET_DIMENSION"
    echo "  Deleting vectordb/context..."
    rm -rf "${SANDBOX_DIR}/data/vectordb/context"
    echo "  Deleted."
  else
    echo "  Dimensions match ($CURRENT_DIM), no deletion needed."

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

The script deletes a PID file under a path derived from SANDBOX_DIR without validating that the path belongs to the expected sandbox. While less destructive than rm -rf, it still mutates filesystem state based on potentially untrusted path input and could interfere with unrelated services if misdirected.

Content

Scanner excerpt · scripts/switch-embedding-model.sh (reported line 149)May include surrounding context.

sh
fi

# BUG-3 fix: clean up stale lock files
rm -f "${SANDBOX_DIR}/data/.openviking.pid" 2>/dev/null
find "${SANDBOX_DIR}/data/vectordb" -name "LOCK" -delete 2>/dev/null
echo "  Cleaned up lock files"

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

The cleanup step removes the backup config file using a path derived from SANDBOX_DIR, again without path validation. The direct impact is smaller than other deletions, but it still reflects unsafe file-operation patterns that could remove unintended files if the sandbox path is manipulated.

Content

Scanner excerpt · scripts/switch-embedding-model.sh (reported line 225)May include surrounding context.

sh
fi

# Cleanup backup
rm -f "${SANDBOX_DIR}/process_dir/ov.conf.bak"

echo ""
echo "=== Switch Complete ==="

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill clearly instructs file modification and shell/process control, but it declares no explicit tool scope such as permissions or allowed-tools. That mismatch weakens policy enforcement and increases the chance an agent will execute high-risk operations like config edits, process restarts, and destructive deletion without least-privilege constraints.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 101)May include surrounding context.

Task 2: Validate Target Embedding Endpoint

bash
curl -s http://127.0.0.1:${LLAMA_PORT}/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"model":"${MODEL_NAME}","input":"test"}' \
  | python3 -c "import sys,json; d=json.load(sys.stdin); print(len(d['data'][0]['embedding']))"

Session Persistence

Medium
Category
Rogue Agent
Confidence
76% confidence
Finding

Using nohup to launch a background service creates persistent state beyond the immediate session and can make process ownership, auditing, and cleanup harder. In a skill context, this is riskier because an agent can leave long-lived processes running even after task completion or failure.

Content

Scanner excerpt · SKILL.md (reported line 140)May include surrounding context.

bash
curl -s --max-time 15 -X POST http://127.0.0.1:8090/api/v1/envs/openviking/exec \
  -H 'Content-Type: application/json' \
  -d '{"cmd":["bash","-c","nohup /root/runtime/openviking/venv/bin/openviking-server --config /workspace/process_dir/ov.conf > /workspace/process_dir/openviking-server.log 2>&1 & sleep 2 && echo started"]}'

Task 6: Verify

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list includes broad English and Chinese phrases for changing an embedding model, which can cause the skill to activate in situations where a user is only discussing configuration rather than explicitly requesting a destructive switch. In this skill's context, activation can modify ov.conf, delete an incompatible vector index, and restart a server, so overbroad matching increases the chance of unintended disruptive actions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The documentation recommends a destructive rm -rf operation that deletes the entire vectordb context, but it does not prominently warn about irreversible data loss, backup requirements, or validation of the target path. In an operational skill that edits live sandbox state, this increases the risk of accidental deletion and service-impacting mistakes, especially if users or automation copy the command without safeguards.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.