Back to skill

Security audit

vietnamese-contract

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a Vietnamese contract generator, but it handles national ID data and can send extracted identity text to any active AI model without clear consent or provider controls.

Review before installing. Use the skill only with explicit consent from the ID holder, avoid remote AI verification for CCCD/CMND text unless the provider and retention policy are acceptable, install dependencies in an isolated pinned environment, and manually delete uploaded ID images and OCR outputs after use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:102
Finding

Unpinned Global and System-Level Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:102-106, SKILL.md:214-217, references/docx-formatting.md:8, references/cccd-ocr-guide.md:24-30, scripts/cccd-ocr.py:23-35
Vulnerability Type: Supply-chain exposure through unpinned dependencies and unsafe package installation
Risk Level: Medium

Vulnerable Code

SKILL.md:102-106:

bash
## Bước 1: Cài đặt môi trường

# Cài docx-js (thư viện tạo file .docx)
npm install -g docx

SKILL.md:214-217:

bash
### Cài đặt (1 lần)

pip install easyocr opencv-python --break-system-packages

references/cccd-ocr-guide.md:24-30:

bash
## Cài đặt

pip install easyocr opencv-python --break-system-packages

Lần đầu chạy sẽ tự tải model tiếng Việt (~100MB).

scripts/cccd-ocr.py:23-35:

python
def check_deps():
    missing = []
    try:
        import easyocr
    except ImportError:
        missing.append("easyocr")
    try:
        import cv2
    except ImportError:
        missing.append("opencv-python")
    if missing:
        print("Thieu thu vien. Chay lenh:")
        print(f"   pip install {' '.join(missing)} --break-system-packages")

Technical Analysis

The documented installation commands retrieve the latest available versions of docx, easyocr, and opencv-python without version constraints, lockfiles, package hashes, or artifact verification. The npm package is installed globally, expanding the scope of filesystem changes beyond this project. The Python command uses --break-system-packages, bypassing protections intended to prevent pip from modifying an externally managed Python environment.

EasyOCR initialization also downloads a language model during first use. The project does not pin or verify the model version, source, checksum, or signature. Consequently, the effective runtime components can change after the skill package itself has been audited.

This is a supply-c ...[truncated 1624 chars]

Remediation
View remediation

Remediation Suggestions

  1. Pin exact reviewed package versions rather than resolving the latest versions:
    bash
    npm install --save-exact docx@REVIEWED_VERSION
    python -m pip install easyocr==REVIEWED_VERSION opencv-python==REVIEWED_VERSION
    
  2. Add and commit appropriate lockfiles, such as package-lock.json and a hash-pinned Python requirements file.
  3. Require integrity verification with Python package hashes and verified npm lockfile integrity metadata.
  4. Install dependencies inside a dedicated virtual environment or disposable container. Remove --break-system-packages.
  5. Use a project-local npm dependency instead of npm install -g.
  6. Disable automatic model downloads in production where possible. Pre-fetch a reviewed model from an approved source and verify its cryptographic checksum before use.
  7. Run package installation and OCR processing under a restricted, non-privileged account with limited filesystem and network access.
  8. Add dependency scanning and periodic review of direct and transitive dependencies.

other

Warning
Location
SKILL.md:228
Finding

Disclosure of Raw National Identity Data to an Arbitrary AI Provider

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:228-264, references/cccd-ocr-guide.md:63-83, references/cccd-ocr-guide.md:126-131, references/cccd-ocr-guide.md:154-156
Vulnerability Type: Sensitive personal data disclosure without mandatory consent, minimization, or provider controls
Risk Level: Medium

Vulnerable Instructions

SKILL.md:228-238:

text
**Bước 2: AI nhận kết quả OCR và cải thiện**

Agent (bất kỳ model nào) nhận JSON từ bước 1, rồi:
- Sửa lỗi OCR: "NGUYÊN" → "NGUYỄN", "TRÍÊT" → "TRIẾT"
- Chuẩn hóa dấu tiếng Việt: "Pham Minh Triet" → "Phạm Minh Triết"
- Kiểm tra logic: ngày sinh hợp lệ? CCCD đủ 12 số? địa chỉ hợp lý?
- Format chuẩn: ngày DD/MM/YYYY, tên IN HOA

**Bước 3: Hiển thị → xác nhận → điền hợp đồng**

SKILL.md:259-264:

text
### Bảo mật

- EasyOCR chạy OFFLINE — ảnh CCCD không gửi ra internet
- Chỉ raw text (không phải ảnh) được gửi cho AI model để verify
- KHÔNG lưu ảnh CCCD sau khi trích xuất
- Xóa dữ liệu trung gian sau khi tạo xong hợp đồng

references/cccd-ocr-guide.md:154-156:

text
- **Lớp 1 (EasyOCR)**: Chạy OFFLINE — ảnh CCCD không gửi ra internet
- **Lớp 2 (AI verify)**: Chỉ gửi TEXT (không phải ảnh) cho AI model
- KHÔNG lưu ảnh CCCD sau khi trích xuất

Technical Analysis

The workflow directs the agent to provide raw OCR JSON to any AI model currently configured in OpenClaw. That JSON can contain a national identity number, full name, date of birth, sex, nationality, hometown, residential address, issue date, issuing authority, and expiration date.

Converting an identity-card image into text does not anonymize the information. The resulting text remains highly sensitive and may uniquely identify the cardholder. The workflow does not require explicit informed consent before external processing, does not distinguish local from remote models, does not allowlist approved providers, and does not require confirmation of ...[truncated 1683 chars]

Remediation
View remediation

Remediation Suggestions

  1. Require explicit, informed opt-in before sending any OCR-derived identity data to a remote AI provider. Explain which fields will be transferred and identify the provider.
  2. Default to local OCR and local validation. Do not treat an available remote model as automatic authorization to process identity data.
  3. Determine whether the configured model is local or remote before transfer. Block remote transfer unless the provider is explicitly approved.
  4. Minimize data sent for correction. For example, validate date syntax locally and mask identity numbers except for the digits strictly required for validation.
  5. Never send unrelated fields together. A name correction should not require transmitting the identity number, full address, and issue information.
  6. Add configurable provider allowlisting and require acceptable retention, training, logging, and data-location policies.
  7. Ask the user to confirm OCR output locally before any optional AI-assisted correction, rather than only after transfer.
  8. Implement deletion controls in code for uploaded images and temporary OCR artifacts, with cleanup in exception and cancellation paths.
  9. Clearly state that textual OCR output remains sensitive personal information; do not imply that excluding the original image makes remote transfer inherently safe.
  10. Avoid placing complete identity records in application logs, command output retained by orchestration systems, or model conversation history.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is presented as a contract-drafting tool but also includes OCR processing of CCCD/CMND identity documents and extraction of sensitive personal data. This capability expansion is dangerous because users and calling systems may invoke it under a benign legal-drafting expectation while it actually processes high-risk PII, increasing privacy, consent, and data-handling exposure.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill is presented as a contract-drafting tool but also includes OCR processing of CCCD/CMND identity documents and extraction of sensitive personal data. This capability expansion is dangerous because users and calling systems may invoke it under a benign legal-drafting expectation while it actually processes high-risk PII, increasing privacy, consent, and data-handling exposure.

Content

No source excerpt is available for this finding.

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill declares that it should always be used for virtually any mention of contracts or legal documents, creating an overly broad trigger surface. Broad auto-invocation is dangerous because it can activate shell use, web lookups, document generation, and even identity-document handling in situations where the user did not request those actions or where a safer, narrower skill would be more appropriate.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The script performs OCR on Vietnamese national ID cards, which is materially outside the skill’s declared contract/legal-document drafting purpose. This scope mismatch is dangerous because it introduces an unexpected identity-document processing capability that can collect highly sensitive personal data users would not reasonably expect from a contract-drafting skill.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The code extracts sensitive identity attributes including ID number, full name, date of birth, sex, nationality, place of origin, residence, and issuance details. In the context of a contract-drafting skill, this unjustified collection of government-ID data increases the risk of privacy violations, identity theft, and misuse of personal information, especially if users are induced to upload ID images for an unrelated task.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill instructs the agent to run shell commands such as npm, pip, python, pandoc, cp, and which, but it declares no explicit tool scope or permissions boundary. In an agent environment, that creates unnecessary command-execution authority and makes it easier for later prompt injection, path misuse, or unsafe file handling to escalate into real system actions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill description specifies generating Vietnamese legal documents and later mandates Vietnamese language and formatting conventions, but it does not indicate that the user can opt into another language or that the locale restriction is a documented, justified limitation. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The file introduces a full OCR-based national ID extraction workflow for CCCD/CMND data, which expands the skill from contract drafting into collection and processing of highly sensitive identity data. That broader capability increases privacy, data-handling, and misuse risk because the skill now guides users to ingest identity documents, normalize the data, and reuse it in generated documents.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The guide instructs the agent to send OCR-extracted CCCD text—including full name, ID number, birth date, and address—to an AI model without a clear, prior warning that this sensitive personal data may be transmitted off-device. Because CCCD data is highly identifying, disclosure to third-party model providers can create significant privacy, compliance, and identity-theft risk if users are not informed and consenting.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The security section states that CCCD images are not sent to the internet, but elsewhere permits AI verification by arbitrary models, which may be remote services receiving OCR-extracted identity text. This creates a misleading security assurance that could cause users to disclose sensitive personal information under false assumptions about locality and confidentiality.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file contains strong natural-language requirements such as 'NGUYÊN TẮC SỐ 1' and 'BẮT BUỘC' for Vietnamese spelling and direct Unicode usage. Under SQP-3, forcing a specific language or locale without user opt-in is a policy concern unless clearly presented as an optional or region-specific constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script processes highly sensitive ID-card images but provides no privacy warning, consent flow, retention guidance, or data-handling instructions. That omission is dangerous because users may supply identity documents without understanding the sensitivity of the operation or how their personal data should be protected, stored, or deleted.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The module docstring states the tool is for Vietnamese spellchecking and presents all user-facing usage text in Vietnamese, which effectively constrains operation to a specific language. The file does not offer user opt-in, alternative locale support, or an explicit policy justification that this locale restriction is required.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/vn-spellcheck.py (reported line 51)May include surrounding context.

python
def extract_text(docx_path):
    """Xuất text từ .docx bằng pandoc"""
    try:
        result = subprocess.run(
            ["pandoc", docx_path, "-t", "plain", "--wrap=none"],
            capture_output=True, text=True, timeout=30
        )

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/vn-spellcheck.py (reported line 56)May include surrounding context.

python
capture_output=True, text=True, timeout=30
        )
        if result.returncode != 0:
            result = subprocess.run(
                ["pandoc", docx_path, "-o", "-"],
                capture_output=True, text=True, timeout=30
            )

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The entire skill reference is written only in Vietnamese and is tightly coupled to Vietnamese legal formatting and Vietnamese law references, without offering any language or locale choice. Under the stated policy, forcing a specific language/locale without user opt-in is a natural-language policy issue unless the locale constraint is explicitly documented and justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The checklist explicitly instructs copying generated output into /mnt/user-data/outputs/, which is a persistent user-data area, without any notice, consent step, or limitation guidance. In a skill that generates legal documents, this can cause unintended retention of sensitive personal or contractual information and increases privacy risk if users do not expect files to be written there.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

This file contains user-facing natural-language guidance exclusively in Vietnamese, and there is no indication that users can choose another language or that the file is explicitly limited to a Vietnam-only audience. Under the language/locale policy rule, forcing a specific language without opt-in can be a policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

All natural-language instructions and output examples are presented only in Vietnamese, and the script description does not indicate that this language restriction is optional or region-justified. Under the policy, forcing a specific language without opt-in can be a locale-policy violation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.