Back to skill

Security audit

trademark_certificate_ocr

Security checks for vulnerabilities and agentic risk

Overview

This OCR skill matches its stated purpose, but it needs Review because it can upload any readable local file to a configurable external API endpoint.

Use this only in a controlled environment with documents approved for third-party OCR. Before running it, verify `config/.env` points to the intended HTTPS Scnet API endpoint, avoid letting an agent choose arbitrary file paths, protect the API key, and consider pinning dependencies.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/main.py:78
Finding

Unrestricted API destination may disclose credentials and uploaded documents

Content
View full analysis

Vulnerability Details

File Location: scripts/main.py:78-112
Vulnerability Type: Unvalidated outbound network destination
Risk Level: Medium

python
config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')

def recognize_with_retry(ocr_type, file_path, config, retry_count=0):
    api_base = config['SCNET_API_BASE']
    api_key = config['SCNET_API_KEY']
    url = f"{api_base}/ocr/recognize"

    mime_type, _ = mimetypes.guess_type(file_path)
    if mime_type is None:
        mime_type = 'application/octet-stream'

    headers = {
        'Authorization': f'Bearer {api_key}'
    }

    try:
        with open(file_path, 'rb') as f:
            files = {
                'file': (os.path.basename(file_path), f, mime_type)
            }
            data = {
                'ocrType': ocr_type,
                'channelTag': "scnetSkills"
            }
            response = requests.post(
                url,
                headers=headers,
                data=data,
                files=files,
                timeout=60
            )

Technical Analysis

The optional SCNET_API_BASE setting is used directly to construct the request URL without validating its scheme, hostname, port, or trust relationship. The request carries both the Scnet bearer credential and the complete contents of the selected document.

Although sending a document to the default Scnet endpoint is necessary for the declared remote OCR functionality, allowing an unrestricted destination exceeds the minimum network privileges required for that functionality. A modified or incorrectly supplied configuration can redirect requests to an unrelated host. The code also does not require HTTPS, so a configured HTTP endpoint would expose the credential and document to interception.

Attack Path

  1. An attacker, compromised deployment mechanism, or unsafe configuration process alters `conf ...[truncated 1098 chars]
Remediation
View remediation

Remediation Suggestions

  • Use a fixed, allowlisted Scnet HTTPS origin for production credentials.
  • Parse the endpoint with a standard URL parser and require the https scheme.
  • Validate the normalized hostname against an explicit allowlist such as api.scnet.cn.
  • Reject URLs containing embedded user information, fragments, unexpected ports, or ambiguous hostnames.
  • Do not send a production Scnet credential to a custom endpoint.
  • If custom endpoints are a required advanced feature, use a separate endpoint-specific credential and require explicit user confirmation before transmitting a document.
  • Log the normalized destination before transmission without logging the API key or document contents.
  • Add tests covering HTTP URLs, lookalike domains, embedded credentials, malformed URLs, and attacker-controlled hosts.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/main.py:91
Finding

Arbitrary readable files can be uploaded without content or path restrictions

Content
View full analysis

Vulnerability Details

File Location: scripts/main.py:91-112
Vulnerability Type: Unrestricted local file upload
Risk Level: Medium

python
if not os.path.isfile(file_path):
    sys.exit(f"Error: file does not exist - {file_path}")

mime_type, _ = mimetypes.guess_type(file_path)
if mime_type is None:
    mime_type = 'application/octet-stream'

headers = {
    'Authorization': f'Bearer {api_key}'
}

try:
    with open(file_path, 'rb') as f:
        files = {
            'file': (os.path.basename(file_path), f, mime_type)
        }
        data = {
            'ocrType': ocr_type,
            'channelTag': "scnetSkills"
        }
        response = requests.post(
            url,
            headers=headers,
            data=data,
            files=files,
            timeout=60
        )

Technical Analysis

The script accepts any path for which os.path.isfile() returns true. It does not constrain the path to an approved input directory, reject symbolic links, enforce a maximum file size, verify the file signature, or allowlist the documented image and PDF formats.

If MIME inference fails, the file is deliberately accepted as application/octet-stream. Consequently, any readable local file can be transmitted to the configured OCR endpoint, including configuration files, private keys, application secrets, or unrelated personal documents.

The ocrType argument is also not validated against the sole documented value, TRADEMARK_REGISTRATION_CERT. While this does not by itself create arbitrary file access, it demonstrates that command-line inputs are trusted rather than constrained to the declared function.

Attack Path

  1. An attacker influences an Agent prompt, workflow argument, or command invocation.
  2. The attacker supplies the path of a sensitive file instead of a trademark image or PDF.
  3. The Skill checks only whether the path is an existing file. 4 ...[truncated 763 chars]
Remediation
View remediation

Remediation Suggestions

  • Validate ocr_type against an explicit allowlist containing only supported OCR modes.
  • Restrict uploads to documented formats such as JPEG, PNG, and PDF.
  • Verify file signatures rather than trusting filename extensions or MIME inference alone.
  • Reject unknown types instead of falling back to application/octet-stream.
  • Canonicalize the path with Path.resolve() and enforce an approved user-selected input directory or workspace boundary.
  • Consider rejecting symbolic links to prevent an approved-looking path from resolving to a sensitive file.
  • Enforce conservative file-size and page-count limits before opening or uploading content.
  • Require explicit user confirmation when a file lies outside the active workspace or approved upload directory.
  • Run the Skill under a dedicated, minimally privileged account that cannot read unrelated secrets.
  • Add tests for credential files, symbolic links, unsupported extensions, spoofed signatures, oversized files, and paths outside the approved directory.

T08 · Insecure Dependencies

Note
Location
SKILL.md:65
Finding

Runtime dependency is installed without version or integrity pinning

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:65-70
Vulnerability Type: Unpinned third-party dependency
Risk Level: Low

bash
pip install requests

Technical Analysis

The installation instructions retrieve the latest resolver-selected release of requests and its transitive dependencies without a lock file, exact version constraints, or cryptographic hashes. The installed code can therefore differ between deployments even when the Skill package itself has not changed.

This creates a supply-chain and reproducibility weakness. A future compromised package release, compromised package index, or unexpected dependency change could introduce malicious behavior or compatibility failures outside the audited source tree.

No evidence was found that the current requests package is malicious, that the project uses a typosquatted package, or that it downloads dependencies from an untrusted custom index. The finding concerns the unsafe dependency installation practice.

Attack Path

  1. A user follows the documented dependency installation command.
  2. pip resolves mutable package versions from its configured package index.
  3. A compromised or maliciously replaced eligible package release is selected.
  4. Package installation or later import executes the compromised dependency code with the user's privileges.
  5. That code can access data and resources available to the Python process.

Impact Assessment

If the dependency supply chain were compromised, malicious package code could execute with the privileges of the user performing installation or running the Skill. The potential scope includes access to local files, environment data, network resources, and the OCR credential available to the process.

Exploitation depends on a separate compromise of the package source or an unsafe package-index configuration; the repository itself does not contain an observed malicious dependency.

Remediation
View remediation

Remediation Suggestions

  • Pin requests and all transitive dependencies to reviewed versions in a requirements or lock file.
  • Generate and verify cryptographic hashes for every distribution.
  • Install with hash enforcement, for example through pip install --require-hashes.
  • Use an approved package index over HTTPS and disable unexpected supplemental indexes.
  • Regularly scan pinned dependencies for known vulnerabilities and update them through a reviewed process.
  • Document the supported Python version and reproduce dependency resolution in continuous integration.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (18)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 20)May include surrounding context.

python
# 获取技能根目录(脚本所在目录的上一级)
SKILL_ROOT = Path(__file__).parent.parent.absolute()
ENV_FILE = SKILL_ROOT / "config" / ".env"

# --- 新增:重试配置 ---
MAX_RETRIES = 3            # 最大重试次数

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 62)May include surrounding context.

md
# --------------------

def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 124)May include surrounding context.

md
# --------------------

def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 29)May include surrounding context.

python
# --------------------

def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · README.md (reported line 38)May include surrounding context.

md
* [Get started with GitLab CI/CD](https://docs.gitlab.com/ee/ci/quick_start/)
* [Analyze your code for known vulnerabilities with Static Application Security Testing (SAST)](https://docs.gitlab.com/ee/user/application_security/sast/)
* [Deploy to Kubernetes, Amazon EC2, or Amazon ECS using Auto Deploy](https://docs.gitlab.com/ee/topics/autodevops/requirements.html)
* [Use pull-based deployments for improved Kubernetes management](https://docs.gitlab.com/ee/user/clusters/agent/)
* [Set up protected environments](https://docs.gitlab.com/ee/ci/environments/protected_environments.html)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill documentation declares no explicit tool scope or permission boundaries even though its described operation requires reading local files, making network requests to a third-party API, and invoking Python from the shell. Missing tool restrictions increases the chance an agent can execute the skill with broader-than-necessary capabilities and without clear operator awareness of what resources will be accessed.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill explains how to pass a local image path for OCR but does not clearly warn that the image contents will be uploaded to a third-party service for processing. This omission can cause users to unknowingly transmit sensitive documents, creating privacy, confidentiality, and compliance risks when handling trademarks or other attached materials.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 54)May include surrounding context.

SCNET_API_KEY=your_scnet_api_key_here

API 基础地址(一般无需修改)

SCNET_API_BASE=https://api.scnet.cn/api/llm/v1

text
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 106)May include surrounding context.

SCNET_API_KEY=your_scnet_api_key_here

API 基础地址(一般无需修改)

SCNET_API_BASE=https://api.scnet.cn/api/llm/v1

text
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)

External Transmission

Medium
Category
Data Exfiltration
Confidence
91% confidence
Finding

The skill documentation directs users to send files to an external OCR endpoint, which means uploaded trademark certificates and related sensitive document contents are transmitted to a third-party service outside the local trust boundary. In this skill context, that is materially relevant because the documents contain potentially sensitive business and personal data, so use of the external API creates confidentiality, compliance, and data handling risk if not explicitly disclosed and controlled.

Content

Scanner excerpt · references/api-docs.md (reported line 4)May include surrounding context.

md
# Sugon-Scnet OCR API 文档摘要

## 接口地址
`POST https://api.scnet.cn/api/llm/v1/ocr/recognize`

## 请求头
- `Content-Type: multipart/form-data`

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The top-level docstring defines the skill entirely in Chinese and specifies its command-line interface and JSON output in that locale, with the rest of the user-facing messages also hardcoded in Chinese. This creates a language/locale constraint without offering user opt-in or documenting why the skill must be Chinese-only.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 45)May include surrounding context.

python
"   b) 配置文件:\n"
            f"      mkdir -p {SKILL_ROOT}/config\n"
            f"      echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
            f"      chmod 600 {ENV_FILE}\n"
            "\n配置完成后重新运行。"
        )
        sys.exit(error_msg)

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 152)May include surrounding context.

python
"   b) 配置文件:\n"
            f"      mkdir -p {SKILL_ROOT}/config\n"
            f"      echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
            f"      chmod 600 {ENV_FILE}\n"
            "\n配置完成后重新运行。"
        )
        sys.exit(error_msg)

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

This skill uploads the provided document to an external OCR service at api.scnet.cn. Because trademark registration certificates may contain sensitive or regulated data, transmitting files off-host creates a real confidentiality and data-governance risk if users are not explicitly informed, if consent is not obtained, or if the endpoint can be overridden to an untrusted host via configuration.

Content

Scanner excerpt · scripts/main.py (reported line 78)May include surrounding context.

python
)
        sys.exit(error_msg)

    config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')
    return config

def recognize_with_retry(ocr_type, file_path, config, retry_count=0):

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The description is written as a Chinese-only capability statement ('支持商标注册证的票据识别。') with no indication that users can choose another language or that the skill is intentionally limited to a Chinese-language or China-specific workflow. Under the policy, language or locale constraints should be opt-in or clearly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This markdown file contains user-facing natural language entirely in Chinese (e.g. the title and feature description) with no indication that the skill is region-specific or that users can choose the language. Under the language/locale policy, forcing a specific language without opt-in can be a policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

This markdown file contains user-facing natural-language documentation entirely in Chinese. Under the language/locale policy rule, forcing a specific language without user opt-in can be a policy concern when no justification or alternative is provided.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

This markdown file is natural-language documentation, and it presents all instructions, headings, and notes exclusively in Chinese. Under the policy rule for language/locale, a skill artifact that effectively forces a specific language without opt-in can be considered a policy violation unless the regional constraint is clearly documented and justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.