Back to skill

Security audit

organization_code_certificate_ocr

Security checks for vulnerabilities and agentic risk

Overview

This OCR skill appears purpose-aligned, but it should be reviewed because it uploads potentially sensitive certificate files and an API token to a configurable remote endpoint without enforcing a trusted HTTPS host.

Review before installing. Use it only for documents you are allowed to send to Scnet, keep config/.env owner-readable only, and do not set SCNET_API_BASE unless it is a trusted HTTPS endpoint you intend to receive both the document and API token.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/main.py:86
Finding

Arbitrary OCR Endpoint Can Receive API Credentials and Sensitive Documents

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (19)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 20)May include surrounding context.

python
# 获取技能根目录(脚本所在目录的上一级)
SKILL_ROOT = Path(__file__).parent.parent.absolute()
ENV_FILE = SKILL_ROOT / "config" / ".env"

# --- 新增:重试配置 ---
MAX_RETRIES = 3            # 最大重试次数

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 62)May include surrounding context.

md
# --------------------

def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 124)May include surrounding context.

md
# --------------------

def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 29)May include surrounding context.

python
# --------------------

def load_config():
    """从 .env 文件加载配置,若文件不存在则抛出友好错误"""
    if not ENV_FILE.exists():
        error_msg = (
            "\n===============================================\n"

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · README.md (reported line 38)May include surrounding context.

md
* [Get started with GitLab CI/CD](https://docs.gitlab.com/ee/ci/quick_start/)
* [Analyze your code for known vulnerabilities with Static Application Security Testing (SAST)](https://docs.gitlab.com/ee/user/application_security/sast/)
* [Deploy to Kubernetes, Amazon EC2, or Amazon ECS using Auto Deploy](https://docs.gitlab.com/ee/topics/autodevops/requirements.html)
* [Use pull-based deployments for improved Kubernetes management](https://docs.gitlab.com/ee/user/clusters/agent/)
* [Set up protected environments](https://docs.gitlab.com/ee/ci/environments/protected_environments.html)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill declares capabilities that imply local file access, network access, and shell execution, but it does not define any explicit tool scope such as permissions or allowed-tools. This weakens containment and reviewability because an agent may invoke broader capabilities than users expect when processing local documents and sending them to an external OCR service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

description 使用中文描述技能能力,且全文面向中文交互,没有说明是否支持其他语言或允许用户选择语言。根据策略,若技能在自然语言层面默认强制特定语言且无用户 opt-in,可能构成语言/地区政策问题。

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

The skill is explicitly designed to send user-provided document images and OCR requests to an external endpoint at api.scnet.cn. External transmission is expected for cloud OCR, but it remains a real security concern because local document contents may include sensitive organizational identifiers and are exported to a third party.

Content

Scanner excerpt · SKILL.md (reported line 54)May include surrounding context.

SCNET_API_KEY=your_scnet_api_key_here

API 基础地址(一般无需修改)

SCNET_API_BASE=https://api.scnet.cn/api/llm/v1

text
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The auto-trigger guidance is broad and says the AI may automatically invoke the skill based on description keywords, without clear boundaries or exclusion rules. In context, this is risky because the skill reads a local file path and transmits document contents to a third-party OCR API, so ambiguous triggering can cause unintended exfiltration of sensitive documents.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

This configuration section again confirms that requests are sent to an external SCNET API base URL. Because the skill handles certificate images and structured OCR output, unintended or poorly understood outbound transfer can expose sensitive business document data outside the local environment.

Content

Scanner excerpt · SKILL.md (reported line 106)May include surrounding context.

md
| 变量名 | 默认值 | 说明 |
|--------|--------|------|
| SCNET_API_KEY | 必需 | Scnet API 密钥 |
| SCNET_API_BASE | https://api.scnet.cn/api/llm/v1 | API 基础地址(一般无需修改) |

### 输出

External Transmission

Medium
Category
Data Exfiltration
Confidence
90% confidence
Finding

The file documents direct transmission of uploaded documents to an external domain, which is a real external data egress path. In the context of OCR for organization code certificates, this increases risk because certificates may include company identifiers, addresses, legal representative names, and other sensitive records that are exposed to a third-party processor.

Content

Scanner excerpt · references/api-docs.md (reported line 4)May include surrounding context.

md
# Sugon-Scnet OCR API 文档摘要

## 接口地址
`POST https://api.scnet.cn/api/llm/v1/ocr/recognize`

## 请求头
- `Content-Type: multipart/form-data`

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documentation instructs callers to upload files containing document images to a third-party OCR endpoint and authenticate with a bearer token, but it provides no warning that user-provided data leaves the local environment. Because this skill processes organization code certificates, the transmitted files and extracted fields can contain sensitive business and personal information, creating a privacy and compliance risk if users are not clearly informed.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 45)May include surrounding context.

python
"   b) 配置文件:\n"
            f"      mkdir -p {SKILL_ROOT}/config\n"
            f"      echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
            f"      chmod 600 {ENV_FILE}\n"
            "\n配置完成后重新运行。"
        )
        sys.exit(error_msg)

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 152)May include surrounding context.

python
"   b) 配置文件:\n"
            f"      mkdir -p {SKILL_ROOT}/config\n"
            f"      echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
            f"      chmod 600 {ENV_FILE}\n"
            "\n配置完成后重新运行。"
        )
        sys.exit(error_msg)

External Transmission

Medium
Category
Data Exfiltration
Confidence
92% confidence
Finding

This skill transmits user-supplied document files and an API bearer token to an external service at api.scnet.cn. External transmission is expected for an OCR skill, but it is still security-relevant because potentially sensitive certificate data leaves the local environment and could be exposed if the endpoint, configuration, or data handling is not tightly controlled.

Content

Scanner excerpt · scripts/main.py (reported line 78)May include surrounding context.

python
)
        sys.exit(error_msg)

    config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')
    return config

def recognize_with_retry(ocr_type, file_path, config, retry_count=0):

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest description is written only in Chinese ('支持组织机构代码证的票据识别。'), which indicates a fixed language/locale presentation without any stated opt-in or alternative. Under the policy, locale or language constraints should be optional or clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
96% confidence
Finding

The changelog content is entirely in Chinese, including the title and feature description, with no indication that language selection is optional or that the skill is specifically intended for a Chinese-only audience. This can violate language/locale policy when a skill implicitly enforces a locale without user opt-in.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

SKILL.md 在 L130 明确声称技能会在 429 限流时自动等待并最多重试 3 次,但本次提供的技能文件只有清单与使用文档,没有可验证的实现代码或重试逻辑。该表述会让使用者误以为技能具备自动恢复能力,属于文档对实际实现的积极性描述与可见实现不一致。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The module docstring presents the skill interface and output description entirely in Chinese, and later user-facing messages and usage text are also only in Chinese. Under the policy, forcing a specific language without opt-in can be a natural-language locale violation unless the locale restriction is explicitly justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.