Back to skill

Security audit

Fixed Amount Invoice Ocr

Security checks for vulnerabilities and agentic risk

Overview

This OCR skill matches its stated purpose, but it can upload any readable local file and the API credential to a configurable remote endpoint without strong scoping or validation.

Install only if you are comfortable sending invoice files to SCNet's OCR service. Do not use it on unrelated local files or highly sensitive documents unless your organization approves that data flow, and keep `SCNET_API_BASE` fixed to the trusted SCNet HTTPS endpoint.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/main.py:82
Finding

Unrestricted local-file upload to a configurable network endpoint

Content
View full analysis

Vulnerability Details

File Location: scripts/main.py:50-55, 82-108
Vulnerability Type: Unrestricted sensitive-file transmission and credential exposure
Risk Level: High

Vulnerable Code

python
env_api_base = os.environ.get('SCNET_API_BASE')
if env_api_base:
    config['SCNET_API_BASE'] = env_api_base

# 3. 设置默认值
config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')
python
api_base = config['SCNET_API_BASE']
api_key = config['SCNET_API_KEY']
url = f"{api_base}/ocr/recognize"

# 检查文件是否存在
if not os.path.isfile(file_path):
    sys.exit(f"错误: 文件不存在 - {file_path}")

# 自动检测 MIME 类型
mime_type, _ = mimetypes.guess_type(file_path)
if mime_type is None:
    mime_type = 'application/octet-stream'

headers = {
    'Authorization': f'Bearer {api_key}'
}

try:
    with open(file_path, 'rb') as f:
        files = {
            'file': (os.path.basename(file_path), f, mime_type)
        }
        data = {
            'ocrType': ocr_type,
            'channelTag': "scnetSkills"
        }
        response = requests.post(url, headers=headers, data=data, files=files, timeout=60)

Technical Analysis

The Skill accepts a command-line file path and verifies only that it refers to a file. It does not enforce the documented image or PDF formats, inspect file content, impose a size limit, restrict the file to an approved directory, or reject known sensitive paths. Consequently, any readable local file can be submitted to the OCR endpoint.

The destination is constructed from SCNET_API_BASE, which can be overridden through the process environment or the local .env configuration. There is no URL-scheme validation, hostname allowlist, or restriction to the documented SCNet service. The request sends both the selected file and the SCNet API credential in the Authorization header. An attacker who can influence the environment, configuration, or invocati ...[truncated 1885 chars]

Remediation
View remediation

Remediation Suggestions

  1. Allowlist the documented endpoint, such as https://api.scnet.cn/api/llm/v1, rather than accepting an unrestricted base URL.
  2. If custom endpoints are operationally necessary, require explicit user approval and validate the URL with a proper parser:
    • Require HTTPS.
    • Reject embedded credentials.
    • Restrict ports and hostnames.
    • Resolve and reject loopback, link-local, private, and metadata-service addresses where appropriate.
    • Revalidate redirects or disable them.
  3. Validate ocr_type against the sole supported value, QUOTA_INVOICE.
  4. Allow only explicitly supported file types and verify their contents using file signatures rather than relying solely on filename extensions or mimetypes.guess_type.
  5. Reject archives unless archive processing is essential and safely constrained.
  6. Enforce a conservative maximum file size before reading or uploading the file.
  7. Restrict file selection to user-approved input locations and reject known credential, configuration, hidden, and system paths.
  8. Present a clear disclosure or confirmation that the selected invoice will be transmitted to a third-party OCR service.
  9. Avoid sending the SCNet credential to any host other than the approved SCNet API hostname.
  10. Add automated tests covering custom-host rejection, cleartext URL rejection, sensitive-path rejection, unsupported file formats, and oversized files.

T08 · Insecure Dependencies

Note
Location
SKILL.md:65
Finding

Unpinned third-party dependency installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:65-70
Vulnerability Type: Unpinned runtime dependency
Risk Level: Low

Vulnerable Code

markdown
### 依赖安装

本技能需要 Python 3.6+ 和 requests 库。请运行以下命令:

```bash
   pip install requests
text

### Technical Analysis

The documented installation command retrieves whichever version of `requests` and its transitive dependencies is currently selected by the configured Python package index. No exact version, integrity hash, locked transitive dependency set, or trusted index configuration is provided.

This makes installations non-reproducible and permits future dependency changes to alter the Skill's behavior without a corresponding review of this project. It also broadens supply-chain exposure if the package index, dependency resolution process, or a future release is compromised.

The package name `requests` is legitimate and there is no evidence that this project intentionally references a malicious, typosquatted, or dependency-confusion package. The finding concerns the unsafe installation practice rather than a confirmed malicious dependency.

### Attack Path

1. A user follows the documented `pip install requests` instruction.
2. `pip` queries its configured package index and resolves the latest compatible package and transitive dependencies.
3. A compromised or unexpectedly modified package release is downloaded without verification against project-maintained hashes.
4. Package installation or later import executes the compromised dependency code with the privileges of the user running the Skill.

### Impact Assessment

A compromised dependency could execute code with the permissions of the installing or invoking account. This could expose the OCR input, environment variables such as `SCNET_API_KEY`, local files available to that account, and network-accessible resources.

Exploitation depends on compromise or unsafe configuration of the external pac
...[truncated 88 chars]
Remediation
View remediation

Remediation Suggestions

  1. Provide a reviewed dependency lock or requirements file containing exact versions.
  2. Include cryptographic hashes and install with:
    bash
    pip install --require-hashes -r requirements.txt
    
  3. Pin and review transitive dependencies as well as the direct requests dependency.
  4. Configure installation to use a trusted package index and HTTPS.
  5. Regularly scan pinned dependencies for known vulnerabilities and update them through a controlled review process.
  6. Document the supported Python versions consistently and test the locked dependency set against each supported version.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (19)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 20)May include surrounding context.

python
# 获取技能根目录(脚本所在目录的上一级)
SKILL_ROOT = Path(__file__).parent.parent.absolute()
ENV_FILE = SKILL_ROOT / "config" / ".env"

# --- 新增:重试配置 ---
MAX_RETRIES = 3            # 最大重试次数

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 61)May include surrounding context.

md
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 123)May include surrounding context.

md
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 29)May include surrounding context.

python
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 32)May include surrounding context.

python
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 45)May include surrounding context.

python
# --------------------

def load_config():
    """从环境变量或 .env 文件加载配置,环境变量优先"""
    config = {}

    # 1. 如果 config/.env 存在,先加载其中的变量

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · README.md (reported line 41)May include surrounding context.

md
* [Get started with GitLab CI/CD](https://docs.gitlab.com/ee/ci/quick_start/)
* [Analyze your code for known vulnerabilities with Static Application Security Testing (SAST)](https://docs.gitlab.com/ee/user/application_security/sast/)
* [Deploy to Kubernetes, Amazon EC2, or Amazon ECS using Auto Deploy](https://docs.gitlab.com/ee/topics/autodevops/requirements.html)
* [Use pull-based deployments for improved Kubernetes management](https://docs.gitlab.com/ee/user/clusters/agent/)
* [Set up protected environments](https://docs.gitlab.com/ee/ci/environments/protected_environments.html)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
95% confidence
Finding

The skill advertises capabilities that imply environment access, local file reads, network access, and shell execution, but it does not declare any explicit tool scope such as permissions or allowed-tools. This creates an authorization and transparency gap: an agent may invoke broader capabilities than a user or platform reviewer expects, especially since the skill processes local files and sends data to a remote API.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest description and user-facing instructions are written exclusively in Chinese, which effectively forces a specific language for users interacting with the skill. The file does not indicate that Chinese is optional, selectable, or limited to a justified region-specific deployment.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 53)May include surrounding context.

SCNET_API_KEY=your_scnet_api_key_here

API 基础地址(一般无需修改)

SCNET_API_BASE=https://api.scnet.cn/api/llm/v1

text
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 105)May include surrounding context.

SCNET_API_KEY=your_scnet_api_key_here

API 基础地址(一般无需修改)

SCNET_API_BASE=https://api.scnet.cn/api/llm/v1

text
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)

External Transmission

Medium
Category
Data Exfiltration
Confidence
91% confidence
Finding

The file hardcodes use of an external API endpoint, confirming that document images/PDFs/archives are transmitted off-system for processing. In the context of fixed-amount invoice OCR, those uploads may contain sensitive financial records, so external transmission is security-relevant even if it is the intended product behavior.

Content

Scanner excerpt · references/api-docs.md (reported line 4)May include surrounding context.

md
# Sugon-Scnet OCR API 文档摘要

## 接口地址
`POST https://api.scnet.cn/api/llm/v1/ocr/recognize`

## 请求头
- `Content-Type: multipart/form-data`

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The document instructs the skill to upload user-supplied files to a third-party OCR endpoint and use the returned extracted document fields, but it does not disclose that raw files and derived sensitive invoice data leave the local environment. For an OCR skill handling invoices, this can expose financial or personal data and creates privacy, compliance, and data-handling risk if users or operators are not explicitly warned.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

The code is hardwired by default to send data to an external endpoint at api.scnet.cn, which means invoice images and metadata leave the local environment. In an OCR skill handling financial documents, this external transmission expands the trust boundary and can expose sensitive business information if users are unaware or if the remote service is misconfigured or compromised.

Content

Scanner excerpt · scripts/main.py (reported line 55)May include surrounding context.

python
config['SCNET_API_BASE'] = env_api_base

    # 3. 设置默认值
    config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')

    # 4. 检查必要配置
    api_key = config.get('SCNET_API_KEY', '')

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · scripts/main.py (reported line 71)May include surrounding context.

python
"2. 配置文件:\n"
            f"   mkdir -p {SKILL_ROOT}/config\n"
            f"   echo 'SCNET_API_KEY=你的真实密钥' > {ENV_FILE}\n"
            f"   chmod 600 {ENV_FILE}\n"
        )
        sys.exit(error_msg)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script uploads the user-supplied file contents to a remote OCR API using requests.post(..., files=files) without any just-in-time disclosure or consent mechanism at the point of transmission. Because OCR inputs often contain sensitive financial data from invoices, silent external transmission creates a real privacy and data-handling risk even if the feature is intended.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest description is written only in Chinese ('支持定额发票识别。'), which imposes a specific language/locale in the skill metadata without any indication of user choice or a documented region-specific justification. This matches the policy category for language or locale constraints in natural-language content.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This markdown file presents its natural-language content primarily in Chinese, with no indication that the language choice is optional or tied to a documented region-specific requirement. Under the stated policy, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

This file contains user-facing natural-language documentation only in Chinese, which can constitute a language/locale policy violation when no user opt-in or justification is provided. The content does not indicate that the skill is intentionally region-specific or that alternative languages are available.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.