T09 · Insecure Skill Coding Practices
- Location
scripts/main.py:86- Finding
Arbitrary OCR Endpoint Can Receive API Credentials and Sensitive Documents
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This OCR skill appears purpose-aligned, but it should be reviewed because it uploads potentially sensitive certificate files and an API token to a configurable remote endpoint without enforcing a trusted HTTPS host.
Review before installing. Use it only for documents you are allowed to send to Scnet, keep config/.env owner-readable only, and do not set SCNET_API_BASE unless it is a trusted HTTPS endpoint you intend to receive both the document and API token.
scripts/main.py:86Arbitrary OCR Endpoint Can Receive API Credentials and Sensitive Documents
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# 获取技能根目录(脚本所在目录的上一级)
SKILL_ROOT = Path(__file__).parent.parent.absolute()
ENV_FILE = SKILL_ROOT / "config" / ".env"
# --- 新增:重试配置 ---
MAX_RETRIES = 3 # 最大重试次数
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# --------------------
def load_config():
"""从 .env 文件加载配置,若文件不存在则抛出友好错误"""
if not ENV_FILE.exists():
error_msg = (
"\n===============================================\n"
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# --------------------
def load_config():
"""从 .env 文件加载配置,若文件不存在则抛出友好错误"""
if not ENV_FILE.exists():
error_msg = (
"\n===============================================\n"
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# --------------------
def load_config():
"""从 .env 文件加载配置,若文件不存在则抛出友好错误"""
if not ENV_FILE.exists():
error_msg = (
"\n===============================================\n"
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
* [Get started with GitLab CI/CD](https://docs.gitlab.com/ee/ci/quick_start/)
* [Analyze your code for known vulnerabilities with Static Application Security Testing (SAST)](https://docs.gitlab.com/ee/user/application_security/sast/)
* [Deploy to Kubernetes, Amazon EC2, or Amazon ECS using Auto Deploy](https://docs.gitlab.com/ee/topics/autodevops/requirements.html)
* [Use pull-based deployments for improved Kubernetes management](https://docs.gitlab.com/ee/user/clusters/agent/)
* [Set up protected environments](https://docs.gitlab.com/ee/ci/environments/protected_environments.html)
The skill declares capabilities that imply local file access, network access, and shell execution, but it does not define any explicit tool scope such as permissions or allowed-tools. This weakens containment and reviewability because an agent may invoke broader capabilities than users expect when processing local documents and sending them to an external OCR service.
description 使用中文描述技能能力,且全文面向中文交互,没有说明是否支持其他语言或允许用户选择语言。根据策略,若技能在自然语言层面默认强制特定语言且无用户 opt-in,可能构成语言/地区政策问题。
The skill is explicitly designed to send user-provided document images and OCR requests to an external endpoint at api.scnet.cn. External transmission is expected for cloud OCR, but it remains a real security concern because local document contents may include sensitive organizational identifiers and are exported to a third party.
SCNET_API_KEY=your_scnet_api_key_here
SCNET_API_BASE=https://api.scnet.cn/api/llm/v1
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)
The auto-trigger guidance is broad and says the AI may automatically invoke the skill based on description keywords, without clear boundaries or exclusion rules. In context, this is risky because the skill reads a local file path and transmits document contents to a third-party OCR API, so ambiguous triggering can cause unintended exfiltration of sensitive documents.
This configuration section again confirms that requests are sent to an external SCNET API base URL. Because the skill handles certificate images and structured OCR output, unintended or poorly understood outbound transfer can expose sensitive business document data outside the local environment.
| 变量名 | 默认值 | 说明 |
|--------|--------|------|
| SCNET_API_KEY | 必需 | Scnet API 密钥 |
| SCNET_API_BASE | https://api.scnet.cn/api/llm/v1 | API 基础地址(一般无需修改) |
### 输出
The file documents direct transmission of uploaded documents to an external domain, which is a real external data egress path. In the context of OCR for organization code certificates, this increases risk because certificates may include company identifiers, addresses, legal representative names, and other sensitive records that are exposed to a third-party processor.
# Sugon-Scnet OCR API 文档摘要
## 接口地址
`POST https://api.scnet.cn/api/llm/v1/ocr/recognize`
## 请求头
- `Content-Type: multipart/form-data`
The documentation instructs callers to upload files containing document images to a third-party OCR endpoint and authenticate with a bearer token, but it provides no warning that user-provided data leaves the local environment. Because this skill processes organization code certificates, the transmitted files and extracted fields can contain sensitive business and personal information, creating a privacy and compliance risk if users are not clearly informed.
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
" b) 配置文件:\n"
f" mkdir -p {SKILL_ROOT}/config\n"
f" echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
f" chmod 600 {ENV_FILE}\n"
"\n配置完成后重新运行。"
)
sys.exit(error_msg)
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
" b) 配置文件:\n"
f" mkdir -p {SKILL_ROOT}/config\n"
f" echo 'SCNET_API_KEY=你的密钥' > {ENV_FILE}\n"
f" chmod 600 {ENV_FILE}\n"
"\n配置完成后重新运行。"
)
sys.exit(error_msg)
This skill transmits user-supplied document files and an API bearer token to an external service at api.scnet.cn. External transmission is expected for an OCR skill, but it is still security-relevant because potentially sensitive certificate data leaves the local environment and could be exposed if the endpoint, configuration, or data handling is not tightly controlled.
)
sys.exit(error_msg)
config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')
return config
def recognize_with_retry(ocr_type, file_path, config, retry_count=0):
The manifest description is written only in Chinese ('支持组织机构代码证的票据识别。'), which indicates a fixed language/locale presentation without any stated opt-in or alternative. Under the policy, locale or language constraints should be optional or clearly justified as region-specific.
The changelog content is entirely in Chinese, including the title and feature description, with no indication that language selection is optional or that the skill is specifically intended for a Chinese-only audience. This can violate language/locale policy when a skill implicitly enforces a locale without user opt-in.
SKILL.md 在 L130 明确声称技能会在 429 限流时自动等待并最多重试 3 次,但本次提供的技能文件只有清单与使用文档,没有可验证的实现代码或重试逻辑。该表述会让使用者误以为技能具备自动恢复能力,属于文档对实际实现的积极性描述与可见实现不一致。
The module docstring presents the skill interface and output description entirely in Chinese, and later user-facing messages and usage text are also only in Chinese. Under the policy, forcing a specific language without opt-in can be a natural-language locale violation unless the locale restriction is explicitly justified.
No suspicious patterns detected.