T09 · Insecure Skill Coding Practices
- Location
scripts/main.py:50- Finding
Unrestricted API Endpoint Can Disclose Credentials and Sensitive Medical Documents
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill appears to perform medical-invoice OCR as described, but it uploads sensitive documents with an API key to a configurable network endpoint without strong scoping or a clear consent warning.
Review this skill before installing if you will process real medical invoices. Use it only when you are comfortable sending the selected file to Scnet, keep SCNET_API_BASE fixed to the official HTTPS endpoint, protect the API key, and avoid relying on automatic invocation for sensitive documents unless the agent asks for confirmation first.
scripts/main.py:50Unrestricted API Endpoint Can Disclose Credentials and Sensitive Medical Documents
SKILL.md:68Unpinned Third-Party Dependency Installation
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# 获取技能根目录(脚本所在目录的上一级)
SKILL_ROOT = Path(__file__).parent.parent.absolute()
ENV_FILE = SKILL_ROOT / "config" / ".env"
# --- 新增:重试配置 ---
MAX_RETRIES = 3 # 最大重试次数
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# --------------------
def load_config():
"""从环境变量或 .env 文件加载配置,环境变量优先"""
config = {}
# 1. 如果 config/.env 存在,先加载其中的变量
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# --------------------
def load_config():
"""从环境变量或 .env 文件加载配置,环境变量优先"""
config = {}
# 1. 如果 config/.env 存在,先加载其中的变量
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# --------------------
def load_config():
"""从环境变量或 .env 文件加载配置,环境变量优先"""
config = {}
# 1. 如果 config/.env 存在,先加载其中的变量
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# --------------------
def load_config():
"""从环境变量或 .env 文件加载配置,环境变量优先"""
config = {}
# 1. 如果 config/.env 存在,先加载其中的变量
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
# --------------------
def load_config():
"""从环境变量或 .env 文件加载配置,环境变量优先"""
config = {}
# 1. 如果 config/.env 存在,先加载其中的变量
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
* [Get started with GitLab CI/CD](https://docs.gitlab.com/ee/ci/quick_start/)
* [Analyze your code for known vulnerabilities with Static Application Security Testing (SAST)](https://docs.gitlab.com/ee/user/application_security/sast/)
* [Deploy to Kubernetes, Amazon EC2, or Amazon ECS using Auto Deploy](https://docs.gitlab.com/ee/topics/autodevops/requirements.html)
* [Use pull-based deployments for improved Kubernetes management](https://docs.gitlab.com/ee/user/clusters/agent/)
* [Set up protected environments](https://docs.gitlab.com/ee/ci/environments/protected_environments.html)
The skill documents use of environment variables, local file access, network calls, and shell execution but does not declare an explicit tool scope such as permissions or allowed-tools. This creates a transparency and policy gap: an agent may invoke capabilities broader than users expect, especially when handling sensitive medical invoice data and local file paths.
The skill processes medical invoices, which commonly contain highly sensitive personal and health-related information, and sends image contents to a third-party OCR API. The documentation explains token setup and API usage but does not clearly and prominently warn users that local files and medical data will leave the device and be transmitted to an external service.
This finding reflects an external transmission endpoint used by the OCR service. External transmission is expected for a cloud OCR skill, but in this context it is security-relevant because the transmitted content may include medical invoices and other sensitive personal data.
SCNET_API_KEY=your_scnet_api_key_here
SCNET_API_BASE=https://api.scnet.cn/api/llm/v1
2. 添加:`SCNET_API_KEY=你的密钥`
3. 设置文件权限为 600(仅所有者可读写)
This is a second instance documenting the same external API base URL. While not malicious by itself, it confirms that the skill relies on third-party network transmission for processing sensitive medical invoice content, making privacy and consent controls important.
| 变量名 | 默认值 | 说明 |
|--------|--------|------|
| SCNET_API_KEY | 必需 | Scnet API 密钥 |
| SCNET_API_BASE | https://api.scnet.cn/api/llm/v1 | API 基础地址(一般无需修改) |
### 输出
The file explicitly references a remote API endpoint for OCR processing, which means uploaded documents and their extracted contents are transmitted outside the agent's local boundary. Because this skill handles medical invoices, the external transmission is more sensitive than ordinary OCR and can expose personally identifiable, billing, and healthcare-adjacent data if users are not properly informed or if the vendor is not appropriately vetted.
# Sugon-Scnet OCR API 文档摘要
## 接口地址
`POST https://api.scnet.cn/api/llm/v1/ocr/recognize`
## 请求头
- `Content-Type: multipart/form-data`
The documentation instructs users to upload medical invoice files to a third-party OCR endpoint but does not clearly warn that sensitive financial and potentially medical personal data will leave the local environment. In the context of medical invoices, this increases privacy, compliance, and data-handling risk because users may unknowingly transmit regulated or confidential information to an external service.
The skill is designed to send files and metadata to an external network endpoint by default. In the context of medical invoice OCR, this transmission materially increases risk because sensitive documents leave the local environment and are processed by a remote service, potentially implicating confidentiality and compliance obligations.
config['SCNET_API_BASE'] = env_api_base
# 3. 设置默认值
config.setdefault('SCNET_API_BASE', 'https://api.scnet.cn/api/llm/v1')
# 4. 检查必要配置
api_key = config.get('SCNET_API_KEY', '')
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
"2. 配置文件:\n"
f" mkdir -p {SKILL_ROOT}/config\n"
f" echo 'SCNET_API_KEY=你的真实密钥' > {ENV_FILE}\n"
f" chmod 600 {ENV_FILE}\n"
)
sys.exit(error_msg)
The script uploads the user-supplied file directly to a third-party OCR API, and this skill is explicitly for medical invoices, which commonly contain highly sensitive personal and financial data. There is no explicit consent flow, privacy notice, destination allowlist enforcement, or data minimization, so users may unknowingly exfiltrate regulated data off-host.
The manifest description is written only in Chinese ("支持识别医疗发票识别"), which indicates a language-specific presentation without any stated user opt-in or justification for restricting the skill to that locale. Under the policy, language or locale constraints should either offer user choice or be clearly documented as region-specific.
This markdown file uses Chinese throughout, including headings and change descriptions, but does not state that the skill is China-specific or that Chinese is an intentional locale choice. Under the policy rule for natural-language violations, forcing a specific language without user opt-in or documented justification can be a concern.
This markdown file references testing, CI/CD, integrations, and deployment workflows, including deployment to Kubernetes, EC2, or ECS, but provides no warning that these actions can affect live systems, infrastructure, or project state. For markdown files, skill descriptions should disclose behaviours that could affect system integrity or user data when such operational actions are described.
The documentation states the AI will automatically trigger the skill based on keywords, which can lead to external transmission of sensitive medical invoice images without clear, explicit opt-in at the moment of use. In the medical context, automatic activation increases the chance of surprise data sharing and weakens informed consent.
Natural-language strings throughout the script, including the module description, usage text, and error messages, are only presented in Chinese. This imposes a specific language on all users without opt-in or locale selection, which matches the policy's language/locale violation criteria.
No suspicious patterns detected.