Back to skill

Security audit

Ragflow Workbench 1.0.0 En

Security checks for vulnerabilities and agentic risk

Overview

This RAGFlow automation skill is coherent, but it handles powerful admin credentials and API keys in ways users should review before installing.

Install only if you are comfortable with a local automation skill that can create an admin account, generate API keys, upload and delete knowledge-base data, and configure models. Before running bootstrap, choose a strong unique admin password, protect or relocate the .env file, avoid logging JSON output containing secrets, use HTTPS for any non-local RAGFlow server, and rotate credentials created with the default workflow.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/bootstrap_admin.py:22
Finding

Privileged credentials are created with insecure defaults, stored in plaintext, and exposed in command output

Content
View full analysis
argparse.Namespace: parser = argparse.ArgumentParser(description="初始化管理员账号并自动生成 RAGFlow API Key。") parser.add_argument("--base-url", default="http://127.0.0.1:9380", help="RAGFlow API 地址") parser.add_argument("--container-name", default="docker-ragflow-cpu-1", help="RAGFlow 容器名") parser.add_argument("--nickname", default="admin", help="管理员昵称") parser.add_argument("--email", default="admin@example.com", help="管理员邮箱") parser.add_argument("--password", default="Admin123456", help="管理员密码") parser.add_argument("--token-name", default="ragflow-skill-key", help="要创建/复用的 API Token 名称") parser.add_argument( "--env-file", default=str(default_env_file()), help="输出 .env 文件路径,默认写入技能目录下 .env", ) parser.add_argument("--json", action="store_true", dest="json_output", help="输出 JSON") return parser.parse_args() ``` ```python # scripts/bootstrap_admin.py:37-63 def run_bootstrap(args: argparse.Namespace) -> dict[str, Any]: encrypted = encrypt_password_via_docker(args.container_name, args.password) register_result = register_user(args.base_url, args.nickname, args.email, encrypted) jwt_token = login_user(args.base_url, args.email, encrypted) api_key = get_or_create_api_token(args.base_url, jwt_token, args.token_name) env_updates = { "RAGFLOW_API_URL": args.base_url.rstrip("/"), "RAGFLOW_API_KEY": api_key, "RAGFLOW_CONTAINER_NAME": args.container_name, "RAGFLOW_ADMIN_EMAIL": args.email, "RAGFLOW_ADMIN_PASSWORD": args.password, } write_env_file(Path(args.env_file), ...[truncated 4579 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/common.py:82
Finding

Authentication credentials and uploaded documents can be transmitted to arbitrary plaintext HTTP endpoints

Content
View full analysis
str: base_url = (cli_base_url or "").strip() or _require_env_var(RAGFLOW_API_URL_ENV) parsed = urllib.parse.urlsplit(base_url) if not parsed.scheme or not parsed.netloc: raise ConfigError( f"Invalid {RAGFLOW_API_URL_ENV}. Use an absolute URL such as http://127.0.0.1:9380." ) return base_url.rstrip("/") def require_api_key(api_key: str | None = None) -> str: api_key = (api_key or "").strip() or _require_env_var(RAGFLOW_API_KEY_ENV) return api_key def resolve_runtime_config(args: Any) -> tuple[str, str]: base_url = resolve_base_url(getattr(args, "base_url", None)) api_key = require_api_key(getattr(args, "api_key", None)) return base_url, api_key ``` ```python # scripts/common.py:165-185 def request_json( url: str, api_key: str, *, method: str = "GET", body: bytes | None = None, content_type: str | None = None, accept: str = "application/json", ) -> dict[str, Any]: headers = {"Authorization": f"Bearer {api_key}"} if accept: headers["Accept"] = accept if content_type: headers["Content-Type"] = content_type request_obj = urllib.request.Request(url, headers=headers, data=body, method=method) try: with urllib.request.urlopen(request_obj, timeout=HTTP_TIMEOUT) as response: return decode_json_response(response.read()) ``` ```python # scripts/bootstrap_common.py:52-76 def http_json( base_url: str, path: str, *, method: str = "GET", body: dict[str, An ...[truncated 6592 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
Findings (50)

Credential Access

High
Category
Privilege Escalation
Confidence
87% confidence
Finding

The README explicitly states that API URL, API key, and container name are written to a .env file after obtaining admin credentials and tokens. Storing active secrets in plaintext env files can expose credentials through source control commits, weak filesystem permissions, backups, or other local users/processes, which is especially relevant for an automation skill managing admin bootstrap and token generation.

Content

Scanner excerpt · README.md (reported line 28)May include surrounding context.

md
| Admin registration | Auto-encrypt password via container and register admin account |
| Auto-login & JWT | Obtain auth token for subsequent operations |
| API Key create/reuse | Auto-create named API token (reuse if exists) |
| .env auto-configuration | Write API URL, API Key, container name to env file |

### 3. Default Model Configuration
| Feature | Description |

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · README.md (reported line 1)May include surrounding context.

md
---
name: ragflow-workbench
description: "RAGFlow Workbench — end-to-end RAG platform on Windows: Docker install check, admin bootstrap and API key generation, default model setup (Embedding/Chat/Rerank), knowledge base CRUD, document upload-parse lifecycle, retrieval/search, chat creation. Triggers when the user mentions: RAGFlow installation or deployment, setting up a local RAG instance on Windows, creating a knowledge base, configuring models for RAG, managing documents in datasets, searching across knowledge bases."
---

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 1)May include surrounding context.

md
---
name: ragflow-workbench
description: "RAGFlow Workbench — end-to-end RAG platform on Windows: Docker install check, admin bootstrap and API key generation, default model setup (Embedding/Chat/Rerank), knowledge base CRUD, document upload-parse lifecycle, retrieval/search, chat creation. Triggers when the user mentions: RAGFlow installation or deployment, setting up a local RAG instance on Windows, creating a knowledge base, configuring models for RAG, managing documents in datasets, searching across knowledge bases."
---

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 20)May include surrounding context.

Determine the RAGFlow environment readiness from the .env file to avoid redundant checks:

text
┌─ Check if .env contains a valid RAGFLOW_API_KEY?
│
├─ ✅ Yes (connection established)
│    Skip environment checks, use API directly

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 55)May include surrounding context.

Determine the RAGFlow environment readiness from the .env file to avoid redundant checks:

text
┌─ Check if .env contains a valid RAGFLOW_API_KEY?
│
├─ ✅ Yes (connection established)
│    Skip environment checks, use API directly

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SKILL.md (reported line 87)May include surrounding context.

Determine the RAGFlow environment readiness from the .env file to avoid redundant checks:

text
┌─ Check if .env contains a valid RAGFLOW_API_KEY?
│
├─ ✅ Yes (connection established)
│    Skip environment checks, use API directly

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/bootstrap_admin.py (reported line 31)May include surrounding context.

python
Determine the RAGFlow environment readiness from the `.env` file to avoid redundant checks:

```
┌─ Check if .env contains a valid RAGFLOW_API_KEY?
│
├─ ✅ Yes (connection established)
│    Skip environment checks, use API directly

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/configure_default_models.py (reported line 22)May include surrounding context.

python
Determine the RAGFlow environment readiness from the `.env` file to avoid redundant checks:

```
┌─ Check if .env contains a valid RAGFLOW_API_KEY?
│
├─ ✅ Yes (connection established)
│    Skip environment checks, use API directly

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/configure_default_models.py (reported line 23)May include surrounding context.

python
Determine the RAGFlow environment readiness from the `.env` file to avoid redundant checks:

```
┌─ Check if .env contains a valid RAGFLOW_API_KEY?
│
├─ ✅ Yes (connection established)
│    Skip environment checks, use API directly

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/configure_default_models.py (reported line 35)May include surrounding context.

python
Determine the RAGFlow environment readiness from the `.env` file to avoid redundant checks:

```
┌─ Check if .env contains a valid RAGFLOW_API_KEY?
│
├─ ✅ Yes (connection established)
│    Skip environment checks, use API directly

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 97)May include surrounding context.

md
uv run python scripts/datasets.py create "Sample Knowledge Base" --json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 113)May include surrounding context.

md
uv run python scripts/datasets.py create "Sample Knowledge Base" --json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 98)May include surrounding context.

md
uv run python scripts/upload.py DATASET_ID /path/to/file.pdf --json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 114)May include surrounding context.

md
uv run python scripts/upload.py DATASET_ID /path/to/file.pdf --json

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · SKILL.md (reported line 129)May include surrounding context.

md
- If `parse_status.py` returns a `progress_msg`, echo it verbatim; when status is `FAIL`, treat it as the primary error and guide the user to [`references/troubleshooting.md`](references/troubleshooting.md)
- `bootstrap_admin.py` and `configure_default_models.py` require Docker containers accessible via `docker exec`

## Output Rules

- Follow `references/output-format.md`
- Use tables for 3+ structured data items

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/command-reference.md (reported line 8)May include surrounding context.

bash
uv venv
.\.venv\Scripts\Activate.ps1
copy .env.example .env

Installation Check & Bootstrap

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/bootstrap_admin.py (reported line 31)May include surrounding context.

python
parser.add_argument(
        "--env-file",
        default=str(default_env_file()),
        help="输出 .env 文件路径,默认写入技能目录下 .env",
    )
    parser.add_argument("--json", action="store_true", dest="json_output", help="输出 JSON")
    return parser.parse_args()

Credential Access

High
Category
Privilege Escalation
Confidence
91% confidence
Finding

Defaulting to a repository-adjacent .env file creates a high likelihood that secrets such as API tokens, passwords, or bootstrap credentials will be stored in a predictable plaintext location. In bootstrap tooling, this context makes the risk more serious because the same file is likely to hold powerful initial-access credentials that are often mishandled, committed, or copied across environments.

Content

Scanner excerpt · scripts/bootstrap_common.py (reported line 27)May include surrounding context.

python
def default_env_file() -> Path:
    return Path(__file__).resolve().parent.parent / ".env"


def read_env_file(path: Path) -> dict[str, str]:

Credential Access

High
Category
Privilege Escalation
Confidence
93% confidence
Finding

This script is designed to read administrator email credentials from a .env file and later use them to authenticate automatically. In an agent/skill context, handling admin credentials from local config increases secret exposure risk, especially if the environment file is broadly accessible, accidentally committed, or consumed by untrusted automation.

Content

Scanner excerpt · scripts/configure_default_models.py (reported line 24)May include surrounding context.

python
parser = argparse.ArgumentParser(description="配置 RAGFlow 默认 Embedding/Chat/Rerank 模型。")
    parser.add_argument("--base-url", help="RAGFlow API 地址,默认从 .env 读取 RAGFLOW_API_URL")
    parser.add_argument("--container-name", help="容器名,默认从 .env 读取 RAGFLOW_CONTAINER_NAME")
    parser.add_argument("--email", help="管理员邮箱,默认从 .env 读取 RAGFLOW_ADMIN_EMAIL")
    parser.add_argument("--password", help="管理员密码,默认从 .env 读取 RAGFLOW_ADMIN_PASSWORD")
    parser.add_argument("--embedding-model", default="bge-m3", help="Embedding 模型名")
    parser.add_argument("--chat-model", default="Gemma4:E2B-IT", help="Chat 模型名")

Credential Access

High
Category
Privilege Escalation
Confidence
98% confidence
Finding

The script reads an administrator password from .env and even falls back to a hardcoded default password later in run_configure(). In an automation skill, storing and consuming plaintext admin passwords from local files materially raises the risk of credential disclosure and unauthorized administrative access if the file is exposed or defaults remain unchanged.

Content

Scanner excerpt · scripts/configure_default_models.py (reported line 25)May include surrounding context.

python
parser.add_argument("--base-url", help="RAGFlow API 地址,默认从 .env 读取 RAGFLOW_API_URL")
    parser.add_argument("--container-name", help="容器名,默认从 .env 读取 RAGFLOW_CONTAINER_NAME")
    parser.add_argument("--email", help="管理员邮箱,默认从 .env 读取 RAGFLOW_ADMIN_EMAIL")
    parser.add_argument("--password", help="管理员密码,默认从 .env 读取 RAGFLOW_ADMIN_PASSWORD")
    parser.add_argument("--embedding-model", default="bge-m3", help="Embedding 模型名")
    parser.add_argument("--chat-model", default="Gemma4:E2B-IT", help="Chat 模型名")
    parser.add_argument("--rerank-model", default="bge-reranker-v2-m3", help="Rerank 模型名")

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file documents a batch document deletion capability but does not include any warning, confirmation requirement, or note about irreversibility. For markdown files, destructive behaviors that may affect user data should be accompanied by clear user-facing warnings.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This markdown command reference includes a delete-documents operation but provides no warning that it permanently removes documents or affects user data. For markdown files, destructive behaviors should be accompanied by clear warnings about impact to data or system integrity.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The guide explicitly states that an API token will be written to a local .env file but does not warn that this file contains sensitive credentials or advise how to protect it. This can lead users to leave secrets in plaintext on disk, commit them to source control, or expose them through backups and local compromise, especially in a quick-start workflow where users copy commands without adding safeguards.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script’s stated purpose is admin bootstrap and API key generation, but it also stores the plaintext admin password in a .env file and returns the generated API key in its output. Persisting and exposing multiple high-value secrets beyond the minimally necessary behavior increases the chance of credential theft through logs, shell history, shared workspaces, or accidental file disclosure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The code writes sensitive values including the admin password and API key into a .env file, which is commonly checked into repositories, copied between environments, or read by multiple processes. In the context of an admin bootstrap utility, these credentials grant broad control and create a durable secret exposure risk if the file is mishandled.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/bootstrap_admin.py:42

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/create_chat.py:106

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/datasets.py:269

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/list_models.py:282

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/parse_status.py:282

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/parse.py:93

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/search.py:265

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/stop_parse_documents.py:87

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/update_dataset.py:143

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/update_document.py:146

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
scripts/upload.py:297