Back to skill

Security audit

RAGFlow

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed RAGFlow management skill that uses the user's RAGFlow API key to manage datasets, documents, parsing, retrieval, and models.

Install only if you trust the RAGFlow server configured in RAGFLOW_API_URL and the RAGFLOW_API_KEY scope. Prefer HTTPS for remote servers, use least-privilege API keys where possible, and review delete/update actions carefully because the skill can modify or remove RAGFlow datasets and documents after user confirmation.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/common.py:84
Finding

Bearer API Key Can Be Transmitted Over an Insecure or Unintended Transport

Content
View full analysis
str: base_url = (cli_base_url or "").strip() or _require_env_var(RAGFLOW_API_URL_ENV) parsed = urllib.parse.urlsplit(base_url) if not parsed.scheme or not parsed.netloc: raise ConfigError( f"Invalid {RAGFLOW_API_URL_ENV}. Use an absolute URL such as http://127.0.0.1:9380." ) return base_url.rstrip("/") ``` ```python def request_json( url: str, api_key: str, *, method: str = "GET", body: bytes | None = None, content_type: str | None = None, accept: str = "application/json", ) -> dict[str, Any]: headers = {"Authorization": f"Bearer {api_key}"} if accept: headers["Accept"] = accept if content_type: headers["Content-Type"] = content_type request_obj = urllib.request.Request(url, headers=headers, data=body, method=method) try: with urllib.request.urlopen(request_obj, timeout=HTTP_TIMEOUT) as response: return decode_json_response(response.read()) ``` ### Technical Analysis The base URL validator only checks that the supplied value contains a scheme and network location. It does not restrict the scheme to HTTP or HTTPS, require TLS for remote hosts, reject embedded user information, or constrain the destination to an expected RAGFlow host. After this minimal validation, every request adds the RAGFlow API key to the `Authorization` header as a Bearer token. Therefore, a plaintext remote URL such as `http://attacker-controlled.example` would receive the API key without transport encryption. The documentat ...[truncated 1579 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Note
Location
scripts/parse_status.py:74
Finding

Parse Status Fetches Metadata for Every Document Before Applying the Requested ID Filter

Content
View full analysis
tuple[list[dict[str, Any]], int]: payload = ensure_success(request_json(_build_documents_url(base_url, dataset_id, page, page_size), api_key)) data = payload.get("data") if not isinstance(data, dict): raise DataError("Response missing data object.") docs = data.get("docs") total = data.get("total") if not isinstance(docs, list): raise DataError("Response missing data.docs.") if not isinstance(total, int): raise DataError("Response missing data.total.") return docs, total def _fetch_all_documents(base_url: str, api_key: str, dataset_id: str) -> list[dict[str, Any]]: all_docs: list[dict[str, Any]] = [] page = 1 total: int | None = None while True: docs, page_total = _fetch_documents_page(base_url, api_key, dataset_id, page, DEFAULT_PAGE_SIZE) if total is None: total = page_total all_docs.extend(docs) if len(all_docs) >= total or not docs: return all_docs[:total] page += 1 ``` ```python def collect_status_payload( dataset_id: str, target_ids: list[str] | None = None, *, base_url: str | None = None, api_key: str | None = None, ) -> dict[str, Any]: resolved_base_url = resolve_base_url(base_url) resolved_api_key = require_api_key(api_key) raw_documents = _fetch_all_documents(resolved_base_url, resolved_api_key, dataset_id) documents = [_normalize_document(raw_doc) for raw_doc in raw_documents] return _build_payload(dataset_id, _select_documents(documents, target_ids)) ``` ...[truncated 1590 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
scripts/parse.py:29
Finding

Unencoded Dataset IDs Can Alter Parse, Stop, and Upload Request Paths

Content
View full analysis
dict[str, object]: url = f"{base_url}/api/v1/datasets/{dataset_id}/chunks" body = json.dumps({"document_ids": document_ids}).encode("utf-8") response = ensure_success( request_json( url, api_key, method="POST", body=body, content_type="application/json", ) ) ``` ```python def stop_parse(dataset_id: str, document_ids: list[str], *, base_url: str, api_key: str) -> dict[str, Any]: url = f"{base_url}/api/v1/datasets/{dataset_id}/chunks" response = ensure_success( request_json( url, api_key, method="DELETE", body=json.dumps({"document_ids": document_ids}).encode("utf-8"), content_type="application/json", ) ) ``` ```python def upload_documents(dataset_id: str, file_paths: list[str], *, base_url: str, api_key: str) -> dict[str, Any]: missing = [path for path in file_paths if not Path(path).exists()] if missing: raise ConfigError("File(s) not found: " + ", ".join(missing)) boundary, body = _build_multipart(file_paths) url = f"{base_url}/api/v1/datasets/{dataset_id}/documents" request_obj = urllib.request.Request(url, data=body, method="POST") request_obj.add_header("Authorization", f"Bearer {api_key}") request_obj.add_header("Content-Type", f"multipart/form-data; boundary={boundary}") try: with urllib.request.urlopen(request_obj, timeout=120) as response: payload = decode_json_response(response.read()) ``` ### Technical Analysis The three operations interpolate `dataset_id` direct ...[truncated 1723 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (9)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The code is clearly limited to dataset operations in scripts/datasets.py: subcommands are list, info, create, and delete. There is no implementation for updating datasets, any document-related operations, parsing control/status, chunk retrieval, or configured model listing. This is a description-behavior mismatch because the declared description represents a much broader skill than this code chunk actually provides. The implemented resource scope (RAGFlow datasets API) is consistent with part of the description, but the supplied chunk does not support many of the declared capabilities.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 43)May include surrounding context.

md
python3 scripts/upload.py DATASET_ID /path/to/file.pdf --json

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 44)May include surrounding context.

md
python3 scripts/upload.py DATASET_ID /path/to/file.pdf --json

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · SKILL.md (reported line 69)May include surrounding context.

md
- If a parse status result includes `progress_msg`, surface it directly. For `FAIL`, treat it as the primary error detail.
- Use `--retrieval-test` only for single-dataset debugging or when the user explicitly asks for that endpoint.

## Output Rules

- Follow `reference.md`.
- Use tables for 3+ items when possible.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill declares access to environment variables and relies on Python scripts that will likely perform file and network operations, but it does not declare an explicit tool scope such as allowed-tools or permissions. This creates an overbroad execution surface where the agent may invoke capabilities beyond what is transparently documented, increasing the risk of unintended data access or outbound requests using the RAGFlow API key.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The delete_datasets function issues an HTTP DELETE request that can irreversibly remove datasets, but the code provides no confirmation prompt and no user-facing warning before execution. In this file, the delete subcommand help text also does not disclose the destructive nature beyond the verb itself.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

This code sends the user-provided query and optional document/dataset identifiers to a remote HTTP API, then prints retrieved chunk content back to stdout. There is no confirmation prompt, warning comment/docstring, or user-facing disclosure in this file indicating that input data will be transmitted externally and that retrieved content may be exposed in output.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The delete_documents function issues a DELETE request that removes documents from the dataset, which is a destructive operation. The code provides no confirmation prompt, no explicit warning message before execution, and no inline comment or docstring disclosing the destructive behavior beyond the command name.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · scripts/common.py (reported line 59)May include surrounding context.

python
return

    for stream_name in ("stdout", "stderr"):
        stream = getattr(sys, stream_name, None)
        if stream is None or not hasattr(stream, "reconfigure"):
            continue
        try:

Static analysis

No suspicious patterns detected.