Back to skill

Security audit

别瞎说 - AI事实核查器/Checktruth - AI Fact-Checker

Security checks for vulnerabilities and agentic risk

Overview

The core fact-checking skill is coherent, but bundled optional scripts can send user content and API keys to external or custom LLM endpoints, so users should review them before use.

Core installation appears suitable for fact-checking workflows, but do not run the optional reference scripts on confidential text unless you accept sending that text to the configured LLM providers. If using the reference code, use a dedicated virtual environment, pin dependencies, set least-privileged API keys, avoid exposing unrelated environment variables, and do not set custom API base URLs unless you fully trust the endpoint.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
reference/src/fetch_answers.py:16
Finding

API credentials and user content can be transmitted to an unrestricted configurable endpoint

Content
View full analysis
str: """Call the OpenAI API.""" api_key = os.getenv("OPENAI_API_KEY") or os.getenv("LLM_API_KEY") api_base = os.getenv("OPENAI_API_BASE", "https://api.openai.com/v1") if not api_key: raise ValueError("Please set OPENAI_API_KEY or LLM_API_KEY") try: import openai client = openai.OpenAI(api_key=api_key, base_url=api_base) response = client.chat.completions.create( model=model, messages=[{"role": "user", "content": question}], temperature=0.1, max_tokens=2000, ) return response.choices[0].message.content except ImportError: import requests headers = { "Authorization": f"Bearer {api_key}", "Content-Type": "application/json" } payload = { "model": model, "messages": [{"role": "user", "content": question}], "temperature": 0.1, "max_tokens": 2000 } resp = requests.post( f"{api_base}/chat/completions", headers=headers, json=payload, timeout=60 ) resp.raise_for_status() return resp.json()["choices"][0]["message"]["content"] ``` The same pattern is used by the processing modules: ```python api_key = os.getenv("LLM_API_KEY") api_base = os.getenv("LLM_API_BASE", "https://api.openai.com/v1") model = os.getenv("LLM_MODEL", "gpt-4o") client = openai.OpenAI(api_key=api_key, base_url=api_base) ``` ### Technical Analysis The op ...[truncated 2598 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
reference/requirements.txt:1
Finding

Unpinned third-party dependencies allow unreviewed package versions to be installed

Content
View full analysis
=1.0.0 zhipuai>=2.0.0 dashscope>=1.0.0 moonshot-sdk>=0.1.0 pyyaml>=6.0 ``` The documented installation command is: ```bash pip install -r requirements.txt ``` ### Technical Analysis Every dependency uses an open-ended minimum-version constraint. No upper bounds, lockfile, hashes, or package integrity constraints are provided. A future release satisfying these ranges can therefore be selected without having been reviewed with the Skill. Python package installation and import can execute package-controlled code. If an upstream package or publishing account is compromised, a malicious version is released, or a future release introduces a security regression, users following the documented installation procedure may install that version automatically. No evidence was found that the named dependencies are currently malicious or that the project uses typo-squatted package names. This finding concerns unsafe dependency management and the resulting supply-chain exposure. ### Attack Path 1. An upstream package publishing account or distribution channel is compromised, or an unsafe future version is published under one of the declared package names. 2. The malicious or vulnerable release retains a version number satisfying the corresponding `>=` constraint. 3. A developer follows `reference/README.md` and runs `pip install -r requirements.txt`. 4. The package resolver selects the affected release because no exact version or trusted hash is required. 5. Package-controlled installation or runtime code executes in the developer's Python environment when installed or imported. ### Impact Assessment Impact is bounded by the account and environment used for installation or execution, ...[truncated 550 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (66)

Tainted flow: 'headers' from os.getenv (line 43, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · reference/src/consistency.py (reported line 53)May include surrounding context.

python
"temperature": 0.1,
            "max_tokens": 1000
        }
        resp = requests.post(
            f"{api_base}/chat/completions",
            headers=headers,
            json=payload,

Tainted flow: 'headers' from os.getenv (line 49, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · reference/src/decompose.py (reported line 59)May include surrounding context.

python
"temperature": 0.1,
            "max_tokens": 2000
        }
        resp = requests.post(
            f"{api_base}/chat/completions",
            headers=headers,
            json=payload,

Tainted flow: 'headers' from os.getenv (line 34, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
93% confidence
Finding

The OpenAI base URL is taken from the OPENAI_API_BASE environment variable and used directly for the POST destination, while the Authorization bearer token is attached to the request. If an attacker can influence the environment, they can redirect the request to an arbitrary host and capture the API key and submitted question data.

Content

Scanner excerpt · reference/src/fetch_answers.py (reported line 44)May include surrounding context.

python
"temperature": 0.1,
            "max_tokens": 2000
        }
        resp = requests.post(
            f"{api_base}/chat/completions",
            headers=headers,
            json=payload,

Tainted flow: 'headers' from os.getenv (line 61, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · reference/src/fetch_answers.py (reported line 72)May include surrounding context.

python
"temperature": 0.1,
        "max_tokens": 2000
    }
    resp = requests.post(
        "https://api.anthropic.com/v1/messages",
        headers=headers,
        json=payload,

Tainted flow: 'url' from os.getenv (line 90, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

The Gemini request URL includes the API key as a query parameter, which is more exposure-prone than using an authorization header because URLs may be captured in logs, proxies, browser history, or monitoring systems. This increases the risk of accidental credential leakage even though the destination is an expected provider endpoint.

Content

Scanner excerpt · reference/src/fetch_answers.py (reported line 100)May include surrounding context.

python
"maxOutputTokens": 2000
        }
    }
    resp = requests.post(url, json=payload, timeout=60)
    resp.raise_for_status()
    data = resp.json()
    return data["candidates"][0]["content"]["parts"][0]["text"]

Tainted flow: 'headers' from os.getenv (line 45, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · reference/src/verify.py (reported line 55)May include surrounding context.

python
"temperature": 0.1,
            "max_tokens": 1000
        }
        resp = requests.post(
            f"{api_base}/chat/completions",
            headers=headers,
            json=payload,

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 463)May include surrounding context.

md
**核心功能零配置**:本 Skill 的所有核心验证逻辑均通过 `SKILL.md` 中的指令、

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · reference/src/consistency.py (reported line 19)May include surrounding context.

python
prompt_path = Path(__file__).parent.parent / "prompts" / "decompose.txt"
    with open(prompt_path, "r", encoding="utf-8") as f:
        prompt = f.read()
    return prompt.replace("{{ANSWER}}", answer)


def call_llm(prompt: str) -> str:

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · reference/src/decompose.py (reported line 19)May include surrounding context.

python
prompt_path = Path(__file__).parent.parent / "prompts" / "decompose.txt"
    with open(prompt_path, "r", encoding="utf-8") as f:
        prompt = f.read()
    return prompt.replace("{{ANSWER}}", answer)


def call_llm(prompt: str) -> str:

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · reference/src/verify.py (reported line 21)May include surrounding context.

python
prompt_path = Path(__file__).parent.parent / "prompts" / "decompose.txt"
    with open(prompt_path, "r", encoding="utf-8") as f:
        prompt = f.read()
    return prompt.replace("{{ANSWER}}", answer)


def call_llm(prompt: str) -> str:

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The instruction requires that every step output progress in Chinese, which imposes a specific language regardless of the user's preferred locale. This is reinforced later by a rule stating all output must use Chinese, with no opt-in or alternative language path documented.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The rule '所有输出使用中文' explicitly requires all outputs to be in Chinese. Under the policy, forcing a specific language without user choice is a locale-policy violation unless clearly justified or made optional.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The file title and the entire document are written exclusively in Chinese, with no indication that other languages are supported or that the Chinese-only requirement is intentional for a region-specific audience. Under the language/locale policy rule, forcing a specific language without user opt-in can be a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The requirements explicitly say the tool will send user questions and candidate answers to multiple external AI APIs, but they do not require user notice, consent, minimization, or handling rules for sensitive content. Because questions and answers may contain personal, confidential, or proprietary data, this creates a real privacy and data-governance risk rather than a purely theoretical issue.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The design explicitly states that user-provided question/answer content will be sent via HTTP to multiple third-party AI APIs using configured API keys, but it does not mention any user notice, consent flow, redaction, or data-handling constraints. This creates a real privacy and data-governance risk because potentially sensitive prompts, answers, or references may be disclosed to external providers without the user's awareness.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The prompt is written entirely in Chinese and instructs the model to act as a checker using Chinese output examples, which imposes a specific language/locale behavior. There is no indication that the user can choose another language or that the locale restriction is justified as region-specific.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The entire prompt is written in Chinese and specifies the required output schema using Chinese field examples, which imposes a specific language/locale behavior on the skill. There is no indication that the user can choose another language or that this locale restriction is justified by a region-specific purpose.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The prompt is written entirely in Chinese and requires the model to produce the JSON field values such as the reason in Chinese, but it does not offer any language choice or explain a justified region-specific requirement. This is a natural-language locale policy issue because it imposes a specific language without user opt-in.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · reference/README.md (reported line 86)May include surrounding context.

md
| DeepSeek | https://platform.deepseek.com/ | `DEEPSEEK_API_KEY` |
| 混元(腾讯) | https://cloud.tencent.com/product/hunyuan | `HUNYUAN_API_KEY` |
| Kimi(Moonshot) | https://platform.moonshot.cn/ | `MOONSHOT_API_KEY` |
| MiniMax | https://api.minimax.chat/ | `MINIMAX_API_KEY` |

---

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

This code sends user content to Tencent Hunyuan over an external API, which creates a real data-exposure boundary crossing. In context, the danger is not the HTTPS endpoint itself but the combination of automatic third-party transmission and the absence of inline consent, minimization, or classification of what data is safe to send.

Content

Scanner excerpt · reference/multi_model_verify.py (reported line 110)May include surrounding context.

python
try:
        from openai import OpenAI
        
        client = OpenAI(api_key=api_key, base_url="https://api.hunyuan.cloud.tencent.com/v1")
        prompt = f"""请综合各方信息,对以下回答进行事实核查,给出平衡、公正的判断。

问题:{question}

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

This path transmits supplied content to Moonshot/Kimi, exposing whatever the user entered to an external service. In a multi-model verification tool, the same content may be replicated across several vendors, amplifying privacy, confidentiality, and regulatory risk if users submit sensitive material.

Content

Scanner excerpt · reference/multi_model_verify.py (reported line 142)May include surrounding context.

python
try:
        from openai import OpenAI
        
        client = OpenAI(api_key=api_key, base_url="https://api.moonshot.cn/v1")
        prompt = f"""请作为事实核查员,验证以下回答。

问题:{question}

External Transmission

Medium
Category
Data Exfiltration
Confidence
89% confidence
Finding

This code sends input to MiniMax's external API, constituting a genuine outbound data transfer to a third party. The security concern is heightened by the script's purpose: users may submit arbitrary factual claims or documents, including confidential internal text, without any guardrail that warns them the data leaves the local environment.

Content

Scanner excerpt · reference/multi_model_verify.py (reported line 170)May include surrounding context.

python
try:
        from openai import OpenAI
        
        client = OpenAI(api_key=api_key, base_url="https://api.minimax.chat/v1")
        prompt = f"""请作为事实核查员,验证以下回答中的事实。

问题:{question}

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script transmits user-supplied question/answer/text content to multiple third-party LLM providers based solely on the presence of API keys, but it does not present a clear runtime warning, consent step, or data-handling notice before sending potentially sensitive content off-host. In a verification tool, users may paste confidential documents, personal data, or proprietary claims, so silent multi-provider exfiltration materially increases privacy and compliance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The module sends the provided answer content to an external LLM service without any consent prompt, warning, or data-classification check. If users pass sensitive, regulated, or proprietary text, this can cause unintended third-party disclosure and privacy/compliance issues.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
84% confidence
Finding

This code transmits prompt content derived from the user's answer to an external network service. In a consistency-checking skill, that makes the context more sensitive because arbitrary answer text may contain confidential or personal information, creating a real data-exfiltration/privacy risk if users are unaware or if the endpoint is not tightly controlled.

Content

Scanner excerpt · reference/src/consistency.py (reported line 53)May include surrounding context.

python
"temperature": 0.1,
            "max_tokens": 1000
        }
        resp = requests.post(
            f"{api_base}/chat/completions",
            headers=headers,
            json=payload,

Static analysis

No suspicious patterns detected.