Back to skill

Security audit

Feishu Article Collector

Security checks for vulnerabilities and agentic risk

Overview

The skill mostly matches its Feishu article-saving purpose, but weak URL validation can make it fetch unintended sites and forward the result to DeepSeek and Feishu.

Review before installing. Use this only in an environment where Feishu Bitable writes and DeepSeek processing of article content are acceptable, and avoid exposing the trigger to untrusted chat participants until URL validation is fixed with strict hostname allowlisting and redirect checks.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/collect.py:66
Finding

Server-Side Request Forgery Through Substring-Based Domain Validation

Content
View full analysis
"{}|\\^`\[\]]+' urls = re.findall(url_pattern, text) for url in urls: if any(domain in url for domain in SUPPORTED_DOMAINS): return url return None ``` The accepted URL is passed to these request sinks: ```python resp = requests.get(url, headers=HEADERS, timeout=15, allow_redirects=True) ``` ### Technical Analysis The code determines whether a URL is permitted by checking whether a supported domain string occurs anywhere in the complete URL. It does not parse the URL and verify its actual hostname. An attacker can therefore place an allowed string in the path, query string, user-information component, or an attacker-controlled hostname. Examples that satisfy the substring check include: ```text http://127.0.0.1/admin?toutiao.com http://169.254.169.254/latest/meta-data/?mp.weixin.qq.com https://toutiao.com.attacker.example/article ``` In addition, `allow_redirects=True` permits an initially accepted URL to redirect to a loopback, private-network, link-local, or cloud metadata address. Redirect destinations are not revalidated. The fetched response is parsed as article content. That content can subsequently be transmitted to DeepSeek for analysis and represented in a Feishu record. Consequently, this issue can combine unauthorized internal resource access with external disclosure of retrieved information. This behavior exceeds the minimum network privileges required by the declared functionality, which only needs to retrieve articles from a small set of documented public domains. ### Attack Path 1. An attacker sends the agent a message containing a crafted HTTP or HTTPS URL. 2. The URL i ...[truncated 1281 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/collect.py:208
Finding

Indirect Prompt Injection Through Untrusted Article Content

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (37)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The README says DeepSeek generates summaries from article content, but it does not warn that article text and possibly metadata are transmitted to a third-party AI service. This is a meaningful privacy and compliance risk because users may submit copyrighted, confidential, or regulated content without realizing it leaves Feishu/OpenClaw and is processed externally.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The documented behavior says the skill collects articles, summarizes them, classifies them, and deduplicates them, but the detected implementation reportedly does not perform those functions and instead includes an undeclared action that grants a user full access to a created Feishu Bitable. A description-behavior mismatch is dangerous because users may authorize or trigger the skill under false assumptions, while hidden permission-granting can expose stored content to unauthorized modification or exfiltration.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The README is entirely written in Chinese and the documented behavior, including the auto-created table name and field labels, is fixed to Chinese without indicating any language or locale choice. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy concern unless the locale restriction is explicitly justified.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README states that the skill will automatically discover or create a Feishu Bitable and store article data, but it does not clearly warn users that triggering the skill causes persistent writes to an external workspace. This can lead to unintended data creation, privacy issues, and operational surprises, especially in shared organizational Feishu environments.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill invokes a Python script and explicitly requires environment secrets plus behavior that implies network, file read, and file write, but it declares no tool scope or permissions boundary. That makes the effective capability set broader and less auditable than the manifest suggests, increasing the risk of over-privileged execution or unintended access to local data, secrets, and external services.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill instructs the agent to pass the user's complete message text directly to a script, while also using external services and persistent storage, but it does not warn users that non-URL text may be processed, transmitted, or stored. In this context, users may include unrelated sensitive data in the same message, so forwarding the full message expands privacy and data-handling risk beyond what is necessary to collect an article link.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The HTTP headers force Accept-Language to prefer zh-CN, and the script overall is written to operate in Chinese without any opt-in or language selection mechanism. This can violate language/locale policy when a skill imposes a specific locale by default rather than offering user choice or clearly documenting a justified regional limitation.

Content

No source excerpt is available for this finding.

Tainted flow: 'url' from requests.post (line 404, network input) → requests.get (network output)

Medium
Category
Data Flow
Confidence
94% confidence
Finding

The script extracts a URL from untrusted message text and fetches it with redirects enabled, but domain validation is done via substring matching (domain in url). An attacker can craft a hostname like mp.weixin.qq.com.evil.com or abuse redirects so the script makes arbitrary outbound requests, creating an SSRF vector and unintended network access from the agent environment.

Content

Scanner excerpt · scripts/collect.py (reported line 113)May include surrounding context.

python
def fetch_wechat_article(url):
    """抓取微信公众号文章"""
    try:
        resp = requests.get(url, headers=HEADERS, timeout=15, allow_redirects=True)
        resp.encoding = "utf-8"
        html = resp.text

Tainted flow: 'url' from requests.post (line 404, network input) → requests.get (network output)

Medium
Category
Data Flow
Confidence
94% confidence
Finding

Like the WeChat fetch path, this function performs a GET to a URL taken from untrusted message input after only weak substring-based domain checks. That allows crafted domains or redirect chains to bypass intended restrictions and turn the collector into an SSRF primitive.

Content

Scanner excerpt · scripts/collect.py (reported line 140)May include surrounding context.

python
def fetch_toutiao_article(url):
    """抓取今日头条文章正文"""
    try:
        resp = requests.get(url, headers=HEADERS, timeout=15, allow_redirects=True)
        resp.encoding = "utf-8"
        html = resp.text

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill sends scraped article content and titles to the external DeepSeek service for summarization without any consent, notice, or policy gate. If users submit private, paywalled, or sensitive content, this causes unauthorized third-party disclosure and may violate privacy, confidentiality, or compliance expectations.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

This duplicate external-transmission finding is substantively valid for the same reason: article content is sent off-platform to DeepSeek. The risk is privacy and confidentiality leakage rather than code execution.

Content

Scanner excerpt · scripts/collect.py (reported line 218)May include surrounding context.

python
{{"summary": "总结内容", "category": "分类标签"}}"""

    try:
        resp = requests.post(
            "https://api.deepseek.com/v1/chat/completions",
            headers={
                "Authorization": f"Bearer {deepseek_api_key}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

This duplicate external-transmission finding is substantively valid for the same reason: article content is sent off-platform to DeepSeek. The risk is privacy and confidentiality leakage rather than code execution.

Content

Scanner excerpt · scripts/collect.py (reported line 218)May include surrounding context.

python
{{"summary": "总结内容", "category": "分类标签"}}"""

    try:
        resp = requests.post(
            "https://api.deepseek.com/v1/chat/completions",
            headers={
                "Authorization": f"Bearer {deepseek_api_key}",

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

The explicit DeepSeek endpoint confirms third-party transmission of article data. In skill context, this is more dangerous because users may assume the tool only fetches and stores content in Feishu, not that it also forwards full content to a separate AI provider.

Content

Scanner excerpt · scripts/collect.py (reported line 219)May include surrounding context.

python
try:
        resp = requests.post(
            "https://api.deepseek.com/v1/chat/completions",
            headers={
                "Authorization": f"Bearer {deepseek_api_key}",
                "Content-Type": "application/json",

Tainted flow: 'url' from requests.post (line 404, network input) → requests.post (network output)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/collect.py (reported line 252)May include surrounding context.

python
def get_tenant_access_token(app_id, app_secret):
    url = f"{BASE_URL}/auth/v3/tenant_access_token/internal"
    resp = requests.post(url, json={
        "app_id": app_id,
        "app_secret": app_secret,
    })

Tainted flow: 'app_token' from requests.post (line 297, network input) → requests.get (network output)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/collect.py (reported line 278)May include surrounding context.

python
if f.get("name") == BITABLE_NAME and f.get("type") == "bitable":
                app_token = f["token"]
                # 获取 table_id
                resp2 = requests.get(
                    f"{BASE_URL}/bitable/v1/apps/{app_token}/tables",
                    headers=headers,
                )

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/collect.py (reported line 288)May include surrounding context.

python
return app_token, table_id

    # 未找到,创建新的多维表格
    resp = requests.post(
        f"{BASE_URL}/bitable/v1/apps",
        headers=headers,
        json={"name": BITABLE_NAME},

Tainted flow: 'app_token' from requests.post (line 297, network input) → requests.get (network output)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/collect.py (reported line 301)May include surrounding context.

python
table_id = data["data"]["app"]["default_table_id"]

    # 重命名默认字段为"文章标题"
    resp = requests.get(
        f"{BASE_URL}/bitable/v1/apps/{app_token}/tables/{table_id}/fields",
        headers=headers,
    )

Tainted flow: 'app_token' from requests.post (line 297, network input) → requests.get (network output)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/collect.py (reported line 355)May include surrounding context.

python
table_id = data["data"]["app"]["default_table_id"]

    # 重命名默认字段为"文章标题"
    resp = requests.get(
        f"{BASE_URL}/bitable/v1/apps/{app_token}/tables/{table_id}/fields",
        headers=headers,
    )

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/collect.py (reported line 308)May include surrounding context.

python
fields_data = resp.json()
    if fields_data.get("code") == 0 and fields_data["data"]["items"]:
        first_field = fields_data["data"]["items"][0]
        requests.put(
            f"{BASE_URL}/bitable/v1/apps/{app_token}/tables/{table_id}/fields/{first_field['field_id']}",
            headers=headers,
            json={"field_name": "文章标题", "type": first_field["type"]},

Tainted flow: 'app_token' from requests.post (line 297, network input) → requests.put (network output)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/collect.py (reported line 308)May include surrounding context.

python
fields_data = resp.json()
    if fields_data.get("code") == 0 and fields_data["data"]["items"]:
        first_field = fields_data["data"]["items"][0]
        requests.put(
            f"{BASE_URL}/bitable/v1/apps/{app_token}/tables/{table_id}/fields/{first_field['field_id']}",
            headers=headers,
            json={"field_name": "文章标题", "type": first_field["type"]},

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/collect.py (reported line 337)May include surrounding context.

python
{"field_name": "已读", "type": 7},
    ]
    for field in fields:
        requests.post(
            f"{BASE_URL}/bitable/v1/apps/{app_token}/tables/{table_id}/fields",
            headers=headers,
            json=field,

Tainted flow: 'app_token' from requests.post (line 297, network input) → requests.post (network output)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/collect.py (reported line 337)May include surrounding context.

python
{"field_name": "已读", "type": 7},
    ]
    for field in fields:
        requests.post(
            f"{BASE_URL}/bitable/v1/apps/{app_token}/tables/{table_id}/fields",
            headers=headers,
            json=field,

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/collect.py (reported line 396)May include surrounding context.

python
},
        "page_size": 1,
    }
    resp = requests.post(api, headers=headers, json=body)
    data = resp.json()
    if data.get("code") != 0:
        return False

Tainted flow: 'api' from requests.post (line 381, network input) → requests.post (network output)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/collect.py (reported line 396)May include surrounding context.

python
},
        "page_size": 1,
    }
    resp = requests.post(api, headers=headers, json=body)
    data = resp.json()
    if data.get("code") != 0:
        return False

Static analysis

No suspicious patterns detected.