Back to skill

Security audit

钉钉 AI 表格跨表格洞察分析

Security checks for vulnerabilities and agentic risk

Overview

The skill is a plausible DingTalk table-analysis tool, but it under-discloses sensitive data handling and has unsafe command execution patterns that warrant careful review before use.

Install only after reviewing this with the DingTalk workspace owner or security admin. Use a narrow read-only token, avoid all-table scans unless explicitly intended, prefer --no-llm for sensitive data, do not print or store tokens in shell startup files, and treat this package as needing fixes for shell command construction, temporary files, dependency pinning, and accurate LLM/data-retention disclosure.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (6)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/analyze_tables.py:208
Finding

Arbitrary Command Execution Through Unsafe Shell Command Construction

Content
View full analysis
"{tmp_file}" 2>&1' result = subprocess.run(cmd_redirect, shell=True, timeout=30) ``` A second vulnerable command-construction path handles paginated records: ```python if cursor: cmd = (f'mcporter --config {config_path} call dingtalk-ai-table.search_base_record ' f'dentryUuid="{doc_id}" sheetIdOrName="{sheet_identifier}" ' f'limit={page_limit} cursor="{cursor}" > "{tmp_file}" 2>&1') else: cmd = (f'mcporter --config {config_path} call dingtalk-ai-table.search_base_record ' f'dentryUuid="{doc_id}" sheetIdOrName="{sheet_identifier}" ' f'limit={page_limit} > "{tmp_file}" 2>&1') subprocess.run(cmd, shell=True, timeout=30) ``` ### Technical Analysis The code constructs command strings by interpolating configuration paths, user-supplied keywords, document identifiers, sheet identifiers, and pagination cursors. These strings are then executed using `shell=True`. Wrapping a value in double quotes does not make it safe for shell execution. A malicious value can terminate the quoted argument or invoke shell substitutions. Relevant input sources include: - The `--keyword` command-line argument. - `DINGTALK_MCP_CONFIG`, which controls `config_path`. - Document and sheet identifiers returned by the MCP service. - Paginat ...[truncated 1332 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/analyze_tables.py:220
Finding

Predictable Temporary Files Permit Symlink Attacks and File Clobbering

Content
View full analysis
"{tmp_file}" 2>&1' result = subprocess.run(cmd_redirect, shell=True, timeout=30) with open(tmp_file, 'r', encoding='utf-8') as f: output = f.read().strip() ``` The same pattern is used for the LLM response: ```python tmp_file = tempfile.mktemp(suffix='.json') try: print(f" 🔄 调用 OpenClaw 大模型(--agent main)...") cmd = f'openclaw agent --agent main --message {shlex.quote(prompt[:10000])} --json --timeout 120 > "{tmp_file}" 2>&1' exit_code = subprocess.call(cmd, shell=True, timeout=130) ``` ### Technical Analysis `tempfile.mktemp()` generates a pathname but does not atomically create and open the file. A race exists between pathname generation and the shell opening that pathname for output redirection. A local attacker can create a file or symbolic link at the generated location before the redirection occurs. Because the shell follows symbolic links, command output may overwrite another file writable by the Skill process. The temporary files can also contain sensitive DingTalk business data, errors, or model responses. Deleting the pathname in a `finally` block does not prevent the initial race or undo an overwrite. ### Attack Path 1. The Skill generates a predictable temporary pathname using `tempfile.mktemp()`. 2. Before the shell opens that pathname, a local attacker creates a symbolic link at the same location. 3. The symbolic link points to another file writable by the Skill process. 4. Shell redirection opens the linked target and truncates or overwrites it with MCP or LLM output. 5. Alternatively, the attacker reads or replaces temporary content before the Skill parse ...[truncated 438 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
scripts/analyze_tables.py:638
Finding

Untrusted Table Content Is Submitted as Instructions to the Privileged Main Agent

Content
View full analysis
"{tmp_file}" 2>&1' ``` The alternate wrapper includes arbitrary field names and values: ```python for record in records: fields = record.get("fields", {}) sample = {} for i, (key, value) in enumerate(fields.items()): if i >= max_fields_per_record: break if isinstance(value, dict): value = value.get("name", value.get("text", str(value))) sample[key] = value table_info["数据示例"].append(sample) data_summary.append(table_info) full_prompt = f"{system_prompt}\n\n{user_prompt}" result = subprocess.run( ["openclaw", "agent", "--agent", "main", "--message", full_prompt[:10000], "--json"], capture_output=True, text=True, timeout=300 ) ``` ### Technical Analysis Table names, field names, and field values are untrusted content. They are inserted directly into a natural-language prompt sent to `openclaw agent --agent main` ...[truncated 1920 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
scripts/analyze_tables.py:493
Finding

Complete Tables Are Downloaded Before Sampling and Data Is Sent to the Default LLM Path

Content
View full analysis
limit_per_sheet: print(f" 📊 {sheet_name}: {len(records)}条 → 随机抽样{limit_per_sheet}条") records = random.sample(records, limit_per_sheet) ``` The sampled data is then included in an LLM request: ```python prompt = f"""请分析以下钉钉 AI 表格数据并生成洞察报告: ... **表格数据摘要**: {json.dumps(data_summary, ensure_ascii=False, indent=2)} ... """ cmd = f'openclaw agent --agent main --message {shlex.quote(prompt[:10000])} --json --timeout 120 > "{tmp_file}" 2>&1' exit_code = subprocess.call(cmd, shell=True, timeout=130) ``` The documentation claims local-only processing: ```markdown - ✅ **本地分析** - 所有分析在本地执行 - ✅ **无数据留存** - 分析结果不上传外部服务 ``` ### Technical Analysis `max_records=None` instructs the pagination function to retrieve every available record in each sheet. Sampling is performed only after the complete dataset has been transferred to and stored in the local process. This contradicts the documented implication that only a limited sample is read. It also increases the consequences of local compromise, temporary-file attacks, memory inspection, logs, crashes, and dependency vulnerabilities. LLM mode is enabled by default. The selected samples are sent to `openclaw agent`, whose configured model backend may be remote. The script cannot guarantee that this processing is local or that the backend retains no data. ### Attack Path 1. A user runs the Skill using its default settings. 2. The Skill enumerates all accessible tables and sheets. 3. Pagination retrieves every record because `max_records` is `None`. 4. The ...[truncated 962 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
references/dependencies.md:9
Finding

Executable Dependencies Are Installed Without Immutable Version or Integrity Pinning

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
references/dependencies.md:58
Finding

Documentation Encourages Plaintext Token Persistence and Full Token Disclosure

Content
View full analysis
> ~/.bashrc source ~/.bashrc ``` It also recommends displaying the complete token during verification: ```bash # 3. 检查环境变量 echo $DINGTALK_MCP_TOKEN ``` ### Technical Analysis Shell startup files are general-purpose plaintext configuration files. Storing a long-lived access token in `.bashrc` or `.zshrc` exposes it to every interactive shell and to any process or script capable of reading that file. Printing the complete environment variable can disclose it through terminal recordings, command output capture, support transcripts, CI logs, remote-session logs, or shoulder surfing. Environment variables may also be inherited by child processes that do not require DingTalk access. ### Attack Path 1. A user follows the permanent configuration recommendation and writes the token to a shell startup file. 2. The token is loaded into all future interactive shell environments. 3. The user follows the verification step and prints the entire token. 4. A local process, another account with relevant file access, a terminal recorder, or a logging system captures the credential. 5. The attacker reuses the token to access DingTalk resources permitted by that credential. ### Impact Assessment The exposed privileges are those granted to `DINGTALK_MCP_TOKEN`. Potential impact includes: - Unauthorized reading of accessible DingTalk AI tables. - Disclosure of organizational business data. - Continued access until the token expires or is revoked. - Abuse from another host if the token is not device-bound. The documentation states that read-only permissions ...[truncated 76 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (55)

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · SECURITY.md (reported line 12)May include surrounding context.

md
- **Sampling limits** - Each table reads maximum 100 records by default to minimize data exposure

### Authentication
- **Token security** - MCP access token stored in local config file (`config/mcporter.json`)
- **No hardcoded credentials** - All sensitive values come from environment or config files
- **Minimal permissions** - Only requires read access to DingTalk AI tables

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/analyze_tables.py (reported line 607)May include surrounding context.

python
- **Sampling limits** - Each table reads maximum 100 records by default to minimize data exposure

### Authentication
- **Token security** - MCP access token stored in local config file (`config/mcporter.json`)
- **No hardcoded credentials** - All sensitive values come from environment or config files
- **Minimal permissions** - Only requires read access to DingTalk AI tables

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The documentation claims a narrower, local, read-sampled analysis workflow, but other sections indicate remote LLM invocation by default, local caching/logging, and potentially broader data collection through pagination across all sheets. This mismatch can cause users to expose more data than expected, including sensitive business records being processed externally or retained locally.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The security section states that analysis is entirely local and that results are not uploaded externally, yet the skill elsewhere says it invokes an LLM through openclaw agent by default. This is dangerous because users may supply confidential table data under a false assumption of local-only processing, leading to unauthorized external disclosure.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The file contains contradictory security claims: one section promises local-only processing and no external uploads, while another describes remote LLM-based analysis. Contradictory security assurances undermine informed consent and can lead organizations to use the skill in regulated or sensitive contexts where external transmission is prohibited.

Content

No source excerpt is available for this finding.

YARA rule 'backdoor_persistence': Backdoor persistence with malicious payloads (shell commands, SSH key injection, hidden root users) [malware]

High
Category
YARA Match
Confidence
75% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · references/dependencies.md (reported line 64)May include surrounding context.

�� (如需要):**

bash
# macOS
brew install python@3.9

# Ubuntu/Debian
sudo apt-get install python3 python3-pip

# Windows
# 从 https://www.python.org/downloads/ 下载

钉钉 AI 表格 MCP Token

作用: 认证访问钉钉 AI 表格

配置方式:

bash
export DINGTALK_MCP_TOKEN="your-token-here"

永久配置 (推荐):

bash
# 添加到 ~/.bashrc 或 ~/.zshrc
echo 'export DINGTALK_MCP_TOKEN="your-token-here"' >> ~/.bashrc
source ~/.bashrc

可选依赖

mcporter CLI

作用: 获取可访问的表格列表(临时方案)

状态: ⚠️ 未来将移除,改用 dingtalk-ai-table 的 list 接口

安装:

bash
npm install -g mcporter

配置:

bash
# 配置文件位于
/home/admin/openclaw/workspace/config/mcporter.json

依赖关系图

text
dingtalk-ai-table-insights
├── dingtalk-ai-table (必需)
│   └── 钉钉 AI 表格 MCP
├── python3 (必需)
└── mcporter (临�

Credential Access

High
Category
Privilege Escalation
Confidence
70% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/dependencies.md (reported line 153)May include surrounding context.

症状:

text
Error: 403 - The access token should be issued by the organization

解决方案:

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The claim ‘不存储到外部服务、本地处理,不传输’ conflicts with earlier instructions to send summarized table data to an LLM via OpenClaw sessions or CLI. This creates a false privacy assurance, increasing the chance that users or admins approve analysis without understanding that data may leave the local trust boundary.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The module claims that data is analyzed locally and not uploaded externally, but the implementation sends table-derived summaries to an LLM agent command. This is dangerous because users may trust the privacy claim and unknowingly expose sensitive business data, project details, or operational metrics to another processing component or external service boundary.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

This is a practical tool-parameter abuse issue because untrusted parameters are concatenated into a shell command that invokes a privileged integration tool. In a skill that processes external table metadata and environment-derived configuration, an attacker could craft input that alters the mcporter invocation or executes arbitrary shell commands.

Content

Scanner excerpt · scripts/analyze_tables.py (reported line 225)May include surrounding context.

python
try:
                # 执行命令,输出重定向到临时文件
                cmd_redirect = f'{cmd} > "{tmp_file}" 2>&1'
                result = subprocess.run(cmd_redirect, shell=True, timeout=30)
                
                # 读取文件内容
                with open(tmp_file, 'r', encoding='utf-8') as f:

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
99% confidence
Finding

This command passes external identifiers and pagination state into a shell-constructed mcporter invocation. Because the skill's purpose is to query attacker-influenceable table names and IDs, the surrounding context makes command injection more realistic and more dangerous than in a closed-input script.

Content

Scanner excerpt · scripts/analyze_tables.py (reported line 427)May include surrounding context.

python
f'limit={page_limit} > "{tmp_file}" 2>&1')
            
            # 执行命令
            subprocess.run(cmd, shell=True, timeout=30)
            
            # 读取并解析 JSON
            with open(tmp_file, 'r', encoding='utf-8') as f:

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · scripts/analyze_tables.py (reported line 726)May include surrounding context.

python
# 使用 --agent main 参数(测试成功)
        cmd = f'openclaw agent --agent main --message {shlex.quote(prompt[:10000])} --json --timeout 120 > "{tmp_file}" 2>&1'
        
        exit_code = subprocess.call(cmd, shell=True, timeout=130)
        
        # 读取输出
        if os.path.exists(tmp_file):

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The changelog explicitly states that the skill sends detailed table distributions, actual record samples, and key business fields to an LLM for analysis, but it does not mention any consent flow, redaction, minimization beyond small sampling, or user-facing disclosure of external data handling. In a cross-table business analytics skill, those samples can contain sensitive operational, personnel, project, or incident data, so silent transmission to a model service creates a real confidentiality and compliance risk.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file’s title, instructions, and support content are all presented in Chinese, and there is no indication that users may choose another language or that the skill is restricted to a Chinese-speaking or region-specific audience. This can violate a language/locale policy when a skill forces a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README presents very broad natural-language activation examples such as analyzing all tables or all sales/project-related tables without defining clear scope, confirmation, or permission boundaries. In a skill that can access multiple DingTalk AI tables, this can lead to unintended over-collection of sensitive business data and overly broad analysis requests being executed when the user did not mean to authorize such wide access.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
82% confidence
Finding

The skill declares executable capabilities and dependency requirements that imply environment access, shell execution, and file I/O, but it does not define an explicit tool scope such as allowed-tools or permissions. That weakens least-privilege boundaries and makes it harder for users or the platform to understand and constrain what the skill may do at runtime.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest description begins in Chinese and presents the skill as a Chinese-language capability, while only a brief English sentence is appended. Under the policy, forcing a specific language or locale without user opt-in is a natural-language policy concern unless the skill explicitly offers a language choice or clearly documents a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The user-facing README content is predominantly Chinese and does not tell users they can choose another language for interaction or output. Because SQP-3 applies to all file types and includes language-policy violations, this is a policy issue unless the locale restriction is explicitly justified or made optional.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The document instructs users to create a persistent configuration file containing a live MCP server URL with embedded key material, but the warning about sensitive credentials is weak and appears only later as a .gitignore example rather than as an explicit secret-handling requirement at the point of use. In this skill context, the configuration grants access to DingTalk AI table data, so accidental disclosure through source control, logs, screenshots, or shared workspaces could expose business data and backend access.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
81% confidence
Finding

This duplicate finding points to the same behavior: storing MCP access details in a reusable workspace config file. Persistent shared config for an analysis skill that reads cross-table business data increases exposure because compromise of one file may enable ongoing unauthorized reads across multiple tables and workflows.

Content

Scanner excerpt · references/architecture.md (reported line 85)May include surrounding context.

bash
# 创建配置文件
mkdir -p <workspace>/config
vim <workspace>/config/mcporter.json

# 填入配置(从钉钉管理员获取)

Session Persistence

Medium
Category
Rogue Agent
Confidence
81% confidence
Finding

This duplicate finding points to the same behavior: storing MCP access details in a reusable workspace config file. Persistent shared config for an analysis skill that reads cross-table business data increases exposure because compromise of one file may enable ongoing unauthorized reads across multiple tables and workflows.

Content

Scanner excerpt · references/architecture.md (reported line 85)May include surrounding context.

bash
# 创建配置文件
mkdir -p <workspace>/config
vim <workspace>/config/mcporter.json

# 填入配置(从钉钉管理员获取)

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · references/dependencies.md (reported line 46)May include surrounding context.

md
brew install python@3.9

# Ubuntu/Debian
sudo apt-get install python3 python3-pip

# Windows
# 从 https://www.python.org/downloads/ 下载

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation tells users to run echo $DINGTALK_MCP_TOKEN, which prints the full credential to the terminal and potentially into shell history, logs, screenshots, or session recordings. This is an information disclosure issue because MCP tokens grant access to DingTalk AI table data and should not be displayed in plaintext unless strictly necessary.

Content

No source excerpt is available for this finding.

File System Enumeration

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code scans file system directories looking for sensitive files. This could be reconnaissance for credential theft.

Content

Scanner excerpt · references/dependencies.md (reported line 146)May include surrounding context.

clawhub install dingtalk-ai-table

验证安装位置

ls -la ~/.openclaw/skills/dingtalk-ai-table/scripts/

text

### 问题:权限不足

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The example 'scan all tables' trigger is broad and maps to ordinary user phrasing without clear confirmation or scope boundaries. In a skill designed to access many authorized AI tables, this can cause over-collection of sensitive business data beyond the user's intended need, increasing privacy and least-privilege risks.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.prompt_injection_instructions

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
references/llm_integration.md:305