Back to skill

Security audit

Mark Heartflow Skill

Security checks for vulnerabilities and agentic risk

Overview

HeartFlow discloses a local memory and reasoning engine, but it also persists user text into system-prompt memory and exposes host-level code execution with weak isolation.

Install only if you intentionally want a local persistent-memory agent layer and trusted local code execution. Avoid enabling the Hermes memory-injection plugin for untrusted or shared environments, do not pass untrusted code to codeExecutor or self-initiated task execution, set SHUTDOWN_TOKEN before using the daemon, and update/audit the npm dependency tree before relying on semantic-search features.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T02 · Agent Memory Poisoning

Error
Location
plugins/heartflow-memory-inject.py:85
Finding

Persistent system-prompt poisoning through unsanitized long-term memory

Content
View full analysis
'); return; } const result = hfm.learn(key, value, ['manual']); if (result.success) { console.log(`已写入 LEARNED 层: [${key}] = ${value}`); } } ``` Stored values are inserted verbatim into the generated prompt suffix. For example, generic learned entries are rendered without escaping or validation: ```js if (other.length > 0) { lines.push(''); lines.push('【其他记忆】'); for (const e of other.slice(0, 5)) { lines.push(` • ${e.value}`); } } ``` The plugin then attaches the resulting content to the system prompt before every processed message: ```python def before_message(self, context): """ 在每次处理用户消息前,注入心虫记忆到系统提示。 """ inject_text = _run_inject() if inject_text: return { "system_prompt_suffix": inject_text, } return {} ``` ### Technical Analysis The memory subsystem does not distinguish trusted instructions from untrusted data. An arbitrary LEARNED memory value is serialized as plain text and subsequently assigned system-prompt privilege through `system_prompt_suffix`. There is no: - Prompt-injection detection or neutralization. - Escaping or structured serialization that prevents memory text from being interpreted as instructions. - Provenance or trust-level enforcement. - Human approval before memory becomes part of the system prompt. - Separation between factual memory, user preferences, conversation text, and exe ...[truncated 1540 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
src/core/code/code-executor.js:438
Finding

Arbitrary host code execution through a non-isolated execution engine

Content
View full analysis
{ const timer = setTimeout(() => { reject(new Error(`执行超时 (${timeout}ms)`)); }, timeout); try { const result = fn(...args); if (result instanceof Promise) { result .then(val => { clearTimeout(timer); resolve(val); }) .catch(err => { clearTimeout(timer); reject(err); }); } else { clearTimeout(timer); resolve(result); } } catch (err) { clearTimeout(timer); reject(err); } }); } ``` Shell input is passed directly to Bash after only blacklist-based filtering: ```js const dangerCheck = checkDangerousCommand(code); if (dangerCheck.dangerous) { return { status: ExecStatus.SANDBOX_BLOCKED, output: '', error: `危险命令被阻止: ${dangerCheck.reason}`, truncated: false, execError: ExecError.SANDBOX }; } try { const result = execSync(code, { timeout, encoding: 'utf-8', maxBuffer: MAX_OUTPUT_LIMIT, shell: '/bin/bash' }); ``` Python input is written to a temporary file and executed directly on the host: ```js try { fs.writeFileSync(tmpFile, code, 'utf-8'); const pythonCmd = this._getPythonCommand(); const result = execSync(`${pythonCmd} "${tmpFile}"`, { timeout, encoding: 'utf-8', maxBuffer: MAX_OUTPUT_LIMIT }); ``` ### Technical Analysis The ordinary `execute()` interface is ...[truncated 2440 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Note
Location
bin/daemon.js:122
Finding

Daemon shutdown authentication fails open when the token is not configured

Content
View full analysis
process.exit(0), 100); } } ``` ### Technical Analysis Authentication is checked only when `SHUTDOWN_TOKEN` contains a truthy value. If the environment variable is absent or empty, the condition is false and the shutdown request is accepted without a token. This behavior contradicts the documented statement that shutdown requires `SHUTDOWN_TOKEN`. The Unix socket is set to mode `0700`, which limits exploitation to the same operating-system user under normal conditions, but it does not repair the fail-open authentication logic. ### Attack Path 1. The daemon starts without a non-empty `SHUTDOWN_TOKEN`. 2. A process running under the same operating-system user connects to `/tmp/heartflow-daemon.sock`. 3. It sends a JSON request such as: ```json {"cmd":"shutdown"} ``` 4. `expectedToken` is falsy, so the authorization rejection is skipped. 5. The daemon schedules `process.exit(0)` and terminates. ### Impact Assessment The issue permits local denial of service against the HeartFlow daemon by another process operating as the same user. It does not provide code execution, access to another user’s account, or privilege escalation. The practical scope is constrained by Unix socket permissions, but repeated shutdown requests could disrupt integrations that depend on daemon availability. ]]>
Remediation
View remediation
process.exit(0), 100); } ``` 2. Refuse to start the daemon’s remote shutdown interface unless a sufficiently strong token is configured. 3. Use constant-time comparison for token verification. 4. Generate a random per-installation or per-process token and store it with owner-only permissions. 5. Consider removing the socket shutdown command and relying on authenticated process management or operating-system signals. 6. Retain restrictive socket ownership and verify the socket is not replaced by another user before listening. 7. Rate-limit failed requests and log shutdown attempts without logging the secret token. ]]>
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (711)

Known Vulnerable Dependency: protobufjs==6.11.6 — 11 advisory(ies): CVE-2026-44294 (protobuf.js: Denial of service from crafted field names in generated code); CVE-2026-44293 (protobuf.js: Code injection through bytes field defaults in generated toObject c); CVE-2026-44289 (protobuf.js: Denial of service through unbounded protobuf recursion) +8 more

Critical
Category
Supply Chain
Confidence
97% confidence
Finding

This lockfile pins protobufjs 6.11.6, which the supplied advisory data identifies as having multiple known vulnerabilities, including denial of service via crafted recursion/field names and code-injection issues in generated code paths. Even though this file is only a dependency manifest, shipping a version with known critical flaws is a real supply-chain risk because downstream code may exercise vulnerable parser or codegen behavior implicitly.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Critical
Category
Not specified by scanner
Confidence
99% confidence
Finding

The whitelist explicitly exposes code execution routes such as codeExecutor.execute, runTests, and sandbox despite the stated purpose being a cognitive engine. Exposing execution primitives through a generic dispatch API can let downstream callers trigger arbitrary or semi-arbitrary code paths, making this a high-consequence capability mismatch even if the executor later attempts sandboxing.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · CHANGELOG.md (reported line 314)May include surrounding context.

md
167|#### 修复类别一:描述与行为不匹配
   168|- `package.json`:description 精确匹配 SKILL.md 认知/自愈引擎描述
   169|- `CHANGELOG.md`:标注 MarkCode 为可选独立组件,避免误导
   170|- `skills/video-generate/SKILL.md`:移除"自动写入 API 密钥到 .env"危险指令
   171|- `skills/zai-vision/SKILL.md`:添加安全警告,说明数据外传风险
   172|- `skills/desktop-agent/SKILL.md`:添加高风险警告,默认禁用说明
   173|- `skills/browser-automation/SKILL.md`:添加安全警告,说明网络访问范围

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · CHANGELOG.md (reported line 329)May include surrounding context.

md
182|- `scripts/awakening-integration.js`:添加安全头部,标注为哲学思考框架
   183|
   184|#### 修复类别三:数据泄露风险
   185|- `plugins/agentmemory/__init__.py`:移除 `_preload_agentmemory_dotenv()`,不再自动读取 .env
   186|- `plugins/agentmemory/__init__.py`:`sync_turn()` 默认禁用,需 `AGENTMEMORY_OBSERVE_ENABLED=1` 才发送数据
   187|- `src/core/autonomy/pdca-engine.js`:`saveTrace()` 截断敏感字段,文件权限 0600
   188|

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
88% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
89% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

Filesystem state persistence combined with workflow routing, psychological inference, multi-agent orchestration, and self-healing integrations creates a much more active runtime than the manifest suggests. This increased complexity and hidden state raise review difficulty and can mask impactful behavior behind a benign identity layer.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec, suspicious.dynamic_code_execution, suspicious.env_credential_access (+2 more)

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/core/code-engine.js:68

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/core/code-verifier.js:385

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/core/code/code-executor.js:283

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/proactive/self-initiator.js:896

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
src/core/code-engine.js:1586

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
src/core/code-verifier.js:366

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
src/core/code/code-executor.js:472

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
src/proactive/self-initiator.js:484

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/core/search/hybrid-search.js:51

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
src/core/search/hybrid-search.js:421

HTTPS certificate verification is disabled.

Warn
Code
suspicious.insecure_tls_verification
Location
src/utils/atomic-write.js:139