Back to skill

Security audit

Keep Protocol

Security checks for vulnerabilities and agentic risk

Overview

The skill is a real agent communication package, but it gives an agent high-impact server bootstrap and messaging powers that are under-scoped and can affect Docker containers, execute mutable remote software, and expose message contents.

Install only in a trusted, local development environment unless you have reviewed and constrained the server startup path. Avoid using keep_ensure_server from an autonomous agent unless you are comfortable with Docker or Go execution, mutable remote artifacts, and possible container removal on the configured port. Do not expose the TCP server to untrusted networks, do not rely on src names as strong authenticated identities, and do not send secrets or sensitive memory because message bodies can be logged and routed to other agents.

Vulnerability Patterns
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T03 · Remote Payload Retrieval and Execution

Error
Location
python/keep/client.py:80
Finding

Mutable Remote Server Artifacts Are Downloaded and Executed

Content
View full analysis
bool: ``` ```python result = subprocess.run( [ "docker", "run", "-d", "--name", f"keep-server-{port}", "-p", f"{port}:9009", docker_image, ], capture_output=True, text=True, timeout=60, ) ``` ```python result = subprocess.run( ["go", "install", "github.com/clcrawford-dev/keep-server@latest"], capture_output=True, text=True, timeout=120, ) ``` ```python if go_bin: subprocess.Popen( [str(go_bin)], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, start_new_session=True, ) ``` The same capability is exposed directly to agents through `keep_ensure_server()` in `python/keep/mcp/server.py:145-176`. ### Technical Analysis The bootstrap function downloads and executes server software identified by mutable `latest` references. Neither the Docker image nor the Go module is pinned to an immutable digest or version. No checksum, signature, provenance attestation, or trusted-key verification is performed before execution. Consequently, the effective code executed by an already-reviewed Skill can change without any change to this repository. Compromise of the upstream repository, container registry, release pipeline, or publisher credentials could turn an otherwise legitimate bootstrap operation into arbitrary code execution. The Docker path executes the downloaded image through a local Docker daemon. In many environments, Docker access is effectively equivalent to powerful host access. The Go fallback installs a binary into the user's Go binary directory and launches it ...[truncated 1146 chars]
Remediation
View remediation
" ``` 2. Replace `@latest` with an audited, explicit Go module version. 3. Verify release signatures or provenance attestations before execution. 4. Maintain an allowlist of trusted image registries, module paths, versions, and digests. 5. Require explicit user confirmation before any download, installation, or process launch. Agent invocation alone should not authorize software installation. 6. Separate availability checking from installation: - `check_server()` should only inspect connectivity. - `install_server()` should be an explicit administrative operation. 7. Run the server with reduced privileges, a read-only filesystem, dropped Linux capabilities, resource limits, and restricted networking. 8. Return the exact artifact version and digest used so operators can audit the resulting environment. ]]>

T05 · Unauthorized Access and Privilege Escalation

Error
Location
python/keep/client.py:117
Finding

Server Bootstrap Forcibly Deletes Unrelated Docker Containers

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
keep.go:141
Finding

Ed25519 Signatures Do Not Bind Public Keys to Agent Identities

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
keep.go:223
Finding

Server Logs Complete Inter-Agent Message Bodies in Plaintext

Content
View full analysis
%s", p.Src, p.Typ, p.Body, p.Dst) ``` The packaged ClawHub implementation contains the same accepted-message logging behavior: ```go log.Printf("From %s (typ %d): %s -> %s", p.Src, p.Typ, p.Body, p.Dst) ``` ### Technical Analysis The server records the full `body` field of accepted packets and also logs the body of unsigned packets. Agent messages can contain user prompts, task details, internal context, generated content, secrets, or information derived from memory systems. Logging is performed without redaction, field-level filtering, sensitivity classification, configurable opt-in behavior, or documented retention controls. Anyone with access to process logs, container logs, centralized logging infrastructure, backups, or diagnostics can recover message contents. Ed25519 signatures provide integrity but no confidentiality. The underlying TCP transport is also unencrypted, although this finding specifically concerns durable plaintext disclosure through logging. ### Attack Path 1. An agent sends a packet containing sensitive task data in `body`. 2. The server accepts and routes the packet. 3. The complete body is written to standard process logs. 4. Container logging drivers or centralized observability systems retain the entry. 5. A user, administrator, compromised log collector, or other party with log access reads the sensitive content. An attacker can also submit unsigned packets containing crafted sensitive-lo ...[truncated 506 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (76)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code does match part of the description: it uses TCP, protobuf serialization, ed25519 signature verification, and routing between named agents. However, several prominently declared capabilities are absent from the supplied code. There is no discovery mechanism, no memory-sharing functionality, and nothing related to MCP tools or latency optimization. The implementation is a straightforward central relay/server with heartbeats and registration-by-first-valid-source packet. Because the declared description significantly overstates the provided functionality, this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The core of the description is partially accurate: the code does implement signed Protobuf packets over TCP, agent registration, routing, and basic discovery. However, key advertised capabilities are not present in this code chunk. There is no MCP tool integration, no explicit mechanism for sub-second latency beyond ordinary TCP handling, and no memory-sharing implementation. The code mainly acts as a signed message relay/router with simple discovery and metrics. Because the declared description includes notable features absent from the implementation, this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
89% confidence
Finding

The core declared protocol behavior is mostly represented: the code signs protobuf packets with ed25519, sends them over TCP, supports discovery, and can listen for routed packets. However, this chunk also performs additional undeclared capabilities that are materially important from a security perspective: it probes ports, invokes Docker and Go through subprocesses, removes and starts Docker containers, installs a Go package, launches a background server process, and persists endpoint metadata in the user's home directory. Those are not implied by the description's simple protocol/client framing and no permissions are declared. Also, the description mentions MCP tools and memory sharing, but this code chunk does not implement those features. Therefore the description does not accurately represent the actual behavior of this code.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description presents a communication protocol/library centered on signed Protobuf packets, authentication, discovery, routing, and memory sharing. The supplied code chunk does not implement or demonstrate those protocol features. Instead, it is specifically a test module for ensure_server() behavior: probing localhost ports, checking for Docker/Go binaries, invoking subprocesses to manage containers or install/run a server, and validating startup behavior with unit and integration tests. While this may support the broader project, the actual code's primary purpose is server bootstrap testing and local environment orchestration, which is materially different from the declared protocol functionality.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents a production protocol/communication skill with authenticated protobuf-over-TCP messaging, discovery, routing, and memory sharing. The supplied code chunk is instead a test harness for discovery-related server behavior. Its primary purpose is validation of server responses and cache behavior, not implementing or exposing the described protocol capabilities. It also accesses the local filesystem for endpoint caching, which is not mentioned in the description. While discovery is related to the declared domain, the code shown is materially different in purpose and omits the headline features claimed in the description (signed protobuf packets, ed25519 auth, routing, memory sharing, MCP tools).

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The memory barter section explicitly instructs users to share institutional knowledge via a packet field without any caution about secrets, personal data, proprietary information, or irreversible propagation to other agents. That creates a strong risk of sensitive data exfiltration or unintended lateral disclosure, especially because agent-to-agent sharing is presented as a normal feature.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · docs/clawhub-publishing.md (reported line 208)May include surrounding context.

" # Must see valid version and server response deactivate rm -rf /tmp/keep-test-sandbox

text

### Phase 3: Public Release

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
90% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · docs/clawhub-publishing.md (reported line 208)May include surrounding context.

" # Must see valid version and server response deactivate rm -rf /tmp/keep-test-sandbox

text

### Phase 3: Public Release

External Script Fetching

High
Category
Supply Chain
Confidence
90% confidence
Finding

Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.

Content

Scanner excerpt · docs/keep-protocol-virality-plan.md (reported line 29)May include surrounding context.

md
| Item | Why Deferred |
|------|--------------|
| Multicast discovery (224.0.0.1) | Blocked by cloud providers, needs Docker hacks |
| `curl \| bash` bootstrap | Supply chain risk, orgs block it |
| Rewards/reputation | Rabbit hole — no abuse data yet to justify complexity |
| Error evangelism | Feels spammy, let simplicity speak for itself |
| Signed binaries (cosign) | Add after adoption proves need |

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
70% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · docs/release-workflow.md (reported line 65)May include surrounding context.

bash
git add -A
git commit -m "chore: bump version to vX.Y.Z"
git push staging main
git push origin main

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · docs/release-workflow.md (reported line 152)May include surrounding context.

md
- Build feature (implementation)
- Test on keep-1 (verification)
- Package and publish to PyPI
- Update SKILL.md for ClawHub
- Community push (X post, etc.)

YARA rule 'agent_skill_mcp_tool_poisoning_metadata': MCP/tool metadata poisoning indicators in tool schemas or skill manifests [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · examples/mcp_tool_definition.py (reported line 10)May include surrounding context.

python
#!/usr/bin/env python3
"""Example: How to expose keep-protocol as an MCP tool for AI agents.

This shows the tool definition schema that an MCP server could use to let
Claude, ChatGPT, Copilot, or other AI agents send signed packets via tool calls.
"""

KEEP_SEND_TOOL = {
    "name": "keep_send",
    "description": (
        "Send a signed packet to another AI agent via the keep protocol. "
        "Uses ed25519 signatures over TCP+Protobuf for authenticated, "
        "low-latency agent-to-agent communication."
    ),
    "input_schema": {
        "type": "object",
        "properties": {
            "dst": {
                "type": "string",
                "description": (
                    "Destination agent or routing target. Examples: "
                    "'bot:weather', 'nearest:planner', 'swarm:research'"
                ),
            },
            "body": {
                "t

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

This code can install and execute external software through Docker and Go toolchains even though the module presents itself as a transport client. That creates a broad supply-chain and arbitrary-code-execution surface inconsistent with the feature's purpose, making it especially risky in automated agent environments.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

The README instructs users to run a container from a mutable :latest tag rather than a pinned version or image digest. This creates a supply-chain risk: a future image update or compromised registry account could cause users to execute unexpected code when following the documented one-liner.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README promotes ensure_server() auto-start behavior that may launch Docker containers or execute go install ...@latest on the host without an explicit warning or consent gate. In an agent skill context, this is more dangerous because an autonomous agent may interpret the documentation as permission to perform system-changing actions, leading to unreviewed code execution and package retrieval from remote sources.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill advertises operational steps that involve shell execution, environment configuration, package installation, and Docker usage, but it does not declare any explicit tool scope or permissions boundaries. In an agent environment, that omission can cause overbroad execution capability and weak operator awareness about what the skill may access or invoke.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The skill encourages agent discovery and listing connected identities without warning that this leaks metadata about other participants on the system or network. In multi-tenant or sensitive environments, enumerating active agent names can aid reconnaissance, targeting, and social-engineering of downstream agents.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

The Docker example pulls and runs a container from a mutable tag, which allows the image content to change over time without notice. That creates a supply-chain risk where users may execute unexpected or malicious code if the registry image is replaced, retagged, or compromised.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
91% confidence
Finding

Using npx clawhub without pinning an exact package version allows whatever version is current in the registry at invocation time to be fetched and executed. In a publishing workflow, that creates supply-chain risk because a compromised or breaking upstream release could run during login, publish, inspect, or whoami operations.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
91% confidence
Finding

This command again relies on unpinned npx clawhub, so the executed code may change between runs or be replaced by a malicious upstream version. Because it is part of auth verification, compromise here could expose credentials or mislead operators about account state.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The guide instructs users to run clawhub publish . and only later notes that the CLI publishes SKILL.md and supporting files in the directory, without an explicit up-front warning that this is a public release action. That can lead to accidental disclosure of sensitive files, draft content, or internal documentation if the working directory contains more than intended.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

Publishing with unpinned npx clawhub is especially risky because the command both fetches executable code and performs a sensitive release action. A malicious or unexpected CLI version could alter uploaded contents, exfiltrate tokens, or publish unintended artifacts.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

The inspect step also executes remote package code without version pinning, preserving the same supply-chain exposure. Even though this is a verification command, a compromised package could still execute arbitrary logic in the publisher's environment.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

The troubleshooting guidance tells users to rerun login with unpinned npx clawhub, reintroducing the same supply-chain risk during a credential-handling operation. Repetition in troubleshooting sections makes unsafe usage more likely because users often copy commands verbatim under time pressure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The plan explicitly proposes client.ensure_server() that checks a port and, if absent, automatically starts infrastructure via Docker or go install. In an agent-to-agent networking SDK, silently pulling images or installing/executing software expands trust boundaries, can bypass operator expectations, and creates supply-chain and unauthorized-execution risk even if framed as convenience.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.exposed_secret_literal, suspicious.obfuscated_code

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
examples/python_raw.py:19

Potential obfuscated payload detected.

Warn
Code
suspicious.obfuscated_code
Location
python/keep/keep_pb2.py:16