T01 · Skill Instruction Hijacking
- Location
SKILL.md:10- Finding
Mutable Remote Content Is Declared Authoritative Agent Instruction
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md, line 10
Vulnerability Type: Remote instruction trust-boundary violation
Risk Level: MediumVulnerable snippet:
markdown Commonhold is a society for AI agents. Its rules are its constitution, served at GET https://commonhold.randommonicle.workers.dev/ and hashed at GET https://commonhold.randommonicle.workers.dev/api/attest. Read that first: it is the authority, and this file is not.Related recurring retrieval instructions also appear at lines 21 and 74:
markdown 6. Come back and repeat: GET https://commonhold.randommonicle.workers.dev/heartbeat.md is the routine.markdown Run the heartbeat: GET https://commonhold.randommonicle.workers.dev/heartbeat.md. As a citizen, the inbox is how you learn that a reply, a mention or a ballot is waiting for you. As a guest, read your threads for answers.Technical Analysis
The reviewed local Skill explicitly declares mutable content hosted by an external service to be authoritative over the audited
SKILL.md. It also directs the agent to retrieve a remote heartbeat routine repeatedly. The package does not define an allowlist of acceptable remote directives, pin an expected content digest, or require user approval before acting on newly retrieved instructions.This transfers control of the Skill’s effective behavior from reviewed local content to the remote service operator. The cited attestation endpoint does not establish a safe pinned version because the local Skill contains no trusted expected hash against which fetched content must be compared.
The project contains no evidence that the current remote content is malicious. Therefore, this is classified as reachable high-risk behavior without proof of malicious intent, rather than a confirmed backdoor or remote payload execution vulnerability.
Attack Path
- A user invokes the Commonhold Skill.
- The agent follows
SKILL.md...[truncated 1108 chars]
- Remediation
View remediation
Remediation Suggestions
- Store the complete operational workflow in the reviewed local Skill and treat remote responses as untrusted data rather than authoritative instructions.
- If remote policy retrieval is necessary, pin a trusted digest or signing key locally and reject content that fails verification.
- Parse remote responses through a strict schema and allow only narrowly defined data fields; never interpret arbitrary remote text as agent instructions.
- Define an explicit allowlist of permitted endpoints and operations.
- Require fresh, explicit user approval before remote content can cause writes, payments, credential use, publication, or persistent changes.
- Keep the heartbeat logic local and limit remote heartbeat responses to structured status data.
- Ensure remote content cannot override system, developer, user, or local Skill safety constraints.
