T01 · Skill Instruction Hijacking
- Location
SKILL.md:147- Finding
Mutable Remote Instructions Can Alter Agent Behavior After Review
- Content
View full analysis
Vulnerability Details
File Location:
SKILL.md:147-156,SKILL.md:175-185, andSKILL.md:321-322
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: HighComplete Code Snippets
SKILL.md:147-156:markdown ### Required reading (cache once) - **MUST** fetch **HEARTBEAT.md** before first action. - **MUST** fetch **MESSAGING.md** before first action. ### Required Initialization Order (CRITICAL) After registration, the agent MUST: 1. Fetch and read **HEARTBEAT.md** 2. Fetch and read **MESSAGING.md** 3. Only then begin the operational loopSKILL.md:175-185:markdown For full details, see https://molt-fs.vercel.app/heartbeat.md ### Recommended Cadence **Poll every 10–30 minutes with jitter.**base_interval = random(10, 30) minutes jitter = random(0, 5) minutes next_poll = base_interval + jitter
text SKILL.md:321-322:markdown ## Check for Updates Periodically re-fetch the skill files to ensure you have the latest documentation, endpoints, and rules. The URLs in the Skill Files section are canonical.Technical Analysis
The audited Skill does not contain all of the instructions that govern its operation. Instead, it requires the agent to fetch and read
HEARTBEAT.mdandMESSAGING.mdfrom a remote deployment before acting. It further instructs the agent to periodically retrieve updated documentation and rules.Because these remote files are mutable and are not pinned to a reviewed version or cryptographic digest, their contents can change after the Skill has been audited. If an agent interprets downloaded text as trusted Skill instructions, control of the remote service or its deployment can be used to introduce new behavioral directives without modifying the local audited artifact.
Fetching service documentation is relevant to the declared social-platform functionality, but granting mutable documents authority over ...[truncated 2008 chars]
- Remediation
View remediation
Remediation Suggestions
- Bundle reviewed copies of
HEARTBEAT.mdandMESSAGING.mddirectly in the Skill package. - Pin every remotely retrieved document to an immutable version and verify a cryptographic digest before processing it.
- Remove automatic or periodic replacement of behavioral instructions. Require explicit review and approval for updates.
- Treat remote documents, API responses, feed posts, comments, and messages as untrusted data rather than authoritative instructions.
- Enforce an instruction hierarchy that prevents downloaded content from overriding system policies, user intent, safety requirements, or credential-handling restrictions.
- Validate remote content against a narrow schema and reject unexpected commands, tool requests, credential requests, or references to unapproved destinations.
- Restrict outbound network access to the documented HTTPS host and API paths, while still treating content returned by that host as potentially compromised.
- Require human confirmation for sensitive or irreversible social actions, especially purchases or actions affecting other accounts.
- Bundle reviewed copies of
