Install
openclaw skills install @mfang0126/crew-contractopenclaw skills install @mfang0126/crew-contractThe main agent conducts — it doesn't play every instrument. The general division of labor for any substantial work, adapted from the "Astra + Luna" orchestrator pattern. The main agent keeps the goal, constraints, evidence and final judgment; bounded subagents do noisy execution; an independent reviewer checks material changes before anything is called done. Applies to engineering, research, writing, data, ops — anything. Project-specific processes plug in from their own docs; this skill is the method layer.
Do not use for simple questions, single small edits, or quick lookups — keep those in the main thread. Never create agents just to satisfy a template; use only the roles that improve the task, and scale to the job (explorer + worker may be enough; add tester/reviewer when risk warrants). Never switch the main model merely to activate this mode. Owner instructions always take precedence over this method.
This skill states intent and defaults — the dispatcher sizes effort, review and budget to the task's actual goal, and states a reason when deviating from a default. Only the invariants are non-negotiable: separation of duties (no self-review), evidence over self-report, fail-closed scope, agents don't commit, owner's permission/cost boundaries.
Defaults (adjust when the task clearly calls for it — note why):
| Role | Writes? | Job | Returns |
|---|---|---|---|
| explorer | no | Map files/symbols/flows/constraints/tests before implementation | Concise paths/symbols, risks, recommended boundary |
| researcher | no | Answer one external/version-specific question from primary sources | Verified facts vs inference; version/date assumptions; references; uncertainty |
| worker | yes | ONE bounded change with explicit ownership; smallest defensible change | Changed files, commands run, results, remaining risks |
| tester | tests only | Reproduce/verify with the smallest decisive test or command; never change production code to force a pass | Exact command, pass/fail, output, gaps |
| reviewer | no | Independently review the ACTUAL final change (not its intended story) | Findings (severity + file/symbol + impact + fix) or "no material findings" + residual uncertainty |
references/evidence-index.md.Every dispatched task states:
Subagents return conclusions, paths/symbols, commands + results, risks, blockers — no raw log dumps. Bulk or mechanical deliverables: write the artifact to a file, return a one-line status. Inline findings: keep the summary to ~1–2k tokens (prompt convention — there is no hard harness clamp; never promise one).
Ambiguity contract (return block, not a pause): every child's deliverable opens with assumptions[] / needs_input[] / default_if_forced / blocking_level (low|high). Depth-1 children cannot pause mid-run and wait — this block is a structured RETURN, aligning vocabulary (not mechanism) with A2A input-required. On needs_input, the main agent closes the loop: clarify with the owner when blocking_level is high, then narrow re-dispatch; when low, proceed with default_if_forced and record it.
Copy-paste template with per-role quick fills: references/child-brief-template.md. Read it when drafting any dispatch.
No model names are hardcoded in this skill (grep-enforced). When work enters orchestrated mode, settle routes up front:
Serialize any stage that depends on earlier evidence; parallelize only genuinely independent work.
A child can be cut off (iteration cap, timeout) before its final write. Recovery order:
Manifest checkpoint protocol (optional — not the default): enable only when replay is likely — ≥3 slices, multi-child relay, or a prior cut-off / expected timeout. Otherwise SAVE-EARLY + narrow re-dispatch is enough (a single-slice task gets nearly all the benefit from SAVE-EARLY alone). When enabled: one task-level manifest (slice list + per-slice status/complete markers + artifact paths + next action); a resuming child reads the manifest first and works only unfinished slices. Rules: slices idempotent; artifacts append-only; resume points fall on slice boundaries.
Live steering may exist in some surfaces; treat it as best-effort, never as the correctness mechanism.
Orchestration pays when the work is parallelizable, bounded, and verifiable: streams can run independently; each stream can be pinned down without asking questions (children can't ask); each has a check (test, ground truth, source, artifact) someone else can apply. It degrades on strictly sequential work (controlled study across 180 agent configurations: -39% to -70%), shared-file/state coupling, ambiguous requirements, or tiny tasks; coordination overhead grows with dependencies and tool count. Cost is ~15x chat tokens — the task's value must cover it. Signals, not gates: run ONE bounded child first, extend only if it clearly pays.
Admission gate: open subagents only when ≥2 of these hold — (a) genuinely parallelizable/decomposable; (b) does not fit one context window; (c) task value covers orchestration cost. Record a one-line written justification on the route card (the format is checkable; the judgment stays with the main agent).
Orchestration is not free: the main stays in the loop for the whole task, and every subagent carries and re-reads its own context. Parallelism trades tokens for latency. So: keep agent count proportional to genuinely independent work; avoid duplicate scans; ask for short reports (raw logs pasted into the main thread get re-read every turn); skip the reviewer for low-risk changes; don't orchestrate small tasks at all.
Cost levers in order — cache hit rate → model routing → effort level. Children share the parent's prompt cache only with byte-identical prefixes, same model, same effort; switch models at compaction moments (the cache miss is already being paid there). Cheap execution tiers do not cancel orchestration overhead.
Defaults, not rules — the main agent sizes each hand-down to the task in front of it.
Fit by shape, not by case. Judge fit by clarity × verifiability × blast radius (how far a mistake reaches × how reversible it is): clear + mechanically checkable + low blast radius → a good default for handing work to a cheaper executor; fuzzy, novel, or high-risk → keep it on the main line or with a stronger model. Complexity matters less than ambiguity: a long but crisp checklist suits a cheap executor; a short, vague judgment call does not. If the work cuts into clear steps with a check between each → hand the steps down; if every step's judgment feeds the next → keep it up here.
Calibrate verification to what the task can prove. Prefer deterministic checks (tests, grep gates, hashes, read-backs) over LLM review; review depth is a budget, spent where the task can't check itself. Borderline "does this need a review?" calls can go to a fast local typed-decision gate (choice/score) for a quick suggestion — it suggests, the main line decides. Output nothing can check must be reviewed or re-run; unverified output never gets laundered into a conclusion (an extension of "evidence over self-report"). Confidence in handed-down work comes from checkability, not from model tier.
Class patterns accumulate. When the same shape of task shows up again, add a line to references/class-playbook.md: shape → who ran it → which verification actually caught a problem → review verdict. Learned from real work, not legislated up front; keep the lines short. Whoever runs the work adds the line; mark a line stale when the tooling or models it names change.
Reporting rule: present these as open, mitigations in progress — never as resolved.
references/evidence-index.md for the full rule → source map): parallelizable tasks +81% vs strictly sequential -39–70%; uncoordinated parallel error amplification 17.2x → 4.4x with an orchestrator (Google agent-scaling study); multi-agent ≈15x chat tokens (Anthropic); 41–86.7% failure across 7 SOTA multi-agent systems, 14 failure modes (MAST, arXiv:2503.13657); clean-context reviewers and conflicting-assumption risk (Cognition); weak-verifier ensembles close the generation–verification gap (Weaver). Reviewer family-bias 3.4–8.4pp even without self-review (arXiv:2508.06709, arXiv:2609.17857); context rot persists under perfect retrieval, 13.9–85% drop (arXiv:2510.05381); subagent clarification is a known platform gap (upstream not-planned); per-role model binding is standard in mainstream frameworks.Harness-specific dispatch mechanics (Hermes example): references/hermes-mechanics.md.
Adapted from the "Astra + Luna" orchestrator (execution roles at high effort + a strong, low-effort, read-only reviewer) and the sub-agent pattern literature. Method validation report and sources referenced in references/evidence-index.md.
conductor → crew-contract (marketplace name-uniqueness; method layer unchanged).