Install
openclaw skills install @clarezoe/typed-decisions-around-llmsDecide WHERE a typed judgment model (TypeSafe's Jev or any System One model) belongs relative to a generative LLM, and who is allowed to authorize what. Use when adding a guardrail, router, verifier, or approval gate around an LLM; when converting a free-text 'LLM-as-judge' step into a typed decision; when a model's own confidence score is being used to authorize its own output; when choosing thresholds, confidence bands, or fallback behavior; or when reviewing an agent/tool-calling pipeline for who holds authority. This is the architecture question — placement, ordering, and authority. For API mechanics, primitive selection, and question wording use the typesafe-ai skill and the live docs instead.
openclaw skills install @clarezoe/typed-decisions-around-llms"Do not ask one model to be the author, router, policy engine, judge, and auditor of its own work."
That is the whole skill. Everything below is how to take those roles apart.
Scope. This is about placement and authority: what sits before the LLM,
what sits after it, and who is allowed to authorize a side effect. It is not
about API calls. For endpoints, SDKs, primitive selection, question wording,
and criteria design, use the typesafe-ai skill and the live docs; do not
restate them here.
Provenance. Distilled from an independent field guide (Sept 2026) built on public docs. Not vendor material, not an endorsement. It dates itself: verify product names, model aliases, limits, and any number against live documentation before deployment. Hardcode none of them.
| Need | Owner |
|---|---|
| Apply exact rules | Code |
| Judge messy text | Typed judgment model |
| Write or reason openly | LLM |
| Authorize side effects | Policy + human |
| Record and replay | Harness |
The anti-pattern this replaces: one prompt that understands the request, invents a plan, chooses a provider, approves its own tool call, judges its own result, and declares completion. That hides many independent judgments inside one transcript. Pulling them out into typed questions makes them visible, testable, and replaceable.
Boundary test. Use a typed judgment when the answer space is known before the request arrives, the state holds enough evidence, and code can act on the returned distribution. Use an LLM when the answer must be invented, explained, or reached through multi-step reasoning. Use deterministic code when the rule is already explicit.
Where a judgment model does NOT belong: drafting the final answer, generating a patch, synthesizing a long plan, replacing an exact parser, or any call whose result cannot change behavior. A decision layer that produces metrics but never alters routing, review, or context is only added latency.
Mental model: not a manager supervising several LLMs — a typed sensor inside a program. Supervision implies authority; a sensor implies measurement. The program owns the workflow.
A favorable semantic answer is evidence, not authority.
A high probability cannot grant a capability the caller does not possess. Run this order and do not reorder it:
proposal
-> schema and type validation
-> identity and capability check
-> path / domain / amount allowlists
-> data-egress policy
-> semantic review (typed judgment)
-> risk-specific threshold
-> user or operator confirmation
-> execute with least privilege
-> immutable receipt
Two rules an agent will otherwise violate:
"Do not ask the generative LLM to produce a proposal and a confidence score that authorizes the same proposal. Put the bounded review in [a judgment model] or deterministic validation, and let policy decide what is sufficient. Independence is not perfection, but it removes one obvious conflict of interest."
Run this audit on any existing pipeline — it finds real bugs, not hypothetical ones. For every automatic action, ask:
Watch for the disguised form: an LLM emits both a label saying an action needs no approval and the cost estimate that the only remaining threshold compares against. Both inputs to the gate come from the thing being gated.
request
-> deterministic validation
-> preflight battery: classify intent, risk, route
-> code: choose context and LLM
-> LLM: generate plan, text, or code
-> postflight battery: verify bounded properties
-> code: pass, retry, review, or block
-> receipt + metrics
Both batteries are optional and do different jobs. Preflight controls what the expensive model sees and which model gets the task. Postflight evaluates the output against narrow properties before code allows a consequential action. Read-only, low-risk flows may need only one side.
Shape: route (Choice over allowed routes), needs_current_data (Noul),
contains_sensitive_data (Noul), complexity (Score).
Shape, all Noul: addresses_task, evidence_supports, unrelated_changes,
needs_clarification.
Treat state as the packet handed to a review panel — not a transcript dump.
| Include | Avoid |
|---|---|
| Current request | Entire chat by default |
| Relevant policy | Unversioned memory dump |
| Candidate action | Hidden global assumptions |
| Verified records | Secrets not needed to judge |
| Source timestamps | Stale facts without dates |
More context is not better: irrelevant context buys distraction, cost, privacy exposure, and harder evaluation.
"Probability is not permission."
Numbers describe model uncertainty about a bounded question. They do not measure business impact and cannot convert an unauthorized action into an authorized one.
| Band | Default behavior |
|---|---|
| High | eligible for the automatic path |
| Medium | confirm, gather state, or review |
| Low | stop, clarify, or fall back |
The goal is not maximum automatic throughput. It is the best completed outcome per unit of cost, latency, and review attention at an acceptable risk level.
Record one receipt per decision:
state_digest, question_version, jev_model, answers, policy_version,
route, llm_provider, latency_ms, cost_estimate, outcome
| Stage | Behavior |
|---|---|
| Offline | replay labeled cases |
| Shadow | log decisions, change nothing |
| Assist | recommend routes to humans |
| Limited | automate a low-risk slice |
| Expand | raise scope only after evidence |
typesafe-ai).Pre-launch: required fields validated; no-match options present; deterministic rules run before AI; secrets minimized; provider routes approved; failure behavior defined; thresholds evaluated on held-out cases; receipts replayable; rollback tested; review staffed for the volume the policy creates.
First week: review low-confidence cases, automated actions, fallbacks, and provider errors daily. Freeze question wording during the first observation window so behavior changes attribute to traffic, not a moving contract. Make exactly one correction at the end of the week, with a measurable hypothesis.