Tp4
High
- Category
- MCP Tool Poisoning
- Confidence
- 96% confidence
- Finding
- The description says this skill is for rehearsing an OpenClaw update and producing compatibility-gap and migration/rollback/post-start verification planning based on package evidence and Kova reports. The supplied code does not do upgrade rehearsal or migration planning. Instead, it is a benchmark runner: it verifies a frozen corpus, executes benchmark arms over fixture prompt payloads, optionally calls a configured advisory adapter, scores the outputs against frozen scoring keys, and emits evaluation artifacts. The module repeatedly states it is 'evaluation-only' and that 'canonical_status_effect' is none. While it is a sibling to rehearsal tooling and imports a rehearsal-related error class, its primary purpose is materially different from the declared description.
