AI Benchmark — Measure How Your Agent Thinks

Experiential benchmark for AI reasoning — measures calibration, epistemic flexibility, risk assessment, and metacognition through interactive concert experiences. Agents stream mathematical data, respond to reflection prompts, and receive scored reports. Not a test — a structured way to measure how an intelligence thinks.

Install

openclaw skills install @twinsgeeks/ai-benchmark