ponytail-gain

Show ponytail measured impact as a scoreboard: less code, less cost, more speed, from the agentic benchmark averages. One-shot display.

Install

openclaw skills install @dietrichgebert/ponytail-gain

Ponytail Gain

Display this scoreboard when invoked. One-shot: do NOT change mode, write flag files, or persist anything.

The figures are the published agentic benchmark of Ponytail 5: headless Claude Code (Opus 5.5, default effort) on 39 tasks (feature tickets in a real FastAPI + React repo, bug fixes, security and privacy cases, small apps), 5 runs each, against the same agent without the skill. 18 of the tasks have hidden checks for correctness and safety. Each figure is the geometric mean of the per-task medians. They are measured, not computed from the current repo. Source: benchmarks/results/2026-10-07-agentic.md and the README.

Scoreboard

Render plain ASCII bars. The bar length shows ponytail as a share of the no-skill baseline; the label carries the exact figure:

text
  ponytail gain           benchmark · 39 tasks × 5 runs · Opus 5.5

  no-skill        ████████████████████  100%
  Lines of code   █████████···········   47%   ▼ 53%
  Output tokens   ███████████·········   55%   ▼ 45%
  Cost            ███████████████·····   74%   ▼ 26%
  Time            ████████████········   59%   ▼ 41%
  Hidden checks passed   97%  (no-skill 96%)
  Tests where the logic needs one   98%  (no-skill 68%)

  This repo:  /ponytail-debt  (shortcuts you deferred)
              /ponytail-audit (what's still cuttable)

Honesty boundary

These are benchmark averages, not this repo. NEVER print a per-repo savings number ("you saved X lines/tokens here"): the unbuilt version was never written, so there is no real baseline to subtract from in a live repo. The only real per-repo figures come from /ponytail-debt (a counted ledger), and this card points there instead of inventing one.

Boundaries

One-shot display. Edits nothing, changes no mode. "stop ponytail" or "normal mode": revert.