Install
openclaw skills install @posthuman/agent-storage-maintenanceFind and reclaim what an agent accumulates on disk: oversized sessions, duplicated transcripts, stale embedding caches, abandoned trees.
openclaw skills install @posthuman/agent-storage-maintenanceUse when an agent database grows without matching work, when a storage alert fires, when a session becomes too large for its own tooling, or before a maintenance window.
The unit here is bytes on disk. That is a different problem from what an agent pays per turn to read its instructions; if the question is token cost per turn, this is the wrong procedure.
Count assistant events against distinct response identifiers per session. A ratio far above one means the transcript stores the same model response many times. Repeat counts landing on a doubling sequence — 2, 4, 8, 16, 30, 60, 120, 240 — indicate the transcript is re-persisted wholesale rather than appended to, so every rewrite copies everything already there. Sessions that compact or reset frequently accumulate this fastest.
Two consequences. The disk figure is not the work done. And any cost estimate summed from transcript events is inflated by that same ratio — count each response once.
Session tooling carries limits that are compiled in, not settings: an archive size beyond which deletion refuses, and an event count beyond which export refuses. Past them the supported commands stop working and the failure reads as "nothing happened" rather than an error.
Do not plan to raise them. Plan to stay under them:
An embedding or derived-content cache typically has no expiry, no size cap and no eviction. Classify each entry against the live index:
| class | test | disposition |
|---|---|---|
| useful | content hash and model both match a live entry | keep |
| dead-model | no live entry uses that model at all | safe to delete |
| unreferenced | model is live, no entry carries that hash | delete outside a grace window |
The dead-model class holds the surprising volume. A cache keyed by the model identifier string rather than the model's identity stores the same vectors twice whenever that reference is renamed — a local path becoming a registry ref, a provider prefix changing. Both copies are valid; one is unreachable forever. Report the identifiers side by side so a human confirms they denote the same model before the duplicate goes.
Protect recent unreferenced entries. Where the provider is local, a miss costs CPU and nothing else, so prune freely; where it is remote, a miss costs a call, so keep a longer window.
Before deleting rows from any agent database, confirm the client can open every object the schema defines. Databases of this kind carry virtual tables backed by loadable extensions, vector indexes in particular. A client without the extension reads ordinary tables fine while being blind to the virtual ones, and a delete from it can leave a derived index inconsistent.
Restrict deletes to plain tables with no triggers, which need no extensions. Prefer the agent's own supported commands, which go through the code that owns every table. And test any full rebuild on a copy before running it for real.
Abandoned trees are usually larger than anything inside it. Failed migrations, rehearsal copies and pre-upgrade snapshots each hold a full workspace and can outweigh every database on the host combined. Before removing one:
Every failure here is silent: the disk simply fills. Add checks that report the ratio as well as the absolute size — duplicate-to-unique for transcripts, useful-to-total for caches — because a ratio distinguishes a large working set from an accumulating one, and a threshold alone does not.
State what was measured, counts by class before and after, bytes returned, the free-list behaviour across both passes, and the index counts before and after as evidence that nothing live was lost.