Install
openclaw skills install @himadriganguly/context-engineeringDiagnose and improve context usage in AI agent sessions. Use when context grows large, instructions are being missed, output quality degrades, token usage rises, or a project needs deliberate context organization.
openclaw skills install @himadriganguly/context-engineeringContext is the information an agent can access while working on a task. As a session grows, more instructions, conversation history, tool output, files, and skills may compete for that working context.
This skill provides a structured procedure for diagnosing context pressure, identifying likely failure modes, applying structural fixes, and verifying the result.
The measurements produced by the bundled scripts are heuristics and proxies, not direct measurements of model attention, reasoning quality, or semantic understanding.
Use this skill as decision support, not as an authoritative diagnostic system.
Before running the bundled scripts:
The scripts are designed to read local files. They do not require network access or third-party Python packages.
Load this skill when one or more of these signals appear:
Do not use this skill as a substitute for:
If the runtime exposes a session directory that matches the documented input layout, run:
python scripts/context_audit.py \
--session-dir PATH \
--window-size WINDOW_SIZE
For example:
python scripts/context_audit.py \
--session-dir ./session-fixture \
--window-size 128000
The --window-size value should correspond to the effective context
capacity relevant to the session being analyzed.
Do not assume that 128,000 tokens is universally correct.
The audit reports:
The audit cannot observe information that is not represented in the supplied session directory.
If a runtime does not expose its session data in the expected layout, do not attempt to guess its private storage format.
Use the estimator instead:
python scripts/token_estimator.py \
--dir . \
--recursive \
--top 20
This gives a static estimate of the largest files in a project.
Use the measurements together with the actual task behavior.
The five working categories used by this skill are:
| Class | Typical signal | Possible cause |
|---|---|---|
| Poisoning | Incorrect information continues to influence later work | An incorrect or obsolete statement entered the working context |
| Clash | The agent alternates between incompatible rules | Conflicting instructions or sources |
| Confusion | Old or irrelevant details influence the current task | Stale context remains active |
| Lost-in-middle | Important information becomes harder to retrieve from long context | Relevant information is buried among unrelated material |
| Overload | Context becomes truncated, crowded, or difficult to manage | Too much information is active at once |
These are practical categories, not established clinical or scientific diagnoses.
A session may have more than one category.
When overload is clearly present, reducing unnecessary context is usually a sensible first intervention because it can simplify other problems.
Prefer structural changes over repeatedly rewriting the same prompt.
Load:
references/progressive-disclosure-patterns.md
when deciding how to reorganize context.
Common interventions include:
Do not blindly apply every intervention.
The correct intervention depends on the runtime, model, task, and evidence available.
Create or update:
context-profile.yaml
in the project root.
Use:
templates/context-profile.yaml
as the starting point.
The profile should describe:
The profile is project-specific. Do not place it inside this installed skill unless the project explicitly wants that.
Re-run the audit using the same measurement configuration:
python scripts/context_audit.py \
--session-dir PATH \
--window-size WINDOW_SIZE \
--compare
Look for meaningful changes in:
Then evaluate the actual task result.
A smaller context is not automatically a better context.
A successful intervention should:
Record useful before/after measurements in the project's own documentation when reproducibility matters.
When reorganizing context, prefer:
Move information that remains true across sessions into reference files.
Examples:
When a session becomes difficult to navigate, create a concise brief containing:
Do not preserve every conversational detail merely for completeness.
For large command output, logs, or documents:
A skill description should clearly communicate:
Avoid descriptions that cause unrelated tasks to trigger the skill.
If a runtime provides conditional activation or tool requirements, use them only in runtime-specific documentation or metadata.
Do not assume that one runtime's activation mechanism exists in another.
For long-running tasks, keep the most important constraints in a durable location or use the runtime's supported mechanism for persistent instructions.
Do not repeatedly duplicate the entire instruction set.
A summary that removes a critical constraint can be worse than the original context.
Preserve decisions, constraints, assumptions, and unresolved questions.
A reference file that is loaded for every task is no longer providing much progressive-disclosure benefit.
Load detailed references when the task requires them.
Large command output, logs, generated files, and source files can dominate context.
Estimate their size before repeatedly loading them.
Moving an instruction to a different position may help temporarily.
Prefer reducing irrelevant context, eliminating contradictions, or improving information organization when those are the underlying causes.
A skill itself can contain substantial instructions and references.
Only load skills relevant to the current task.
The audit's instruction-survival measurement is based on textual matching.
The stale-content measurement is based on lexical overlap.
The repetition measurement is based on n-gram counting.
None of these measure semantic understanding, model attention, or reasoning quality directly.
Use the metrics to detect trends and to compare before/after states, not to make absolute claims about a session's quality.
The degradation thresholds suggested by this skill are conservative starting points. They may not fit every model, runtime, or task.
Tighten them for quality-critical work. Loosen them for exploratory work. The correct threshold is the one that catches the failure before it affects output.
A context that is nearly full tends to perform worse than one with room to spare.
Leave room for the model's own reasoning. Do not allocate the last portion of the window to content that is not essential.
The audit reads the session directory it is given. If the directory is incomplete, the measurements will be incomplete. If it contains irrelevant files, the measurements will be skewed.
Always confirm that the supplied directory matches the runtime's actual session state before drawing conclusions.
Load these on demand, as the corresponding step is reached:
| File | Load when |
|---|---|
references/degradation-signals.md | Classifying the problem (Step 2) |
references/progressive-disclosure-patterns.md | Applying a structural fix (Step 3) |
references/token-budget-framework.md | Rebuilding the context profile (Step 4) |
Both scripts use only the Python standard library. They require Python 3.9 or later. No external dependencies are required.
scripts/context_audit.pyProduce a quantitative audit of a session directory.
python scripts/context_audit.py \
--session-dir PATH \
--window-size WINDOW_SIZE
python scripts/context_audit.py \
--session-dir PATH \
--window-size WINDOW_SIZE \
--json
python scripts/context_audit.py \
--session-dir PATH \
--window-size WINDOW_SIZE \
--compare
scripts/token_estimator.pyEstimate the token cost of text, a file, or a directory.
python scripts/token_estimator.py --text "string"
python scripts/token_estimator.py --file path/to/file
python scripts/token_estimator.py --dir . --recursive --top 20
The estimator uses a calibrated character-based heuristic. Treat the results as estimates, not exact token counts.
Five steps for diagnosing and improving context usage:
The procedure is intended to be applied thoughtfully, not mechanically.
Measurements are proxies. Thresholds are defaults. The correct intervention depends on the runtime, model, task, and evidence available.