Back to skill

Security audit

agent-cognitive-states

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a coherent agent self-monitoring aid, but it under-controls persistent memory/logging and can cause sensitive conversation details, including credentials, to be saved.

Review before installing. Use this only if you are comfortable with agent memory and local logs, and do not let it persist credentials, tokens, private keys, session cookies, personal data, or sensitive environment details. If using the cron guardian, keep it user-scoped, protect log files, add retention limits, and document how to remove the scheduled job.

Vulnerability Patterns
  • System PersistenceInstalls backdoors, hooks, services, or scheduled tasks that survive the run
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
references/detection-heuristics.md:68
Finding

Sensitive credentials may be persisted as cognitive-state memory

Content
View full analysis
Remediation
View remediation

T06 · System Persistence

Note
Location
templates/guardian-cronjob.yaml:27
Finding

Optional guardian creates recurring execution without lifecycle or log-retention safeguards

Content
View full analysis
> /var/log/agent-cognitive.log 2>&1 ``` ```yaml # 5. Logs full report to file for post-hoc analysis log_file: "~/.agent-cognitive-states.log" ``` ### Technical Analysis The template openly documents an optional task that runs every ten minutes. Recurring execution is related to the declared Active Guardian feature and is not concealed; the template also does not install the task automatically. Consequently, this is not evidence of a malicious backdoor. However, registration creates execution that survives the original Skill run and potentially the original Agent session. The project provides no corresponding disable or uninstall procedure, ownership guidance, least-privilege requirements, or retention policy. The crontab example also appends all output to `/var/log/agent-cognitive.log`. This conflicts with the template’s statement that healthy execution is silent because `scripts/self_check.py` always prints a human, Markdown, or JSON report before exiting. Thus, following the example produces a log entry every ten minutes regardless of whether an alert threshold is reached. No rotation or size cap is configured. The `/var/log` destination may require elevated setup permissions on many systems. The template does not itself elevate privileges, and no privilege-escalation mechanism was found, but recommending a system log location can encourage deployment with broader permissions than the local monitoring ...[truncated 1716 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (9)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The code substantially aligns with the declared purpose: it detects and reports multiple cognitive-state degradation conditions including context fatigue, attention drift, memory debt, confidence erosion, and skill staleness. It does not perform risky undeclared actions or access sensitive resources. However, there is a mild but material description/behavior mismatch in two respects: first, the script includes an additional cognitive-state detector, Context Fragmentation, that is not mentioned in the declared description. Second, although the description says 'detect, report, and mitigate,' the implementation only detects and reports, with mitigation represented as textual recommendations rather than actual remediation. Because the primary purpose is still close, this is a limited mismatch rather than a severe one.

Content

No source excerpt is available for this finding.

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · scripts/self_check.py (reported line 209)May include surrounding context.

python
signal = "; ".join(signal_parts)
    impact = "Following outdated instructions produces errors and wasted effort"
    action = "Patch skill with corrected commands/paths before continuing"
    
    return CognitiveState(
        name="Skill Staleness",

Ssd 3

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This section normalizes persisting 'critical facts,' decisions, and other conversation details to memory before they are lost. In a self-monitoring skill, that context makes the issue more dangerous because the retention is framed as routine reliability behavior, which can lead agents to store sensitive user information without necessity, minimization, or consent.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

Directing the agent to write a brief session summary to memory or a file introduces session persistence of potentially sensitive conversational content. In this skill's context, persistence is presented as a standard mitigation step, which increases the likelihood of storing confidential requirements, decisions, or environment details without clear boundaries or user approval.

Content

Scanner excerpt · SKILL.md (reported line 208)May include surrounding context.

md
### Context Fatigue → Consolidate & Split
1. Persist all critical facts, decisions, and TODO state to memory
2. Write a brief session summary to memory or a file
3. Suggest `/new` or session split for remaining work
4. If user wants to continue: prioritize ruthlessly, ignore tangent topics

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs the agent to scan recent exchanges for preferences, decisions, corrections, and environment facts, then batch-write them to memory. That creates a concrete over-collection pattern: broad semantic extraction from conversation into persistent storage, increasing the risk of retaining sensitive personal, project, or environmental data beyond what is needed.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
72% confidence
Finding

The template establishes recurring execution of a script that reads active session state and, in the non-Hermes example, derives parameters from a session log under /tmp and writes results to a persistent log file. This creates a session-persistence/data-retention surface: session metadata may be repeatedly accessed and stored beyond the immediate interaction, which can expose operational context or sensitive user-derived data if log locations or temporary files are not adequately protected.

Content

Scanner excerpt · templates/guardian-cronjob.yaml (reported line 7)May include surrounding context.

yaml
# during autonomous/long-running work and alerts the user if degraded.
#
# Usage with Hermes Agent cronjob system:
#   hermes cron create --schedule "every 10m" --script scripts/self_check.py
#
# Or as a YAML config for other cron systems:

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly recommends appending cognitive state data to a local log file without warning about persistence, retention, or access controls. Even if the example payload is small, operational use can easily expand to include conversation-derived state, timestamps, and workflow details that may be sensitive on shared systems or developer machines.

Content

No source excerpt is available for this finding.

Ssd 3

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

Passive logging without interrupting the conversation encourages silent retention of state data and normalizes background collection. While the shown example is limited to cognitive-state metadata, the surrounding skill context promotes persistence of conversation-derived information, which makes silent logging a privacy and auditability concern.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

This manifest-style YAML template describes usage with a cron system but does not clearly constrain when this skill should be activated beyond generic recurring scheduling. Because it lacks explicit trigger boundaries or exclusion conditions, it could be adopted in broader contexts than intended for autonomous monitoring.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.