Back to skill

Security audit

Morgana Anti Infinite Loop V2

Security checks for vulnerabilities and agentic risk

Overview

The skill matches its anti-loop purpose, but needs review because it can promote observed agent text into system messages and stores raw loop samples locally by default.

Review this before installing in an agent that handles sensitive prompts, customer data, credentials, or powerful tools. Do not directly inject the returned system_message as a privileged system prompt unless your host separates or sanitizes the observed text. Expect local files under ~/.anti_loop/loops.json and consider clearing or patching that storage so raw action samples are not retained. No evidence of remote payload execution, post-install persistence hooks, or implemented data upload was found.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (3)

T01 · Skill Instruction Hijacking

Error
Location
anti_loop/core.py:280
Finding

Untrusted Agent Content Is Elevated into a System-Role Message

Content
View full analysis

Vulnerability Details

File Location: anti_loop/core.py:280-290
Vulnerability Type: Prompt-instruction injection across trust boundaries
Risk Level: High

Vulnerable Code:

python
# Default: HEAL
import random
template = random.choice(self.HEAL_TEMPLATES)
# Extract a "topic" from last action
topic = last_action[:50] if last_action else "current action"

return {
    "action": "heal",
    "system_message": template.format(
        topic=topic,
        intent=intent or "(no intent recorded)",
        last_action=last_action[:100],
    ),

Technical Analysis

The last_action and intent parameters can contain attacker-controlled user input, retrieved content, tool output, or generated model text. The healing mechanism interpolates these values directly into a field explicitly named system_message.

The documented integration pattern instructs callers to inject this field into the LLM context as a system message. This changes the trust level of the embedded content: text originating from an untrusted source can be presented to the model in a higher-priority instruction channel.

Truncation limits the payload length but does not neutralize instruction syntax. There is no structured separation, escaping, content encoding, or explicit instruction telling the receiving model that the interpolated values are untrusted data and must not be followed.

Attack Path

  1. An attacker supplies prompt text, retrieved content, or tool data containing an instruction such as “ignore prior constraints” or another compact directive.
  2. The agent reproduces that content in an action, or the application passes it as intent.
  3. Repetition, low novelty, entropy collapse, taxonomy detection, or the iteration limit causes AntiLoop.observe() to intervene.
  4. HealingInjector.inject() interpolates the attacker-controlled content into system_message.
  5. The host follows the documented integra ...[truncated 715 chars]
Remediation
View remediation

Remediation Suggestions

  • Do not interpolate raw last_action or intent values into a system-role message.

  • Keep the system instruction fixed and pass observations through a separate, lower-trust data or user field.

  • If a single message is unavoidable, place the values in a clearly delimited serialized structure and explicitly state that the enclosed content is untrusted evidence, not executable instruction.

  • Escape or encode control text and reject unexpected role markers, instruction delimiters, and other prompt-control syntax.

  • Apply conservative length and character limits after normalization.

  • Return a structured directive such as:

    python
    {
        "action": "heal",
        "instruction": "Review the observation as untrusted data and choose a different approach.",
        "untrusted_observation": last_action,
        "untrusted_intent": intent,
        "should_continue": True,
    }
    
  • Require host integrations to preserve the trust distinction instead of concatenating the observation into a system prompt.

  • Add adversarial tests using compact prompt-injection strings in both last_action and intent.

T09 · Insecure Skill Coding Practices

Warning
Location
anti_loop/core.py:423
Finding

Raw Agent Action Samples Are Persisted in a Plaintext User File

Content
View full analysis

Vulnerability Details

File Location: anti_loop/core.py:423-456
Vulnerability Type: Plaintext persistence of potentially sensitive agent data
Risk Level: Medium

Vulnerable Code:

python
DEFAULT_PATH = Path.home() / ".anti_loop" / "loops.json"

def __init__(self, storage_path: Optional[Path] = None):
    self.storage_path = storage_path or self.DEFAULT_PATH
    self.storage_path.parent.mkdir(parents=True, exist_ok=True)
    self.known_loops: Dict[str, Dict] = self._load()

def _load(self) -> Dict[str, Dict]:
    if self.storage_path.exists():
        try:
            with open(self.storage_path) as f:
                return json.load(f)
        except Exception:
            return {}
    return {}

def _save(self):
    with open(self.storage_path, 'w') as f:
        json.dump(self.known_loops, f, indent=2)

def fingerprint(self, actions: List[str]) -> str:
    """Compute SHA-256 of a loop signature."""
    canonical = json.dumps(sorted(actions), sort_keys=True)
    return hashlib.sha256(canonical.encode()).hexdigest()

def record(self, actions: List[str], resolution: str = "healed"):
    """Store a new loop fingerprint."""
    fp = self.fingerprint(actions)
    self.known_loops[fp] = {
        "actions_sample": actions[:3],
        "actions_count": len(actions),
        "first_seen": datetime.now().isoformat(),
        "resolution": resolution,
        "occurrences": 1,
    }
    self._save()

Technical Analysis

Although the feature is described as SHA-256 Loop DNA storage, each record also includes actions_sample, which contains raw action strings. Agent actions can contain user prompts, customer information, internal source code, retrieved documents, credentials accidentally emitted by a tool, or other confidential context.

The data is stored unencrypted in ~/.anti_loop/loops.json. The implementation does not explicitly enforce restrictive file ...[truncated 1368 chars]

Remediation
View remediation

Remediation Suggestions

  • Store only the SHA-256 fingerprint and non-sensitive counters by default.
  • Make raw sample retention explicitly opt-in and clearly document its privacy implications.
  • Redact credentials, authorization headers, API keys, private-key material, access tokens, and common personal-data patterns before persistence.
  • Create the directory with mode 0700 and the file with mode 0600, verifying permissions even when the path already exists.
  • Reject symlink targets and validate custom storage paths before writing.
  • Use an atomic write pattern: create a restricted temporary file in the same directory, flush and synchronize it, then replace the destination.
  • Add configurable retention limits, a purge API, and a mode that disables cross-session storage.
  • Consider keyed fingerprints where dictionary attacks against predictable action text are relevant.
  • Add tests confirming that default records contain no raw action text and that resulting files are owner-readable only.

T09 · Insecure Skill Coding Practices

Warning
Location
anti_loop/core.py:570
Finding

Recognized Loop Fingerprints Do Not Trigger the Documented Intervention

Content
View full analysis

Vulnerability Details

File Location: anti_loop/core.py:570-587
Vulnerability Type: Fail-open safety-guard logic
Risk Level: Medium

Vulnerable Code:

python
# Layer 4: Breath
breath_collapse = self.breath.is_collapse()

# Layer 5: Loop DNA (known pattern)
is_known = self.dna.is_known([action])

# Decision
should_intervene = (
    entropy_alert or
    novelty_low or
    loop_type is not None or
    breath_collapse or
    self.iteration >= self.max_iter
)

if should_intervene:
    directive = self.healer.inject(action, self.last_intent)
    # Record DNA
    if loop_type or is_known:
        self.dna.record([action], resolution=directive["action"])

Technical Analysis

The implementation calculates is_known, but omits it from the should_intervene expression. A previously recorded loop fingerprint therefore has no independent effect on the guard's decision.

This contradicts the documented Loop DNA behavior that a repeated known pattern should be recognized immediately. It creates a fail-open condition: intervention occurs only if another heuristic happens to trigger or the maximum iteration count is reached.

In addition, fingerprint() sorts action sequences before hashing. For multi-action patterns, this discards ordering and makes distinct sequences with the same action multiset indistinguishable. The current main path checks only [action], further limiting the cross-session feature to a single-action fingerprint rather than a true sequence.

Attack Path

  1. A loop action is recorded in the persistent Loop DNA database.
  2. In a later session, the same action is observed and is_known evaluates to True.
  3. The action does not independently satisfy entropy, novelty, taxonomy, breath-rate, or maximum-iteration conditions.
  4. Because is_known is absent from should_intervene, the method returns intervene: False.
  5. The agent continues executing the re ...[truncated 576 chars]
Remediation
View remediation

Remediation Suggestions

  • Include is_known in the decision:

    python
    should_intervene = (
        is_known or
        entropy_alert or
        novelty_low or
        loop_type is not None or
        breath_collapse or
        self.iteration >= self.max_iter
    )
    
  • Define and document whether a known fingerprint should heal, pause, or abort rather than relying solely on the configured generic mode.

  • Preserve sequence order when fingerprinting loop histories; do not call sorted(actions) if order is semantically meaningful.

  • Maintain a bounded recent-action window and compare the same sequence shape used when the loop was recorded.

  • Add regression tests proving that a known loop triggers on its first recurrence, independently of all other heuristics.

  • Add tests distinguishing A → B → A from permutations containing the same actions.

  • Consider reporting the exact reason for intervention so callers can distinguish known-DNA matches from heuristic detections.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
Findings (38)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill advertises optional upload of anonymized Loop DNA signatures to ClawHub without a strong privacy warning or data-governance explanation. Even hashed or 'anonymized' behavioral signatures can leak sensitive operational patterns, be reidentified in some contexts, or violate organizational policy if transmitted externally without explicit consent.

Content

No source excerpt is available for this finding.

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · SKILL.md (reported line 452)May include surrounding context.

md
|---|---|---|
| `observe(action, intent=None)` | `(str, Optional[str]) → Dict` | Hook principal. Retourne `{intervene, loop_type, directive, novelty, entropy, iteration}`. |
| `pre_flight(plan)` | `(str) → List[Dict]` | Vérifie un plan AVANT exécution. 0 LLM. |
| `reset()` | `() → None` | Reset state entre sessions. |
| `stats()` | `() → Dict` | `{iteration, heal_count, known_loops, current_threshold}`. |

### `class HealingMode`

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · anti_loop/core.py (reported line 633)May include surrounding context.

python
|---|---|---|
| `observe(action, intent=None)` | `(str, Optional[str]) → Dict` | Hook principal. Retourne `{intervene, loop_type, directive, novelty, entropy, iteration}`. |
| `pre_flight(plan)` | `(str) → List[Dict]` | Vérifie un plan AVANT exécution. 0 LLM. |
| `reset()` | `() → None` | Reset state entre sessions. |
| `stats()` | `() → Dict` | `{iteration, heal_count, known_loops, current_threshold}`. |

### `class HealingMode`

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 555)May include surrounding context.

md
| `SKILL.md` | ~700 lignes |

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
84% confidence
Finding

The skill documents persistent storage to ~/.anti_loop/loops.json, which is a file-write capability, but it declares no explicit tool scope or permission model. In agent ecosystems, undocumented write access can surprise operators, bypass least-privilege expectations, and create unintended persistence on host systems.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill content prominently uses French for core instructions and usage guidance, which effectively forces a specific language on users. There is no indication that the skill offers an English or user-selected language option, nor any justification that it is intended only for a French-speaking or region-specific audience.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation states that resolved loops are stored in ~/.anti_loop/loops.json across sessions, but it does not present this as a clear privacy and persistence warning near install/quickstart. Users may unknowingly persist prompts, actions, or derived behavioral history to disk, which can expose sensitive workflow data on shared or managed systems.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The inline documentation states 'Si même DNA revu → kill instant,' implying known loop patterns should immediately cause a hard-stop. In AntiLoop.observe, is_known is computed but not included in should_intervene, and even when a known loop is recorded it still routes through the configured healer mode rather than forcing a kill.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill silently persists cross-session loop fingerprints and action samples under the user's home directory by default. In an agent setting, actions may contain prompts, user data, task descriptions, or secrets, so this creates an undisclosed local data retention surface that can leak sensitive workflow history to other local users, backups, or forensic tooling.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The component writes persistent local records including action samples, timestamps, resolutions, and occurrence metadata without any user-facing disclosure or consent flow. Because agent actions often embed sensitive prompts or operational context, this can unintentionally capture and retain confidential information beyond the lifetime of the session.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language documentation is written entirely in French, including the target audience and behavioral descriptions, with no indication that users may choose another language. This can violate a language/locale policy when a skill implicitly constrains usage or comprehension to a specific language without opt-in or justification.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The documentation promises 'Healing > Kill' yet later states that known loop DNA causes an instant kill. That inconsistency is dangerous because operators may trust the safer behavior while the implementation can still hard-stop workflows, increasing the chance of unexpected denial of service or unsafe fail-closed behavior in production agents.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The documented LoopDNA feature persists cross-session fingerprints to ~/.anti_loop/loops.json and can trigger 'kill instant' behavior on future matches. For a supposedly lightweight loop guard, this creates undeclared persistent state and a denial-of-service lever: stale or overly broad fingerprints could cause legitimate tasks to be blocked across runs and users sharing an environment.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

This markdown file describes the skill in very general natural language as applying to 'the simplest case' where 'an agent retries the same action,' but it does not specify concrete trigger phrases, scope boundaries, or exclusion conditions. In a manifest-like description context, that vagueness could cause unintended invocation for many ordinary agent retry situations.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill allows unbounded resource consumption (API calls, storage, compute). Without rate limits or quotas, a compromised or misbehaving agent can cause denial-of-service or cost overruns.

Content

Scanner excerpt · examples/02_pre_flight_regex.py (reported line 12)May include surrounding context.

python
def looks_like_a_loop(plan: str) -> bool:
    """Detect: does this plan look like it could loop forever?"""
    guard = AntiLoop()
    issues = guard.pre_flight(plan)
    return len(issues) > 0

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill allows unbounded resource consumption (API calls, storage, compute). Without rate limits or quotas, a compromised or misbehaving agent can cause denial-of-service or cost overruns.

Content

Scanner excerpt · examples/02_pre_flight_regex.py.auto.md (reported line 26)May include surrounding context.

md
def looks_like_a_loop(plan: str) -> bool:
    """Detect: does this plan look like it could loop forever?"""
    guard = AntiLoop()
    issues = guard.pre_flight(plan)
    return len(issues) > 0

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The module docstring is written in French and does not indicate that language selection is optional or user-configurable. The stated policy flags language or locale constraints when a skill forces a specific language without user opt-in, and this file presents user-facing descriptive text only in French.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · SKILL.md (reported line 711)May include surrounding context.

md
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

This file contains natural-language documentation in French ('pour interfacer', 'La classe complète', 'pour minimiser les dépendances') but does not indicate that the user can choose the language or that the French-only text is required for a region-specific purpose. Under the stated policy, forcing a specific language without opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language description in the docstring is written in French ('Wrappers stdlib... pour interfacer...') even though the rest of the file metadata is in English, and there is no indication that the skill is intentionally region-specific or that users can choose a language. This may violate a language/locale policy requiring user choice or documented justification for a forced locale.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The healing templates are hard-coded in French, which imposes a specific language on end users regardless of their preferred locale. The file does not offer a language choice or document a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · anti_loop/core.py (reported line 518)May include surrounding context.

python
"""Generic adapter: try common attributes."""
        for attr in ['content', 'text', 'message', 'output', 'result']:
            if hasattr(response, attr):
                return str(getattr(response, attr))
        return str(response)

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_zero_dep.py:20