Back to skill

Security audit

Agent Loop Engineering

Security checks for vulnerabilities and agentic risk

Overview

This skill is an advanced autonomous coding workflow tool that is disclosed, bounded to project-local work, and includes clear stop rules.

Install this only if you want an agent to continue authorized coding work with project-local edits and verification without asking about every reversible step. Keep the target, acceptance criteria, write scope, and protected boundaries explicit, and do not use it on repositories containing secrets, production/customer data, or tasks involving paid resources, public deployment, destructive Git, or account sessions unless you are ready to stop and approve those gates manually.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (8)

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · references/en/migration.md (reported line 37)May include surrounding context.

md
- Stop routine writes to duplicate status journals.
- Keep one current Work Order only when it owns stable scope.
- Keep one final independent QA decision for Standard/Full work.
- Archive only at a natural release/month boundary; do not delete history.

If authority fingerprint changes, stop execution and align before refreshing the packet.

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · scripts/state-tools-lib.mjs (reported line 293)May include surrounding context.

js
export function atomicWriteInside(rootReal, destination, content) {
  const target = assertInside(rootReal, destination, "write path")
  if (existsSync(target)) throw new Error(`Refusing to overwrite existing file: ${target}`)
  const parent = assertInside(rootReal, dirname(target), "write parent")
  if (!statSync(parent).isDirectory()) throw new Error(`Write parent is not a directory: ${parent}`)
  const temporary = assertInside(rootReal, join(parent, `.${basename(target)}.${process.pid}.${Date.now()}.tmp`), "temporary path")

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list includes broad, everyday phrases such as 'keep going' and 'continue where we left off' that can match ordinary conversational intent rather than an explicit request to invoke an autonomous coding loop. In this skill’s context, unintended activation is risky because the skill is designed to continue execution autonomously, make local decisions, and perform verification loops, which could cause the agent to take actions the user did not clearly authorize.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
78% confidence
Finding

The skill explicitly instructs the agent to continue by default and make ordinary reversible decisions 'without asking the Owner to supervise each loop,' which delegates meaningful autonomy to the agent. Although bounded by policy text, this still increases the chance of unauthorized or mistaken actions if scope, authority, or environmental conditions are misinterpreted, especially when combined with broad triggers and resumable execution.

Content

Scanner excerpt · SKILL.md (reported line 10)May include surrounding context.

md
Version: 2.2.0

Use this skill as the execution plane for authorized software work. Continue by default while useful progress remains inside scope. Make ordinary reversible project-local decisions, diagnose failures, repair them, and verify real behavior without asking the Owner to supervise each loop.

Respond in the user's language. Keep persistent state factual, compact, and free of hidden reasoning.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill enables implicit invocation while providing only broad, high-level usage language, which increases the chance the agent will auto-activate this autonomous coding loop in situations that are not sufficiently scoped or authorized. Because this skill is designed to continue work autonomously with persistent state and repair behavior, accidental invocation can expand execution beyond user intent and reduce meaningful human confirmation at the point of action.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script emits detailed workspace metadata, files_read, docs_directory, output_path, conflicts, warnings, and route/authority analysis to stdout when --json is used, and also prints summary paths in human-readable mode. In an agent skill context, stdout is commonly captured into logs, orchestration traces, chat transcripts, or telemetry, so this can disclose sensitive repository structure and authority file paths without an explicit disclosure gate or redaction step.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The document presents all instructions in English and Chinese as a fixed policy format, but it does not state that language selection is optional or explain a region-specific requirement for this locale behavior. Under the policy rule, fixed language or locale handling without opt-in or justification can be a natural-language policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The file is explicitly a zh-CN template and the visible section headings are bilingual with a Chinese locale-specific path, but the template does not state that Chinese output is optional, user-selected, or required for a region-specific purpose. This can be a natural-language locale policy concern because it may implicitly steer skill behavior toward a fixed language without documented opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.