Back to skill

Security audit

consensus-persona-respawn

Security checks for vulnerabilities and agentic risk

Overview

The skill’s purpose is mostly clear, but it can persist untrusted ledger text into future persona instructions, which needs careful review before installation.

Install only if you trust the consensus state writers for the board and are comfortable with this skill automatically creating persistent replacement personas from ledger history. Review or sanitize decision red_flags before respawn, and prefer a version that validates learned fields, bounds text length, and requires approval before activating generated persona sets.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T02 · Agent Memory Poisoning

Error
Location
src/index.mjs:35
Finding

Persistent persona poisoning through unsanitized decision-ledger content

Content
View full analysis

Vulnerability Details

File Location: src/index.mjs:35-51, 64-77, 132-133
Vulnerability Type: Persistent agent memory poisoning
Risk Level: High

Vulnerable Code:

js
function buildLearningSummary(personaId, decisions){
  const patterns = new Map();
  let considered = 0;
  for (const d of decisions) {
    const votes = d.votes || d.response?.votes || [];
    const final = d.final_decision || d.response?.final_decision;
    const v = votes.find((x)=>x.persona_id===personaId);
    if (!v || !final) continue;
    considered += 1;
    const opposite = (final==='APPROVE'&&v.vote!=='YES') || (final==='BLOCK'&&v.vote!=='NO') || (final==='REWRITE'&&v.vote!=='REWRITE');
    if (opposite && Number(v.confidence||0) > 0.85) patterns.set('high_confidence_mismatch', (patterns.get('high_confidence_mismatch')||0)+1);
    for (const rf of (v.red_flags||[])) patterns.set(`red_flag:${rf}`, (patterns.get(`red_flag:${rf}`)||0)+1);
  }
  return { source_decisions: considered, mistake_patterns: [...patterns.entries()].sort((a,b)=>b[1]-a[1]).map(([k,v])=>`${k}:${v}`) };
}

function mutatePersona(oldPersona, learning){
  const top = learning.mistake_patterns.slice(0,3);
  return {
    ...oldPersona,
    persona_id: `persona_${crypto.randomUUID().slice(0,8)}`,
    name: `${oldPersona.name} v2`,
    bias: `Adjusted from ledger mistakes (${top.join(', ') || 'none'})`,
    non_negotiables: [...new Set([...(oldPersona.non_negotiables||[]), 'Validate high-confidence disagreement'])],
    failure_modes: [...new Set([...(oldPersona.failure_modes||[]), 'Overconfidence without evidence'])],
    reputation: 0.55
  };
}

const pw = await writeArtifact(board_id, 'persona_set', updated, statePath);

Technical Analysis

Decision artifacts are treated as trusted learning material even though their votes[].red_flags[] values are not validated for type, length, permitted vocabulary, or instruction-like ...[truncated 2084 chars]

Remediation
View remediation

Remediation Suggestions

  1. Define and enforce a strict schema for every decision artifact before processing it, including votes, persona_id, vote, confidence, and red_flags.
  2. Require each red flag to be a bounded string from an allowlisted taxonomy, such as stable identifiers rather than arbitrary natural language.
  3. Apply conservative limits to the number of decisions, votes, red flags, and characters processed.
  4. Store learned patterns as structured identifiers and counts. Do not interpolate ledger text directly into persona instruction-bearing fields such as bias or non_negotiables.
  5. Generate display text from trusted templates mapped to allowlisted pattern identifiers.
  6. Track artifact provenance and accept learning inputs only from authenticated, authorized writers for the same board.
  7. Require review or policy approval before activating automatically generated persona sets.
  8. Add adversarial tests covering instruction-like red flags, non-string values, oversized values, repeated poisoning attempts, and cross-board artifact isolation.
  9. Ensure downstream persona consumers treat all learned or historical text as quoted evidence rather than executable instructions.
Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (4)

Known Vulnerable Dependency: fast-uri==3.1.0 — 7 advisory(ies): CVE-2026-13676 (fast-uri vulnerable to host confusion via failed IDN canonicalization); CVE-2026-18446 (fast-uri vulnerable to host confusion via backslash authority introducer); CVE-2026-75975 (fast-uri vulnerable to server-side request forgery via malformed IPv6 normalizat) +4 more

High
Category
Supply Chain
Confidence
87% confidence
Finding

The lockfile includes fast-uri 3.1.0, and the listed advisories indicate URI parsing flaws that can lead to host confusion and potentially SSRF or policy bypass when untrusted URLs are validated or normalized with this library. In this skill's consensus/governance context, dependencies like ajv may process structured external data, so if any code path uses affected URI handling for allowlists, schema formats, or outbound request decisions, the consequences could be significant.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: esbuild==0.27.3 — 1 advisory(ies): GHSA-g7r4-m6w7-qqqr (esbuild allows arbitrary file read when running the development server on Window)

Low
Category
Supply Chain
Confidence
92% confidence
Finding

The lockfile pins esbuild 0.27.3, which is reported as affected by GHSA-g7r4-m6w7-qqqr involving arbitrary file read through the development server on Windows. This is a real supply-chain risk, but in this skill context the package appears to be used as a build/runtime tool via tsx rather than as an exposed dev server, so exploitability is limited unless the skill or its tooling launches esbuild's dev server in a vulnerable Windows environment.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
85% confidence
Finding

The dependency uses a caret range, which allows automatic installation of newer semver-compatible releases. This can increase supply-chain risk because a compromised or malicious upstream release could be pulled into future installs without explicit review or lockstep version control.

Content

Scanner excerpt · package.json (reported line 10)May include surrounding context.

json
"demo": "node --import tsx run.js --input ./examples/input.json"
  },
  "dependencies": {
    "consensus-guard-core": "^1.1.15",
    "tsx": "^4.20.3"
  },
  "license": "MIT",

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
85% confidence
Finding

The tsx dependency is also specified with a caret range, permitting unreviewed semver-compatible updates to be resolved during installation. While common in JavaScript projects, this still creates a supply-chain exposure window if an upstream package release is compromised.

Content

Scanner excerpt · package.json (reported line 11)May include surrounding context.

json
},
  "dependencies": {
    "consensus-guard-core": "^1.1.15",
    "tsx": "^4.20.3"
  },
  "license": "MIT",
  "engines": {

Static analysis

No suspicious patterns detected.