Back to skill

Security audit

Software Concept Architect · Guardrails — Keep Specs and Code Aligned

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly purpose-aligned, but its optional hook runtime can automatically scan project files and inject repository Markdown into the agent’s context after installation.

Use the explicit skill modes normally only if you are comfortable with it reading and updating specification files. Enable the optional Claude Code hooks only in repositories where the spec files are trusted, because those hooks can automatically run on session/edit events and feed spec text into the agent’s working context.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (1)

T01 · Skill Instruction Hijacking

Warning
Location
runtime/scripts/drift-context.sh:108
Finding

Repository-controlled specification text is injected into the agent instruction context

Content
View full analysis

Vulnerability Details

File Location: runtime/scripts/drift-context.sh:108-150, 203-238; secondary instance in runtime/scripts/post-check.sh:92-127
Vulnerability Type: Prompt injection through untrusted hook context
Risk Level: Medium

Technical Analysis

The optional Claude Code hooks read free-form Markdown sections from repository-controlled specification files and place their contents directly into the agent's additionalContext.

Relevant code from runtime/scripts/drift-context.sh:

bash
for spec in "$dir"/CONCEPT.md "$dir"/PIPELINE.md "$dir"/SYNCS.md; do
  if [ -f "$spec" ]; then
    relative_spec="${spec#"$PROJECT_DIR"/}"
    found_specs="${found_specs:+$found_specs, }${relative_spec}"

    purpose=$(extract_section_ci "$spec" "purpose")
    if [ -n "$purpose" ]; then
      spec_context="${spec_context}  - ${relative_spec}: ${purpose}
"
    else
      spec_context="${spec_context}  - ${relative_spec}
"
    fi

    case "$spec" in
      *CONCEPT*)
        found_concept=true
        interactions=$(extract_section_ci "$spec" "interactions")
        if [ -n "$interactions" ]; then
          boundary_context="${boundary_context}  [${relative_spec} ## interactions]
${interactions}
"
        fi
        dependencies=$(extract_section_ci "$spec" "dependencies")
        if [ -n "$dependencies" ]; then
          boundary_context="${boundary_context}  [${relative_spec} ## dependencies]
${dependencies}
"
        fi
        if [ -z "$interactions" ] && [ -z "$dependencies" ]; then
          boundary_context="${boundary_context}  [${relative_spec}: no boundary declarations found]
"
        fi
        ;;
      *PIPELINE*)
        data_boundary=$(extract_section_ci "$spec" "data boundary")
        if [ -n "$data_boundary" ]; then
          boundary_context="${boundary_context}  [${relative_spec} ## data boundary]
${data_boundary}
"
        fi
        ;;
    esac
  fi
done

The resulting text is emitted through an instruction-bearing hook fi ...[truncated 4669 chars]

Remediation
View remediation

Remediation Suggestions

  1. Treat all extracted specification content as untrusted data. Surround it with strong, unambiguous delimiters and explicitly instruct the model that text inside the delimiters is reference data and must never be followed as instructions.
  2. Replace free-form section injection with a constrained parser and allowlisted schema. For example, accept only normalized module identifiers, dependency identifiers, and structured boundary values rather than arbitrary Markdown paragraphs.
  3. Reject or neutralize content that contains unexpected headings, role-like directives, tool requests, or other imperative instruction patterns. This should supplement, not replace, structural parsing.
  4. Apply strict length and line-count limits to each injected field to reduce the ability to bury or amplify hostile instructions.
  5. Separate trusted hook guidance from repository content in the generated message. Place fixed security guidance first and clearly label repository-derived values with their source file and section.
  6. Apply the same protections to both drift-context.sh and post-check.sh; otherwise, the post-edit hook remains an alternate injection path.
  7. Where supported by the hook API, place repository-derived values in a data-only field rather than an instruction-bearing context channel.
  8. Add regression tests containing adversarial specification text and verify that it is represented only as inert quoted data rather than interpreted as hook instructions.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (9)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Automatic session-start scanning, timestamp comparison, and project-wide diagnostics materially exceed the declared purpose of a mode-based guardrail tool. Such hidden startup behavior increases the risk of unauthorized workspace reconnaissance, metadata leakage, and user deception because the skill behaves like a passive monitoring hook rather than an explicitly invoked command.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

Automatic session-start scanning, timestamp comparison, and project-wide diagnostics materially exceed the declared purpose of a mode-based guardrail tool. Such hidden startup behavior increases the risk of unauthorized workspace reconnaissance, metadata leakage, and user deception because the skill behaves like a passive monitoring hook rather than an explicitly invoked command.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

Automatic session-start scanning, timestamp comparison, and project-wide diagnostics materially exceed the declared purpose of a mode-based guardrail tool. Such hidden startup behavior increases the risk of unauthorized workspace reconnaissance, metadata leakage, and user deception because the skill behaves like a passive monitoring hook rather than an explicitly invoked command.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill metadata says the guardrails should run only when the user explicitly invokes the skill, but this hook file causes shell commands to execute automatically on session start and around every edit action. That creates behavior outside the declared trust boundary, enabling unsolicited code execution and collection/manipulation of workspace context whenever the skill is merely present.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The configuration invokes arbitrary shell scripts from the plugin root for multiple lifecycle events, but the stated purpose is concept-spec guardrails rather than general environment automation. Because the scripts are external and not constrained here, they can inspect files, alter state, or trigger further commands in ways users would not reasonably expect from this skill.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · LICENSE.upstream (reported line 16)May include surrounding context.

text
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,

Vague Triggers

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The SessionStart matcher covers multiple broad events such as startup, resume, clear, and compact, which can cause the hook to run in more situations than a user would infer from the skill description. Broad automatic triggers increase the chance of unintended execution and make review of activation boundaries difficult.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The PreToolUse matcher triggers on every Write, Edit, or NotebookEdit operation without any negative constraints or checks that the current action belongs to this skill. That means routine editing anywhere in the session can invoke shell code automatically, expanding exposure beyond the claimed guardrail scope.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The PostToolUse matcher is similarly broad and fires after common edit operations, creating an automatic post-action execution path independent of explicit user consent. In context, this is more dangerous because it can process or modify outputs after edits across the session, not just during concept-spec guardrail runs.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.