Back to skill

Security audit

The Agent Incident Response Playbook: Detect, Contain, Recover, and Learn When AI Agent Systems Fail

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent incident-response guide, but its examples can automatically freeze budgets and cancel escrows using production API credentials with insufficient safeguards.

Review carefully before installing or following the examples. Use a sandbox first, pin and verify any dependency, create least-privilege GreenHelix keys scoped to specific agents and actions, add webhook signature verification and replay protection, and avoid enabling auto-containment or fleet shutdown until approval gates and rollback procedures are in place.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:2342
Finding

Unauthenticated Webhook Can Trigger Privileged Financial Containment

Content
View full analysis
= 95: anomaly = AnomalyResult( detected=True, signal="budget_exhaustion", severity="critical", agent_id=agent_id, details=payload, ) responder.auto_contain(agent_id, anomaly) elif utilization >= 80: responder.oncall_callback( severity="WARNING", message=f"Budget at {utilization}% for {agent_id}", ) elif event_type == "reputation_change": score_delta = float(payload.get("score_delta", 0)) if score_delta < -0.15: anomaly = AnomalyResult( detected=True, signal="reputation_drift", severity="critical" if score_delta < -0.30 else "warning", agent_id=agent_id, details=payload, ) responder.escalate( agent_id, anomaly, EscalationTier.AUTO_CONTAIN if score_delta < -0.30 else EscalationTier.ALERT_HUMAN, ) return jsonify({"status": "processed"}) ``` ### Technical Analysis The webhook endpoint accepts an arbitrary JSON request and trusts its `event_type`, `agent_id`, and `payload` fields without authenticating the sender. It does not verify a GreenHelix web ...[truncated 2364 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:2553
Finding

Unpinned Third-Party Package Installation Creates Supply-Chain Exposure

Content
View full analysis
Remediation
View remediation
``` 2. Publish a lock file or hashed requirements file and install with hash enforcement: ```text pip install --require-hashes -r requirements.txt ``` 3. Record hashes for all transitive dependencies, not only the top-level package. 4. Document the expected package publisher and trusted package index. Where practical, configure pip to use only an organization-controlled or explicitly trusted index. 5. Verify package provenance and signatures where supported, and scan downloaded distributions before deployment. 6. Install and run the dependency in an isolated virtual environment or container under a non-privileged account. 7. Scope `GREENHELIX_API_KEY` to the minimum required operations and agents so that a dependency compromise has limited impact. 8. Use automated dependency monitoring and require security review before updating the pinned package or lock file. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (14)

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly presents itself as non-executable educational material, but later includes runnable code that imports libraries, reads environment credentials, and instructs the user to run it. This mismatch can cause users or downstream systems to treat the skill as inert documentation when it actually facilitates authenticated external actions, increasing the risk of accidental execution with real credentials.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

This code establishes authenticated outbound communication to an external API endpoint using a bearer token. In context, that means any user who runs the example may transmit credentials and operational data to a third-party service, which is security-relevant and more dangerous because the skill markets the examples as practical incident tooling.

Content

Scanner excerpt · SKILL.md (reported line 191)May include surrounding context.

md
self,
        api_key: str,
        agent_id: str,
        base_url: str = "https://api.greenhelix.net/v1",
        cost_anomaly_std_devs: float = 2.0,
        reputation_drop_threshold: float = 0.15,
    ):

External Transmission

Medium
Category
Data Exfiltration
Confidence
87% confidence
Finding

This section again configures authenticated requests to an external GreenHelix API. Because the associated class performs containment actions, execution can send sensitive identifiers and trigger remote state changes, making the external transmission operationally significant.

Content

Scanner excerpt · SKILL.md (reported line 610)May include surrounding context.

md
self,
        api_key: str,
        agent_id: str,
        base_url: str = "https://api.greenhelix.net/v1",
    ):
        self.api_key = api_key
        self.agent_id = agent_id

External Transmission

Medium
Category
Data Exfiltration
Confidence
86% confidence
Finding

The recovery class transmits authenticated requests to an external API to restore budgets, submit metrics, and clear incidents. These are legitimate operations, but they still constitute credentialed external transmission with potential to alter production financial controls if run carelessly.

Content

Scanner excerpt · SKILL.md (reported line 971)May include surrounding context.

md
self,
        api_key: str,
        agent_id: str,
        base_url: str = "https://api.greenhelix.net/v1",
    ):
        self.api_key = api_key
        self.agent_id = agent_id

External Transmission

Medium
Category
Data Exfiltration
Confidence
86% confidence
Finding

The forensics class sends authenticated requests externally to retrieve audit, billing, identity, and claim-chain data. This can expose sensitive operational metadata to an external service and should be treated as deliberate data egress, especially in post-incident contexts where logs and identities are sensitive.

Content

Scanner excerpt · SKILL.md (reported line 1293)May include surrounding context.

md
self,
        api_key: str,
        agent_id: str,
        base_url: str = "https://api.greenhelix.net/v1",
    ):
        self.api_key = api_key
        self.agent_id = agent_id

External Transmission

Medium
Category
Data Exfiltration
Confidence
88% confidence
Finding

This external API configuration underpins fleet-wide automated response, including full shutdown behavior. Because one deployment can affect multiple agents, the external authenticated transmission here has elevated impact if misconfigured, triggered incorrectly, or abused through bad inputs.

Content

Scanner excerpt · SKILL.md (reported line 2106)May include surrounding context.

md
self,
        api_key: str,
        fleet_agent_ids: list[str],
        base_url: str = "https://api.greenhelix.net/v1",
        oncall_callback: Optional[Callable] = None,
    ):
        self.api_key = api_key

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The webhook handler is described as capable of automatically freezing agents and cancelling escrows based on incoming alert data, without strong warnings about the consequences or the trust boundary around webhook events. If deployed as shown, malformed, spoofed, or misclassified events could trigger unauthorized containment actions and operational denial of service across agents.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 2529)May include surrounding context.

md
All companion guides plus this playbook are available as a bundle. Each guide introduces production-ready Python classes that compose into a complete agent commerce platform: building (P1), cost management (P2), reputation (P3), audit trails (P4), trust verification (P5), marketplace strategy (P6), multi-agent patterns (P7), security hardening (P8), testing and observability (P9), SaaS automation (P10), compliance (P11), and incident response (P12).

For the full API reference and tool catalog (all 128 tools), visit the GreenHelix developer documentation at [https://api.greenhelix.net/docs](https://api.greenhelix.net/docs).

---

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill includes copy-pastable code that can freeze budgets, cancel escrows, and alter operational state, yet the code is presented as normal workflow material without an immediate, explicit destructive-action warning. In an incident-response context, operators may paste and run it quickly under stress, causing irreversible or fleet-wide financial disruption if pointed at production instead of a sandbox.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The docstring says the pipeline uses methods like check_reputation_drift(baseline_score) and get_spending_breakdown(), and later code treats detector results as dictionaries via .get(...). Earlier in the guide, IncidentDetector.check_reputation_drift takes no baseline argument, returns an AnomalyResult object, and no get_spending_breakdown method is defined, so the documentation actively misstates what the code relies on.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The inline documentation explicitly states that containment publishes an incident event through IncidentContainment.publish_incident_event. In the earlier implementation of IncidentContainment, no such public method exists; event publication occurs internally via _execute('publish_event', ...), so the comment contradicts the provided code surface.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The docstring says recovery uses IncidentRecovery.update_reputation and IncidentRecovery.publish_recovery_event. Earlier in the guide, the implemented methods are restore_reputation and clear_incident, with no update_reputation or publish_recovery_event methods, so the documentation conflicts with the code that is actually defined.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The docstring says the collector uses IncidentForensics.get_verified_claims() and describes timeline reconstruction based on additional analytics, but the provided IncidentForensics class implements get_audit_trail, assess_impact, build_timeline, and run_post_mortem only. This is an active contradiction between documentation and the code API presented in the same skill.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · SKILL.md (reported line 855)May include surrounding context.

md
# Step 2: Cancel escrows (unlock funds)
        results.append(self.freeze_escrows())

        # Step 3: Verify identity (expand scope if compromised)
        results.append(self.verify_agent_identity())

        elapsed = round(time.time() - start_time, 2)

Static analysis

No suspicious patterns detected.