T05 · Unauthorized Access and Privilege Escalation
- Location
references/eks-pod-action-guide.md:91- Finding
Cluster-wide administrator access granted to the RBAC provisioning Lambda
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a Review item because it can automatically change AWS infrastructure and leaves overbroad or persistent EKS permissions, though I found no evidence of hidden exfiltration or backdoor behavior.
Install only for approved sandbox or tightly controlled AWS accounts. Before use, review the generated CloudFormation template and IAM policies, require explicit approval for deploy/delete steps, add CloudWatch stop conditions, use a dedicated CloudFormation service role, and avoid EKS pod-action workflows until cluster-admin Lambda access and retained shared RBAC are narrowed or lifecycle-managed.
references/eks-pod-action-guide.md:91Cluster-wide administrator access granted to the RBAC provisioning Lambda
references/eks-pod-action-guide.md:40Persistent shared EKS RBAC is neither ownership-validated nor removed
The documented behavior includes stack inspection, output-directory renaming, and post-deployment processing that go beyond simple configuration preparation. While less severe than direct resource mutation, this still represents hidden operational behavior and can confuse users about what the skill will do on local filesystems and remote AWS state.
The documented behavior includes stack inspection, output-directory renaming, and post-deployment processing that go beyond simple configuration preparation. While less severe than direct resource mutation, this still represents hidden operational behavior and can confuse users about what the skill will do on local filesystems and remote AWS state.
The documented behavior includes stack inspection, output-directory renaming, and post-deployment processing that go beyond simple configuration preparation. While less severe than direct resource mutation, this still represents hidden operational behavior and can confuse users about what the skill will do on local filesystems and remote AWS state.
Referenced artifact was not completely inspected
- `references/msk-guide.md` — Amazon MSK (broker reboot via SSM
Referenced artifact was not completely inspected
- `references/msk-guide.md` — Amazon MSK (broker reboot via SSM
Referenced artifact was not completely inspected
- `references/msk-guide.md` — Amazon MSK (broker reboot via SSM
Referenced artifact was not completely inspected
- `references/msk-guide.md` — Amazon MSK (broker reboot via SSM
The script goes beyond preparing or generating AWS FIS experiment configuration and actively performs CloudFormation deployment and teardown operations. In the context of a skill advertised as configuration preparation, this introduces unauthorized infrastructure mutation and can create or remove resources in a live AWS account, especially dangerous if invoked by an agent without clear user confirmation.
On deployment failure, the script automatically deletes the CloudFormation stack and waits for deletion to complete, which is a destructive action that can remove partially created resources and erase evidence needed for debugging or recovery. In a skill whose stated purpose is only to prepare FIS experiments, automatic cleanup materially increases the blast radius because a malformed or adversarial template could trigger infrastructure changes followed by silent teardown.
The README explicitly says the skill will automatically deploy CloudFormation stacks, delete failed stacks, and retry with self-healing. In a skill that creates IAM roles, Lambda custom resources, FIS templates, tagging resources, and SSM automation, this behavior can cause real infrastructure mutations and repeated destructive actions without a strong upfront warning, confirmation gate, or blast-radius notice. The context makes this more dangerous because the skill operates directly against live AWS accounts and prepares fault-injection infrastructure by design.
The README explicitly says the skill will deploy CloudFormation and automatically delete/retry failed stacks, which goes beyond a user-safe interpretation of merely preparing or generating configuration. In an agent setting, this scope expansion is dangerous because users may invoke the skill expecting read/write local artifact generation, while the skill performs live infrastructure mutations in AWS with potentially broad IAM effects.
The document describes infrastructure-modifying behavior such as CloudFormation deployment, deletion, and auto-repair, but does not lead with a strong safety warning or consent boundary. In a high-privilege AWS context, lack of prominent warning increases the chance of accidental destructive changes and confused-deputy execution by an agent on behalf of a user.
This section documents validate/deploy/delete/retry behavior, including automatic cleanup of failed stacks, which materially changes the operational risk of the skill. Automatic remediation loops can repeatedly mutate infrastructure and destroy partially created resources, increasing blast radius if the generated template or target selection is wrong.
Automatic management of EKS RBAC via Lambda-backed custom resources extends the skill into Kubernetes authorization changes, not just FIS template preparation. Because RBAC objects are shared and not removed on stack deletion, this can leave persistent privilege artifacts in clusters that outlive the experiment setup.
The workflow claims to prepare or generate an experiment configuration, but it also mandates deploying CloudFormation stacks and modifying templates until deployment succeeds. This is dangerous because users invoking a seemingly preparatory skill may trigger real infrastructure changes, including repeated mutation attempts, without a clearly separated approval boundary.
The skill instructs automatic CloudFormation deployment and deletion-on-failure behavior without an explicit warning at the top that it will modify AWS resources. Hidden mutation is dangerous because it can create, alter, and tear down infrastructure unexpectedly, potentially affecting availability, audit trails, and cost.
The 'never deliver untested configuration' rule forces deployment and validation as part of normal operation, expanding the skill from document generation into an execution pipeline. That coupling is risky because safety review, cost review, and blast-radius review can be bypassed when testing is treated as mandatory and automatic.
Defaulting stop conditions to 'none' for fault injection experiments means experiments may continue without alarm-based interruption or rollback triggers. In a chaos-testing context, this increases the chance that a mis-scoped or misconfigured experiment causes prolonged disruption before operators manually intervene.
The instruction to self-heal on CloudFormation errors and 'Do not ask the user to fix CFN errors' grants the skill broad autonomy to diagnose, modify templates, delete failed stacks, and retry up to five times. That autonomy is risky in infrastructure workflows because the agent may make incorrect remediation decisions or compound damage without human review.
template and iterate until it succeeds (Step 6). The user should receive a
working, deployed experiment template ready to start.
- **Self-heal on CFN errors.** Read stack events, diagnose, fix the template,
delete the failed stack, retry. Do not ask the user to fix CFN errors.
- **Verify FIS action availability** (`aws fis list-actions` /
`aws fis get-action`) before generating templates. Don't fabricate action
IDs.
This guide instructs preparation of a highly disruptive AZ power interruption scenario but does not place a clear, explicit user-facing safety warning at the point where actions are selected and configuration is generated. In an operational chaos-engineering skill, that omission increases the risk that a user will generate or later run broad-impact fault injection against production-like resources without fully appreciating blast radius, prerequisites, and rollback expectations.
The base template normalizes a default of StopConditions: - Source: 'none' for FIS experiments without prominently warning that fault-injection actions may continue until natural completion even if service health degrades. In a chaos-engineering skill that helps users generate disruptive AWS FIS experiments, this increases the chance operators deploy unsafe experiments lacking an emergency abort mechanism, which can worsen outages or prolong service disruption.
The guide enumerates destructive and stress-inducing EKS pod fault actions but does not pair that scope with a clear warning that these actions intentionally disrupt workloads and should only be run in approved test environments with rollback/monitoring in place. In a chaos-engineering/FIS skill, such actions are expected, but omission of impact warnings increases the chance of accidental production disruption or misuse by less experienced operators.
The cleanup section provides manual kubectl delete commands for shared RBAC resources and notes they should be removed only after all experiments are gone, but it does not strongly warn that deleting them prematurely will break other active or future experiments in the namespace. Because the names are fixed and shared, an operator may mistakenly remove permissions relied on by unrelated stacks, causing experiment failures or operational confusion.
The guide instructs users how to perform disruptive ElastiCache fault-injection actions such as AZ power interruption, primary reboot, and failover testing, but it does not place a clear upfront safety warning that these actions can reduce availability, trigger failovers, and should not be run casually against production systems. In a chaos-engineering skill, that omission is security-relevant because users may treat the workflow as routine automation and launch experiments without adequate environment scoping, approval, or blast-radius controls.
The guide instructs users to reboot an MSK broker as a fault injection action but does not clearly warn that this intentionally causes temporary service degradation, broker unavailability, leader re-election, and possible client-visible disruption. In a chaos engineering context this action is expected, but omitting an explicit operational warning increases the risk that a user runs it against production or without validating resilience prerequisites, causing avoidable availability impact.
Detected: suspicious.generated_source_template_injection