T09 · Insecure Skill Coding Practices
- Location
SKILL.md:128- Finding
Shell Command Injection Through Untrusted Experiment Metadata
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is for a real AWS fault-injection workflow, but it reaches broadly into EKS clusters and logs and depends on an unpinned runtime skill, so it should be reviewed carefully before installation.
Install only if you are comfortable with an agent using your AWS and Kubernetes permissions to inspect regional EKS workloads and collect logs. Use least-privilege AWS/IAM and Kubernetes RBAC, avoid Secret-read permissions where possible, review the resolved app-service-log-analysis skill before use, validate experiment metadata before running commands, and treat generated reports/logs as sensitive artifacts.
SKILL.md:128Shell Command Injection Through Untrusted Experiment Metadata
SKILL.md:225Overbroad Cross-Cluster Discovery and Kubernetes Secret Inspection
SKILL.md:216Unpinned Runtime Skill Dependency Can Inject Unreviewed Instructions
The referenced kubeconfig handling is security-sensitive because kubeconfig files often contain cluster endpoints, certificate material, role assumptions, and authentication configuration that can facilitate further access if copied or mishandled. Combined with automatic discovery of all accessible clusters and background log collection, the skill broadens credential and environment exposure beyond what is necessary to execute a single prepared experiment.
2. **Reads README.md** to extract experiment metadata (scenario, region, AZ, duration, affected resources, stack name).
3. **Displays experiment actions** — queries the deployed FIS experiment template via AWS CLI to extract and display all action IDs. Log collection is always enabled.
4. **Pre-experiment health check** — verifies every target resource in the FIS template (RDS, MSK, ElastiCache, EKS, EC2, etc.) is in its service's healthy baseline state before any log collection or experiment start. In interactive mode, unhealthy resources require explicit user override. In non-interactive mode, polls every 60 seconds for up to 10 minutes before aborting.
5. **Discovers EKS apps across all clusters and starts log collection** — loads `app-service-log-analysis` skill to discover ALL EKS clusters in the target region, generates isolated kubeconfig files (never overwrites `~/.kube/config`), deep-scans all accessible clusters in parallel for application dependencies (env vars, ConfigMaps, Secrets, ExternalName, etc.), and starts background `kubectl logs -f` **before the experiment starts**. If kubectl is not available, skips app logs but still collects managed service logs via AWS CLI.
6. **Enforces safety** — presents a clear impact warning with affected resources, monitored applications, managed service log status, and post-baseline duration, requires explicit user confirmation before starting.
7. **Starts the experiment** only after explicit user confirmation.
8. **Monitors progress** — polls experiment status every 30-60 seconds, records timestamps for each status change and per-service events. Displays per-app error counts and recovery signals during each poll cycle.
The referenced kubeconfig handling is security-sensitive because kubeconfig files often contain cluster endpoints, certificate material, role assumptions, and authentication configuration that can facilitate further access if copied or mishandled. Combined with automatic discovery of all accessible clusters and background log collection, the skill broadens credential and environment exposure beyond what is necessary to execute a single prepared experiment.
2. **Reads README.md** to extract experiment metadata (scenario, region, AZ, duration, affected resources, stack name).
3. **Displays experiment actions** — queries the deployed FIS experiment template via AWS CLI to extract and display all action IDs. Log collection is always enabled.
4. **Pre-experiment health check** — verifies every target resource in the FIS template (RDS, MSK, ElastiCache, EKS, EC2, etc.) is in its service's healthy baseline state before any log collection or experiment start. In interactive mode, unhealthy resources require explicit user override. In non-interactive mode, polls every 60 seconds for up to 10 minutes before aborting.
5. **Discovers EKS apps across all clusters and starts log collection** — loads `app-service-log-analysis` skill to discover ALL EKS clusters in the target region, generates isolated kubeconfig files (never overwrites `~/.kube/config`), deep-scans all accessible clusters in parallel for application dependencies (env vars, ConfigMaps, Secrets, ExternalName, etc.), and starts background `kubectl logs -f` **before the experiment starts**. If kubectl is not available, skips app logs but still collects managed service logs via AWS CLI.
6. **Enforces safety** — presents a clear impact warning with affected resources, monitored applications, managed service log status, and post-baseline duration, requires explicit user confirmation before starting.
7. **Starts the experiment** only after explicit user confirmation.
8. **Monitors progress** — polls experiment status every 30-60 seconds, records timestamps for each status change and per-service events. Displays per-app error counts and recovery signals during each poll cycle.
The workflow explicitly imports another skill for multi-cluster EKS discovery and kubeconfig isolation before experiment execution. In this context, the danger is not the word 'kubeconfig' itself but the automatic expansion into all accessible clusters and dependency/log inspection, which can collect privileged configuration and operational data unrelated to the intended experiment scope.
Step 4: Discover EKS apps + start log collection [BEFORE experiment]
├── Check kubectl availability
├── kubectl available → load app-service-log-analysis skill, execute its:
│ ├── Multi-Cluster EKS Discovery and Kubeconfig Isolation
│ ├── Step 3 (Deep Scan for application dependencies)
│ ├── Step 3.5 (Managed service log detection)
│ └── Step 4 (Real-time log collection)
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
2. **读取 README.md** 提取实验元数据(场景、Region、AZ、时长、受影响资源、Stack 名称)。
3. **展示实验操作** — 通过 AWS CLI 查询已部署的 FIS 实验模板,提取并展示所有 action ID。日志收集始终启用。
4. **实验前资源健康检查** — 在启动任何日志收集或实验之前,验证 FIS 模板中每个目标资源(RDS、MSK、ElastiCache、EKS、EC2 等)均处于所属服务的健康基线状态。交互式会话下,不健康资源需要用户显式 override;非交互式会话下每 60 秒轮询一次、最长等待 10 分钟后中止。
5. **跨集群发现 EKS 应用并启动日志收集** — 加载 `app-service-log-analysis` skill,发现目标 Region 中所有 EKS 集群,为每个集群生成独立的 kubeconfig 文件(绝不覆盖 `~/.kube/config`),并行深度扫描所有可访问集群的应用依赖(环境变量、ConfigMap、Secret、ExternalName 等),在**实验启动前**启动后台 `kubectl logs -f`。如果 kubectl 不可用,跳过应用日志但仍通过 AWS CLI 收集托管服务日志。
6. **强制安全确认** — 展示清晰的影响警告(受影响资源、监控应用列表、托管服务日志状态、实验后基线时长),要求用户明确确认后才启动。
7. **启动实验** — 仅在用户明确确认后执行。
8. **监控进度** — 每 30-60 秒轮询实验状态,记录每次状态变更和各服务事件的时间戳。每次轮询同时显示各应用错误/警告计数和恢复信号。
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
2. **读取 README.md** 提取实验元数据(场景、Region、AZ、时长、受影响资源、Stack 名称)。
3. **展示实验操作** — 通过 AWS CLI 查询已部署的 FIS 实验模板,提取并展示所有 action ID。日志收集始终启用。
4. **实验前资源健康检查** — 在启动任何日志收集或实验之前,验证 FIS 模板中每个目标资源(RDS、MSK、ElastiCache、EKS、EC2 等)均处于所属服务的健康基线状态。交互式会话下,不健康资源需要用户显式 override;非交互式会话下每 60 秒轮询一次、最长等待 10 分钟后中止。
5. **跨集群发现 EKS 应用并启动日志收集** — 加载 `app-service-log-analysis` skill,发现目标 Region 中所有 EKS 集群,为每个集群生成独立的 kubeconfig 文件(绝不覆盖 `~/.kube/config`),并行深度扫描所有可访问集群的应用依赖(环境变量、ConfigMap、Secret、ExternalName 等),在**实验启动前**启动后台 `kubectl logs -f`。如果 kubectl 不可用,跳过应用日志但仍通过 AWS CLI 收集托管服务日志。
6. **强制安全确认** — 展示清晰的影响警告(受影响资源、监控应用列表、托管服务日志状态、实验后基线时长),要求用户明确确认后才启动。
7. **启动实验** — 仅在用户明确确认后执行。
8. **监控进度** — 每 30-60 秒轮询实验状态,记录每次状态变更和各服务事件的时间戳。每次轮询同时显示各应用错误/警告计数和恢复信号。
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
2. **读取 README.md** 提取实验元数据(场景、Region、AZ、时长、受影响资源、Stack 名称)。
3. **展示实验操作** — 通过 AWS CLI 查询已部署的 FIS 实验模板,提取并展示所有 action ID。日志收集始终启用。
4. **实验前资源健康检查** — 在启动任何日志收集或实验之前,验证 FIS 模板中每个目标资源(RDS、MSK、ElastiCache、EKS、EC2 等)均处于所属服务的健康基线状态。交互式会话下,不健康资源需要用户显式 override;非交互式会话下每 60 秒轮询一次、最长等待 10 分钟后中止。
5. **跨集群发现 EKS 应用并启动日志收集** — 加载 `app-service-log-analysis` skill,发现目标 Region 中所有 EKS 集群,为每个集群生成独立的 kubeconfig 文件(绝不覆盖 `~/.kube/config`),并行深度扫描所有可访问集群的应用依赖(环境变量、ConfigMap、Secret、ExternalName 等),在**实验启动前**启动后台 `kubectl logs -f`。如果 kubectl 不可用,跳过应用日志但仍通过 AWS CLI 收集托管服务日志。
6. **强制安全确认** — 展示清晰的影响警告(受影响资源、监控应用列表、托管服务日志状态、实验后基线时长),要求用户明确确认后才启动。
7. **启动实验** — 仅在用户明确确认后执行。
8. **监控进度** — 每 30-60 秒轮询实验状态,记录每次状态变更和各服务事件的时间戳。每次轮询同时显示各应用错误/警告计数和恢复信号。
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
2. **读取 README.md** 提取实验元数据(场景、Region、AZ、时长、受影响资源、Stack 名称)。
3. **展示实验操作** — 通过 AWS CLI 查询已部署的 FIS 实验模板,提取并展示所有 action ID。日志收集始终启用。
4. **实验前资源健康检查** — 在启动任何日志收集或实验之前,验证 FIS 模板中每个目标资源(RDS、MSK、ElastiCache、EKS、EC2 等)均处于所属服务的健康基线状态。交互式会话下,不健康资源需要用户显式 override;非交互式会话下每 60 秒轮询一次、最长等待 10 分钟后中止。
5. **跨集群发现 EKS 应用并启动日志收集** — 加载 `app-service-log-analysis` skill,发现目标 Region 中所有 EKS 集群,为每个集群生成独立的 kubeconfig 文件(绝不覆盖 `~/.kube/config`),并行深度扫描所有可访问集群的应用依赖(环境变量、ConfigMap、Secret、ExternalName 等),在**实验启动前**启动后台 `kubectl logs -f`。如果 kubectl 不可用,跳过应用日志但仍通过 AWS CLI 收集托管服务日志。
6. **强制安全确认** — 展示清晰的影响警告(受影响资源、监控应用列表、托管服务日志状态、实验后基线时长),要求用户明确确认后才启动。
7. **启动实验** — 仅在用户明确确认后执行。
8. **监控进度** — 每 30-60 秒轮询实验状态,记录每次状态变更和各服务事件的时间戳。每次轮询同时显示各应用错误/警告计数和恢复信号。
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
2. **读取 README.md** 提取实验元数据(场景、Region、AZ、时长、受影响资源、Stack 名称)。
3. **展示实验操作** — 通过 AWS CLI 查询已部署的 FIS 实验模板,提取并展示所有 action ID。日志收集始终启用。
4. **实验前资源健康检查** — 在启动任何日志收集或实验之前,验证 FIS 模板中每个目标资源(RDS、MSK、ElastiCache、EKS、EC2 等)均处于所属服务的健康基线状态。交互式会话下,不健康资源需要用户显式 override;非交互式会话下每 60 秒轮询一次、最长等待 10 分钟后中止。
5. **跨集群发现 EKS 应用并启动日志收集** — 加载 `app-service-log-analysis` skill,发现目标 Region 中所有 EKS 集群,为每个集群生成独立的 kubeconfig 文件(绝不覆盖 `~/.kube/config`),并行深度扫描所有可访问集群的应用依赖(环境变量、ConfigMap、Secret、ExternalName 等),在**实验启动前**启动后台 `kubectl logs -f`。如果 kubectl 不可用,跳过应用日志但仍通过 AWS CLI 收集托管服务日志。
6. **强制安全确认** — 展示清晰的影响警告(受影响资源、监控应用列表、托管服务日志状态、实验后基线时长),要求用户明确确认后才启动。
7. **启动实验** — 仅在用户明确确认后执行。
8. **监控进度** — 每 30-60 秒轮询实验状态,记录每次状态变更和各服务事件的时间戳。每次轮询同时显示各应用错误/警告计数和恢复信号。
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
2. **读取 README.md** 提取实验元数据(场景、Region、AZ、时长、受影响资源、Stack 名称)。
3. **展示实验操作** — 通过 AWS CLI 查询已部署的 FIS 实验模板,提取并展示所有 action ID。日志收集始终启用。
4. **实验前资源健康检查** — 在启动任何日志收集或实验之前,验证 FIS 模板中每个目标资源(RDS、MSK、ElastiCache、EKS、EC2 等)均处于所属服务的健康基线状态。交互式会话下,不健康资源需要用户显式 override;非交互式会话下每 60 秒轮询一次、最长等待 10 分钟后中止。
5. **跨集群发现 EKS 应用并启动日志收集** — 加载 `app-service-log-analysis` skill,发现目标 Region 中所有 EKS 集群,为每个集群生成独立的 kubeconfig 文件(绝不覆盖 `~/.kube/config`),并行深度扫描所有可访问集群的应用依赖(环境变量、ConfigMap、Secret、ExternalName 等),在**实验启动前**启动后台 `kubectl logs -f`。如果 kubectl 不可用,跳过应用日志但仍通过 AWS CLI 收集托管服务日志。
6. **强制安全确认** — 展示清晰的影响警告(受影响资源、监控应用列表、托管服务日志状态、实验后基线时长),要求用户明确确认后才启动。
7. **启动实验** — 仅在用户明确确认后执行。
8. **监控进度** — 每 30-60 秒轮询实验状态,记录每次状态变更和各服务事件的时间戳。每次轮询同时显示各应用错误/警告计数和恢复信号。
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
步骤 4: 发现 EKS 应用 + 启动日志收集 [实验启动前完成]
├── 检查 kubectl 是否可用
├── kubectl 可用 → 加载 app-service-log-analysis skill,执行其:
│ ├── Multi-Cluster EKS Discovery and Kubeconfig Isolation
│ ├── Step 3(深度扫描应用依赖)
│ ├── Step 3.5(托管服务日志检测)
│ └── Step 4(实时日志收集)
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
步骤 4: 发现 EKS 应用 + 启动日志收集 [实验启动前完成]
├── 检查 kubectl 是否可用
├── kubectl 可用 → 加载 app-service-log-analysis skill,执行其:
│ ├── Multi-Cluster EKS Discovery and Kubeconfig Isolation
│ ├── Step 3(深度扫描应用依赖)
│ ├── Step 3.5(托管服务日志检测)
│ └── Step 4(实时日志收集)
The skill is described as only verifying deployed infrastructure and executing prepared FIS experiments, but this reference also includes CloudFormation stack and resource deletion commands. In an agent setting, exposing cleanup commands broadens the operational scope from execution to destructive teardown, creating a real risk that a user prompt, tool-selection bug, or prompt-injection path could cause irreversible deletion of alarms, dashboards, templates, or the entire stack.
The documentation exposes standalone deletion of the experiment template, CloudWatch alarms, and dashboard even though those actions are not required to execute an experiment. That unnecessary capability increases the attack surface and violates least privilege at the documentation/workflow level, making it easier for an agent or operator to perform destructive actions unrelated to the requested task.
The skill states that application and managed-service logs are collected by default, including deep scans of EKS dependencies and background log streaming, but it does not present an explicit privacy or sensitive-data handling warning. These logs and configuration sources can contain secrets, credentials, tokens, PII, or internal topology details, so default collection and report persistence can expose sensitive data beyond the minimum needed for the experiment.
The usage examples include broad triggers such as 'run the chaos experiment' and 'check if the stack is deployed and run the experiment' for a skill that ultimately starts a real AWS FIS experiment against production resources. Even though the README states explicit confirmation is required, ambiguous activation phrases increase the chance the agent invokes this high-impact skill during a routine troubleshooting conversation, creating unsafe escalation toward destructive actions.
Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
10. **Results saved to file.** The experiment results report is written to a timestamped local markdown file, keeping terminal output concise while preserving a full record.
11. **Cleanup is offered, not forced.** After the experiment, cleanup commands are suggested but never executed without confirmation.
## Directory Structure
The example triggers include broad phrases such as '启动 FIS 实验' and '检查 Stack 是否已部署并运行实验', which can plausibly overlap with ordinary discussion or diagnostic requests. In this skill's context, an unintended invocation is more dangerous than usual because the skill is designed to start real AWS Fault Injection Simulator experiments against live resources after only conversational confirmation.
The skill’s stated purpose is to execute an already-prepared FIS experiment, but it additionally requires broad application discovery and log collection across accessible EKS clusters. That expands scope from targeted experiment execution into environment-wide reconnaissance and data collection, increasing the chance of unnecessary access to unrelated workloads and sensitive operational data.
The instruction to discover all EKS clusters in the region, verify access, and deep-scan them in parallel is unjustified for running a single prepared experiment. This creates a clear over-privileged reconnaissance path that can inventory unrelated clusters, applications, and dependencies beyond what the user asked to operate on.
The start-experiment command actively launches a fault injection experiment, but the reference presents it like a routine command without an explicit warning about production impact, blast radius, and prerequisite authorization. In a skill designed for agent use, that omission can materially increase the chance of accidental disruption because the command triggers intentional service faults.
The cleanup section documents destructive commands, including full stack deletion, without strong warnings that the actions are irreversible and will remove operational resources. In an agent-assisted workflow, missing caution text makes accidental execution more likely and reduces the friction that should accompany teardown operations.
The manifest emphasizes running a prepared experiment and confirming infrastructure is already deployed, but does not mention creating local artifacts. Step 10 instructs the skill to write a markdown report directly into the experiment directory, which is additional state-changing behavior beyond the manifest's stated scope.
The polling example hard-codes TZ=Asia/Shanghai for displayed timestamps, which imposes a specific locale/timezone behavior on all users. The policy allows locale constraints only when the user is given a choice or the restriction is clearly documented and justified, neither of which is present here.
Detected: suspicious.generated_source_template_injection