Context-Inappropriate Capability
Medium
- Confidence
- 91% confidence
- Finding
- The skill instructs the agent to use shell execution for benchmark steps that are not strictly necessary to measure generic model performance. Even though the sample commands are simple and scoped, embedding exec capability into a benchmarking workflow normalizes local command execution and expands the attack surface if the skill is reused, modified, or run in a less restricted environment.
