Context-Inappropriate Capability
Medium
- Confidence
- 92% confidence
- Finding
- The manifest describes this skill as verifying math-heavy code for algorithmic correctness and numerical stability, which implies analytical review. However, the workflow instructs the agent to run shell commands like git inspection, pytest with benchmarks, and Jupyter notebook execution, adding code-execution capability that is not clearly justified by the stated purpose alone.
