Context-Inappropriate Capability
Medium
- Confidence
- 91% confidence
- Finding
- The benchmarking example actively executes an arbitrary binary supplied as tool_path via subprocess.run. In an auditing/benchmarking skill, this expands scope from passive analysis into code execution, which can run untrusted programs and trigger side effects on the host system if the path is attacker-controlled or insufficiently constrained.
