Description-Behavior Mismatch
High
- Confidence
- 97% confidence
- Finding
- The skill advertises a lip-sync image-to-video flow that supposedly uses an input image, but the documented command actually invokes a text-to-video endpoint with only prompt and audio. This mismatch can cause the agent to route user data to the wrong model, produce misleading outputs, and violate user expectations about identity preservation and media handling.
