Natural-Language Policy Violations
Medium
- Confidence
- 93% confidence
- Finding
- The document embeds specific residential community names tied to real budget examples, while relying only on a prose instruction not to reveal them in downstream outputs. In an LLM-based skill, such natural-language guardrails are brittle: the model may quote, summarize, or leak these identifiers when asked for examples, calibration details, or source provenance, causing unintended disclosure of identifiable location-linked information.
