T09 · Insecure Skill Coding Practices
- Location
scripts/schema_extractor.py:48- Finding
Unsanitized Entity Names in Generated Cypher and RDF/Turtle
- Content
View full analysis
1: # Normalize: remove possessives, plurals normalized = clean_word.rstrip("'s") if not normalized.endswith("s"): entities.add(normalized) else: # Try singular form entities.add(normalized[:-1]) ``` ```python def to_cypher_labels(self) -> str: """Generate Neo4j CREATE CONSTRAINT statements.""" output = [] for entity in sorted(self.entities): output.append(f"CREATE CONSTRAINT ON (n:{entity}) ASSERT n.id IS UNIQUE;") return "\n".join(output) ``` ```python def to_turtle_rdf(self) -> str: """Generate basic RDF/Turtle schema.""" output = ["@prefix ex: ."] output.append("") # Define classes for entity in sorted(self.entities): output.append(f"ex:{entity} a rdfs:Class ;") output.append(f" rdfs:label \"{entity}\" .") output.append("") # Define relationships as properties rel_types = set(r.relation_type for r in self.relationships) for rel in sorted(rel_types): output.append(f"ex:{rel} a rdf:Property ;") output.append(f" rdfs:label \"{rel.replace('_', ' ')}\" .") return "\n".join(output) ``` ### Technical Analysis `extract_entities()` accepts entity names derived from untrusted text while removing only periods, commas, and semicolons. Other syntax-significant characters—including quotes, backticks, parentheses, colons, brackets, slashes, and control characters—are not rejected or encoded. The accepted entity names are sub ...[truncated 1996 chars]- Remediation
View remediation
str: if not IDENTIFIER_RE.fullmatch(value): raise ValueError(f"Invalid schema identifier: {value!r}") return value ``` 2. **Separate display labels from machine identifiers** Generate a normalized identifier for Cypher labels and RDF resource names while retaining the original text only as a display label. 3. **Apply format-specific encoding** Do not assume one sanitization routine is safe for both formats: - Use Neo4j-supported identifier quoting and escaping where dynamic labels are unavoidable. - Escape RDF string literals correctly. - Percent-encode or otherwise safely construct RDF IRIs and prefixed names. - Prefer an established RDF library such as `rdflib` rather than manually concatenating Turtle syntax. 4. **Validate before serialization** Perform a second validation step inside `to_cypher_labels()` and `to_turtle_rdf()` so callers cannot bypass extraction and directly modify `self.entities`. 5. **Limit downstream database privileges** Execute generated schema statements only through a database account with the minimum required schema permissions. Never use an administrative account for unreviewed generated output. 6. **Add adversarial tests** Include tests covering quotes, backticks, parentheses, brackets, colons, slashes, newlines, Unicode control characters, and excessively long identifiers. Tests should verify that unsafe identifiers are rejected or safely encoded and that the resulting Cypher and Turtle parse as intended. ]]>
