Back to skill

Security audit

Knowledge Graph - Kg Schema From Text

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent knowledge-graph schema helper with a small code-quality risk in generated Cypher/RDF output, but no evidence of hidden access, exfiltration, persistence, or destructive behavior.

Reasonable to install for local schema-drafting use. Treat its generated Cypher and RDF/Turtle as drafts: review or validate identifiers before running them against Neo4j or importing them into graph tooling, especially when the source text is untrusted.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/schema_extractor.py:48
Finding

Unsanitized Entity Names in Generated Cypher and RDF/Turtle

Content
View full analysis
1: # Normalize: remove possessives, plurals normalized = clean_word.rstrip("'s") if not normalized.endswith("s"): entities.add(normalized) else: # Try singular form entities.add(normalized[:-1]) ``` ```python def to_cypher_labels(self) -> str: """Generate Neo4j CREATE CONSTRAINT statements.""" output = [] for entity in sorted(self.entities): output.append(f"CREATE CONSTRAINT ON (n:{entity}) ASSERT n.id IS UNIQUE;") return "\n".join(output) ``` ```python def to_turtle_rdf(self) -> str: """Generate basic RDF/Turtle schema.""" output = ["@prefix ex: ."] output.append("") # Define classes for entity in sorted(self.entities): output.append(f"ex:{entity} a rdfs:Class ;") output.append(f" rdfs:label \"{entity}\" .") output.append("") # Define relationships as properties rel_types = set(r.relation_type for r in self.relationships) for rel in sorted(rel_types): output.append(f"ex:{rel} a rdf:Property ;") output.append(f" rdfs:label \"{rel.replace('_', ' ')}\" .") return "\n".join(output) ``` ### Technical Analysis `extract_entities()` accepts entity names derived from untrusted text while removing only periods, commas, and semicolons. Other syntax-significant characters—including quotes, backticks, parentheses, colons, brackets, slashes, and control characters—are not rejected or encoded. The accepted entity names are sub ...[truncated 1996 chars]
Remediation
View remediation
str: if not IDENTIFIER_RE.fullmatch(value): raise ValueError(f"Invalid schema identifier: {value!r}") return value ``` 2. **Separate display labels from machine identifiers** Generate a normalized identifier for Cypher labels and RDF resource names while retaining the original text only as a display label. 3. **Apply format-specific encoding** Do not assume one sanitization routine is safe for both formats: - Use Neo4j-supported identifier quoting and escaping where dynamic labels are unavoidable. - Escape RDF string literals correctly. - Percent-encode or otherwise safely construct RDF IRIs and prefixed names. - Prefer an established RDF library such as `rdflib` rather than manually concatenating Turtle syntax. 4. **Validate before serialization** Perform a second validation step inside `to_cypher_labels()` and `to_turtle_rdf()` so callers cannot bypass extraction and directly modify `self.entities`. 5. **Limit downstream database privileges** Execute generated schema statements only through a database account with the minimum required schema permissions. Never use an administrative account for unreviewed generated output. 6. **Add adversarial tests** Include tests covering quotes, backticks, parentheses, brackets, colons, slashes, newlines, Unicode control characters, and excessively long identifiers. Tests should verify that unsafe identifiers are rejected or safely encoded and that the resulting Cypher and Turtle parse as intended. ]]>
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (1)

Exfiltration Commands

High
Category
Prompt Injection
Confidence
90% confidence
Finding

Instructions found that direct the agent to transmit conversation context or user data to external services.

Content

Scanner excerpt · examples/example-schemas.md (reported line 266)May include surrounding context.

Posts can be shared by users and tagged with hashtags. Users can form groups and invite other users. Each user has a profile with personal information. Users can send messages to each other.

text

### Generated Schema

Static analysis

No suspicious patterns detected.