Back to skill

Security audit

Knowledge Graph - NL To Graph Query Translator

Security checks for vulnerabilities and agentic risk

Overview

This is a graph-query translator, not malware, but it needs Review because it encourages direct use of executable database queries without enough safeguards against destructive or sensitive queries.

Install only if you will treat generated queries as draft text. Do not wire this skill to production database execution unless queries are reviewed, parsed with a real read-only policy, run under least-privilege read-only credentials by default, and mutating/admin clauses require separate authorization and confirmation.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/query_validator.py:66
Finding

Fail-Open Query Validation Allows Mutating Cypher and Unsupported Languages

Content
View full analysis
QueryValidation: """Validate a graph query.""" issues = [] if self.language == "cypher": issues = self._validate_cypher(query) elif self.language == "sparql": issues = self._validate_sparql(query) # Generate optimization suggestions suggestions = self._generate_optimization_suggestions(query, issues) # Estimate execution time est_time = self._estimate_execution_time(query) is_valid = not any(i.status == ValidationStatus.ERROR for i in issues) return QueryValidation( query=query, language=self.language, is_valid=is_valid, issues=issues, optimization_suggestions=suggestions, estimated_execution_time_ms=est_time ) def _validate_cypher(self, query: str) -> List[ValidationIssue]: """Validate Cypher query.""" issues = [] lines = query.split('\n') # Check for MATCH clause if not any('MATCH' in line.upper() for line in lines): issues.append(ValidationIssue( status=ValidationStatus.ERROR, message="Missing MATCH clause", suggestion="Cypher queries must start with MATCH, WITH, or UNWIND" )) # Check for RETURN clause if not any('RETURN' in line.upper() for line in lines): issues.append(ValidationIssue( status=ValidationStatus.ERROR, message="Missing RETURN clause", suggestion="Cypher queries must have a RETURN clause" )) # Check for LIMIT (optional but recommended) if not any('LIMIT' in line.upper() for line in lines): if not any('COUNT' in line.upper() for line in lines): issues.append(ValidationIssue( status=Validat ...[truncated 4012 chars]
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (12)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared purpose centers on query generation from natural language, but the supplied code chunk is a validator/optimizer for already-written graph queries. Its primary functions are checking for clauses like MATCH/RETURN or SELECT/WHERE, warning about traversal/cardinality issues, suggesting optimizations, estimating execution time, and modifying queries to add LIMIT or bounded traversals. These are materially different capabilities from NL-to-Cypher/SPARQL translation, so this is a clear description-behavior mismatch.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The architecture explicitly extends a natural-language query translator into an API-facing system that executes generated database queries against Neo4j and RDF stores. That expansion increases the trust boundary and attack surface substantially: untrusted user input can flow through translation into live query execution, creating risk of unauthorized data access, query injection, and expensive or destructive operations if execution is enabled without strict controls.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

Documenting direct query execution for a skill whose stated purpose is only translation creates capability creep and can mislead integrators into granting broader permissions than necessary. In this context, the mismatch is dangerous because a component expected to produce strings may instead be deployed with database connectivity and execution privileges, magnifying the impact of translation errors or prompt-driven manipulation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file includes healthcare-domain queries that retrieve identifiable patient treatment and medication information, such as a named patient and prescribed medications. The document presents these examples without any warning or disclosure about handling sensitive medical data, which fits the markdown-specific missing-warning criterion for privacy-impacting behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The finance-domain examples expose analysis of account holders, transactions, intermediaries, fraud rings, and credit exposure. Because the markdown description omits any warning that these examples may involve sensitive financial data and access controls, it lacks user disclosure for privacy- and integrity-relevant behavior.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language prompt says the example should show all employees under Alice's management, direct and indirect, and the returned levels_down implies the query computes path depth. However, the Cypher snippet only matches subordinate and never binds a path or relationship collection named relationships, so the documented intent contradicts what the code as written can do.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

This API reference repeatedly presents the output as an executable Cypher/SPARQL query and shows direct printing/use of result.query, but it does not warn users that generated queries should be reviewed before execution. In a skill whose purpose is translating natural language into database queries, omission of that warning increases the likelihood that consumers will automatically run model-generated queries, which can cause unintended reads, expensive operations, or data modification if unsafe query forms are produced or accepted downstream.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The reference guide includes CREATE, MERGE, SET, and DELETE examples even though the skill is described as translating natural-language questions into queries, which implies read-oriented behavior. In an NL-to-query setting, including write-capable examples increases the chance that an agent or downstream model will generate destructive or state-changing Cypher against a live database without the user appreciating the risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The markdown presents data-creation and deletion examples without any warning that these statements modify live graph data. In a translator skill, this can normalize unsafe output and lead users or agents to run mutating queries in production, causing unintended data corruption or deletion.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The example loads spaCy's en_core_web_sm, which implicitly fixes the pipeline to English. Because this is natural-language guidance in a markdown file and no user opt-in or explicit English-only scope is provided nearby, it can violate the language/locale policy requirement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The example FILTER (LANG(?label) = "en") enforces an English-only result set in natural-language content. Because the document does not note that this is merely an example or offer a locale choice, it can be read as endorsing a fixed-language constraint without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The DBpedia endpoint example filters abstracts to "en", which again imposes an English-only locale choice in instructional text. Under the policy, fixed language selection should be user-selectable or explicitly justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.