Back to skill

Security audit

huawei-cloud-mrs-spark-sql-check

Security checks for vulnerabilities and agentic risk

Overview

This skill is a local Spark SQL checker whose scripts and rules match its stated purpose and do not show credential access, network transfer, persistence, or destructive behavior.

Install this if you want an offline MRS Spark SQL review helper and are comfortable with Chinese-language rule/report text. Provide only SQL or files you intend the checker to read, because the CLI can read a user-supplied file path and include the analyzed SQL in output.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (17)

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The code clearly supports part of the declared description: syntax-oriented parsing for MRS Spark SQL, statement classification, AST construction, and syntax error reporting. However, the description claims broader capabilities—'comprehensive SQL statement checking,' 'specification compliance,' and 'performance risk detection.' In this chunk, there are no rule-engine checks, no compliance validation against specs, and no performance heuristics beyond minor feature extraction flags like SELECT * or missing WHERE. Those extracted fields could support later analysis, but this code alone does not actually perform the declared comprehensive review behavior. Therefore the description overstates the implemented functionality.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The supplied code is a lexer/tokenizer, not a full SQL checking skill. It tokenizes input into typed tokens and reports lexical errors such as unexpected characters. There is no parser, no rules engine for specification compliance, and no analysis for performance risks. While tokenization could be a supporting component of a larger SQL review tool, this code chunk by itself materially underdelivers relative to the declared description's primary capabilities.

Content

No source excerpt is available for this finding.

Memory Manipulation

High
Category
Memory Poisoning
Confidence
85% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · SKILL.md (reported line 26)May include surrounding context.

md
**Typical Use Cases**:
- "Check this Spark SQL: SELECT * FROM t1"
- "Does this CREATE TABLE USING PARQUET follow Spark specification?"
- "Validate the syntax of this INSERT OVERWRITE statement"
- "Review my Spark SQL for specification compliance"

## Check Modes

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 227)May include surrounding context.

md
[spark_sql_checker.py](scripts/spark_sql_checker.py)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 228)May include surrounding context.

md
[spark_sql_parser.py](scripts/spark_sql_parser.py)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 248)May include surrounding context.

md
| [Keywords](rules/keywords.py) | Spark SQL keyword definitions |

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 249)May include surrounding context.

md
| [Grammar Rules](rules/grammar_rules.py) | Statement type grammar definitions |

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · scripts/spark_sql_parser.py (reported line 504)May include surrounding context.

python
return node

    # ============================================================
    # DELETE Statement Parser
    # ============================================================

    def _parse_delete(self):

YARA rule 'exploit_framework': Exploit framework components and payloads [hacktools]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · scripts/spark_sql_parser.py (reported line 862)May include surrounding context.

python
DE
        if "CASCADE" in node.children.get("actions", "").upper():
            node.children["has_cascade"] = True

        node.tokens = [start_token]
        node.raw_text = self._get_raw_text(start_token)
        return node

    # ============================================================
    # DROP Parser
    # ============================================================

    def _parse_drop(self):
        node = ASTNode("DropStmt")
        start_token = self._current()

        self._advance()  # skip DROP

        # Object type
        obj_types = ["TABLE", "INDEX", "VIEW", "DATABASE", "SCHEMA", "FUNCTION",
                    "MATERIALIZED", "PROCEDURE", "PACKAGE"]
        obj_type = None
        for ot in obj_types:
            if self._current().is_keyword(ot):
                obj_type = ot
                self._advance()
                break
        node.children["object_type"] = obj_type or "UNKNOWN"

        # IF EXISTS
        if self._current().is_keyword("IF"):

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The MULTI_TOKEN_STATEMENTS table includes mappings for CREATE ROLE and DROP ROLE at L578 and L590, implying these statement types are recognized by the grammar layer. However, STATEMENT_RULES contains no definitions for either statement type anywhere in the file, so the documented detection table contradicts the actual rule set available for validation.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Lines L579-L580 map CREATE DATABASE and CREATE SCHEMA to a CREATE DATABASE statement type, which indicates the parser should recognize those statements. But STATEMENT_RULES and CLAUSE_ORDER_MAP define no CREATE DATABASE entry, so the code comments and structure indicating support conflict with the actual implemented validation rules.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This YAML rule file uses Chinese-only names, descriptions, categories, and suggestions throughout, with no indication that the skill is region-specific or that users can opt into this locale. That creates a natural-language policy concern because the skill effectively forces a specific language without documented user choice or justification.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This YAML rule file uses Chinese throughout its title, rule names, descriptions, and fix suggestions, which indicates the skill is constrained to a specific language/locale. The file does not offer a language choice or explain that the skill is intentionally region- or language-specific, so it appears to force a locale without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This code embeds Chinese natural-language strings throughout violations, report headers, CLI usage, and markdown output, which effectively forces a specific language for users. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly documented and justified.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

L220 states the report contains a 'Large SQL interception section' with INTERCEPT001-INTERCEPT007 violations, but the rest of the document only describes a tokenizer/parser/checker implementing syntax and specification checks, with no interception stage or interception rules defined. This is an active documentation contradiction about what the skill reports, not just an omission.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · references/acceptance-criteria.md (reported line 5)May include surrounding context.

md
## Functional Requirements

1. **Syntax Check**: The skill must correctly identify syntax errors in Spark SQL statements, including but not limited to:
   - Invalid keywords
   - Reserved keywords used as identifiers
   - Missing required clauses or keywords

Static analysis

No suspicious patterns detected.