Back to skill

Security audit

Rag Evaluator

Security checks for vulnerabilities and agentic risk

Overview

This skill is a local RAG logging/export CLI, but its metadata presents it as a different Python SDK product, and it persists potentially sensitive prompt/evaluation data in plaintext.

Review this as a local plaintext logging utility, not as a Python SDK. Do not log secrets, regulated data, proprietary prompts, retrieved documents, or credentials unless you are comfortable with them being stored under ~/.local/share/rag-evaluator and included in full-history exports. Treat CSV/JSON exports cautiously, especially before opening them in spreadsheet tools or sharing them.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/script.sh:64
Finding

Unescaped User-Controlled Data in JSON and CSV Exports

Content
View full analysis
> "$out" printf ' {"type":"%s","time":"%s","value":"%s"}' "$name" "$ts" "$val" >> "$out" done < "$f" echo "\n]" >> "$out" ;; csv) echo "type,time,value" > "$out" for f in "$DATA_DIR"/*.log; do [ -f "$f" ] || continue local name=$(basename "$f" .log) while IFS='|' read -r ts val; do echo "$name,$ts,$val" >> "$out"; done < "$f" done ``` A representative path through which user-controlled data reaches these export operations is: ```bash # scripts/script.sh:126-132 local input="$*" local ts=$(date '+%Y-%m-%d %H:%M') echo "$ts|$input" >> "$DATA_DIR/configure.log" local total=$(wc -l < "$DATA_DIR/configure.log") echo " [Rag Evaluator] configure: $input" echo " Saved. Total configure entries: $total" _log "configure" "$input" ``` Equivalent input-storage logic is repeated for the other domain commands. ### Technical Analysis Command arguments are stored in log files without validation and later inserted directly into JSON and CSV output. For JSON exports, quotation marks, backslashes, control characters, and embedded newlines are not JSON-escaped. An attacker can therefore terminate the intended string and inject additional JSON properties or objects, or simply produce invalid JSON. For CSV exports, fields are not enclosed and escaped according to CSV rules. Commas, quotation marks, and newlines can alter the exported row structure. More importantly, values beginning with spreadsheet formula characters such as `=`, `+`, `-`, or `@` remain active. When the CSV is opened in a spreadsheet application, the value may be evalua ...[truncated 2109 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (8)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill metadata claims a Python SDK for agent observability and evaluation, but the file describes a Bash-style local logging CLI that stores user inputs on disk. This mismatch can mislead users and downstream agents into invoking the skill under false assumptions, increasing the chance that sensitive prompts, evaluation notes, or operational data are written locally and exported without informed consent.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest advertises one class of functionality, while the documented behavior is a different tool that logs arbitrary inputs to local files and supports export/search. This is dangerous because users may provide confidential model prompts, benchmark data, or internal experiment details believing they are using an observability SDK rather than a persistent plaintext logging utility.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The file header names the skill 'Ragaai Catalyst' but the body documents a different tool, 'Rag Evaluator.' This identity mismatch weakens trust boundaries and can cause operators or automated systems to install or run a skill different from what they intended, especially in environments where skill names drive approval or allowlisting.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill encourages logging prompts, evaluations, costs, and other free-form inputs, but does not prominently warn that these entries are persistently stored and exportable. In the context of RAG and prompt evaluation, such inputs often contain proprietary prompts, retrieved content, API usage details, or sensitive test data, so silent persistence materially increases confidentiality risk.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file behavior materially differs from the stated package purpose: instead of implementing a Python observability/evaluation SDK, it provides a Bash CLI that records arbitrary user-supplied text into persistent local logs and supports later export/search. This mismatch is dangerous because users may trust the package metadata and unintentionally disclose prompts, datasets, secrets, or operational details to local storage through functionality they did not expect.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The script creates a persistent application data directory, writes command history and per-command logs, and can aggregate them into export files, none of which is clearly justified by the advertised SDK role. In a skill/package context, this increases the chance of silent collection and secondary exposure of sensitive evaluation data, prompts, or credentials pasted by users during normal operation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The export function traverses all log files and writes aggregated historical content into new files in JSON, CSV, or text formats, expanding the number of copies of potentially sensitive user data. Without clear disclosure, scoping, or filtering, this broad export behavior increases exposure risk by making retention more durable and easier to access, share, or back up unintentionally.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

At this code location, arbitrary user input is appended directly to a persistent log file and mirrored into a history log without any warning, consent flow, or redaction. This is dangerous because users may enter prompts, internal documents, tokens, or other sensitive content assuming a transient CLI interaction, but the script retains that data on disk for later discovery or exfiltration.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.