Back to skill

Security audit

Data Validator Pro

Security checks for vulnerabilities and agentic risk

Overview

This is a straightforward local data-quality toolkit with no hidden access, persistence, network behavior, or credential handling found.

Reasonable to install for local data-quality work. Pin pandas and numpy if you need reproducible or audited deployments, and avoid accepting arbitrary regex validation schemas from untrusted users without timeouts or pattern limits.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/schema_validator.py:35
Finding

Unbounded Regular Expression Evaluation Enables Denial of Service

Content
View full analysis

Vulnerability Details

File Location: scripts/schema_validator.py:35-40
Vulnerability Type: Regular Expression Denial of Service (ReDoS)
Risk Level: Medium

python
# Regex
if "regex" in rules and (series.dtype == object or pd.api.types.is_string_dtype(series)):
    pattern = re.compile(rules["regex"])
    invalid = series[~series.astype(str).apply(lambda x: bool(pattern.match(x)) if pd.notna(x) else True)]
    if not invalid.empty:
        errors.append({"column": col, "error": "regex_mismatch", "count": len(invalid), "regex": rules["regex"]})

Technical Analysis

SchemaValidator compiles a regular expression taken directly from the supplied schema and applies it to every value in the selected column. Python's standard re engine can exhibit catastrophic backtracking for ambiguous nested quantifiers, and this implementation imposes no pattern-complexity restriction, input-length limit, row limit, or matching timeout.

An attacker who can control schema rules can provide a pathological expression such as ^(a+)+$. A long input consisting of repeated a characters followed by a nonmatching character can then require exponentially increasing CPU time. Applying the expression across many rows amplifies the resource consumption.

Attack Path

  1. The attacker obtains the ability to submit or influence a validation schema.
  2. The attacker defines a pathological regular expression, such as ^(a+)+$, for a string column.
  3. The attacker supplies one or more long values such as "aaaaaaaa...!".
  4. The application invokes SchemaValidator.validate().
  5. Lines 37-38 compile and repeatedly evaluate the expression without a timeout.
  6. Catastrophic backtracking consumes excessive CPU and delays or stalls the worker process.

Impact Assessment

Successful exploitation does not grant additional system privileges, data access, or code execution. Its impact is limited to availability: an affe ...[truncated 289 chars]

Remediation
View remediation

Remediation Suggestions

  • Treat validation schemas, especially regular expressions, as trusted configuration and prevent untrusted users from supplying arbitrary patterns.
  • Prefer a linear-time regular-expression engine when expressions must be user-controlled.
  • If the selected engine supports deadlines, enforce a strict per-match timeout and handle timeout failures as validation errors.
  • Reject oversized patterns and values before evaluation, and impose limits on the number of rows processed per request.
  • Validate patterns against a restrictive allowlist and reject dangerous constructs such as ambiguous nested quantifiers.
  • Consider replacing arbitrary regular expressions with predefined, reviewed validation rules for common formats.
  • Add regression tests using pathological patterns and long nonmatching strings to verify that validation remains within defined CPU and latency limits.
Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (5)

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The documentation headings and content are presented in Chinese and English together, but there is no indication that language choice is optional or user-selectable. Under the stated policy, forcing a specific language or locale without opt-in can be a natural-language policy issue.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

The dependency is only lower-bounded (pandas>=2.0.0), so builds are not reproducible and may pull different versions over time, including vulnerable or breaking releases. In a data-processing skill, pandas is a core library and broad version drift can affect both security posture and reliability, though this file alone does not prove exploitation.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
pandas>=2.0.0
numpy>=1.24.0

Unverifiable Dependency: pandas has 1 known advisory(ies) (CVE-2020-13091 (** DISPUTED ** pandas through 1.0.3 can unserialize and execute commands from an)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
78% confidence
Finding

The manifest references pandas without pinning to a specific release, so it is impossible to verify from this file whether deployment will resolve to a version affected by known advisories. While the cited advisory is disputed and may not apply to modern releases, the inability to prove a safe version is still a supply-chain hygiene weakness.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
93% confidence
Finding

The dependency is unpinned (numpy>=1.24.0), which allows non-deterministic installs and can introduce vulnerable or incompatible releases without code changes. Because numpy is foundational in data tooling, unexpected upgrades can propagate broadly through the runtime environment.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
pandas>=2.0.0
numpy>=1.24.0

Unverifiable Dependency: numpy has 16 known advisory(ies) (CVE-2014-1859 (Numpy arbitrary file write via symlink attack); CVE-2021-41495 (NumPy NULL Pointer Dereference); CVE-2021-33430 (NumPy Buffer Overflow (Disputed)) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
84% confidence
Finding

The manifest does not pin numpy, so there is no reliable way to determine whether the installed version is exposed to any of the package's historical advisories. Given numpy's broad attack surface and numerous downstream uses, unverifiable resolution increases supply-chain uncertainty even if no specific vulnerable version is shown here.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.