T09 · Insecure Skill Coding Practices
- Location
scripts/write.py:9- Finding
Dataset Name Validation Bypass Enables Writes Outside the Intended Data Directory
- Content
View full analysis
Vulnerability Details
File Location:
scripts/write.py:9-67
Vulnerability Type: Improper path validation and path traversal
Risk Level: HighVulnerable Code
python def create_dataset(dataset_name, fields): try: df = pd.DataFrame({ "_id": [uuid.uuid4().__str__()], "dataset_name": [dataset_name], "fields": [json.dumps(fields)], "_created_at": [datetime.now()] }) metadata_path = os.path.join(get_data_path(), "metadata.lance") lance.write_dataset(df, metadata_path, mode="append") return create_response("create_dataset", "success", None, None) except Exception as e: return create_response("create_dataset", "error", None, str(e)) def check_dataset_exists(dataset_name): try: metadata_path = os.path.join(get_data_path(), "metadata.lance") metadata_ds = lance.dataset(metadata_path) metadata_df = metadata_ds.to_table().to_pandas() matching_rows = metadata_df[metadata_df['dataset_name'] == dataset_name] if matching_rows.empty: raise ValueError(f"Dataset {dataset_name} does not exist in metadata.") fields_data = matching_rows.iloc[0]['fields'] if isinstance(fields_data, str): return json.loads(fields_data) else: return fields_data except Exception as e: raise ValueError(f"Error checking dataset metadata: {str(e)}") def validate_field_count(fields, new_data): if len(fields) != len(new_data): raise ValueError(f"Number of fields mismatch: expected {len(fields)}, got {len(new_data)}") def append_to_dataset(new_data, dataset_name): try: fields = check_dataset_exists(dataset_name) validate_field_count(fields, new_data) # Create a dictionary with _id, _updated_at, and the field data data_dict = { ...[truncated 3689 chars]- Remediation
View remediation
Remediation Suggestions
-
Import and invoke the existing validator in every write operation:
python from manage import ( get_data_path, create_response, check_dataset_exists, validate_dataset_name, ) -
Remove the duplicate
check_dataset_exists()implementation fromwrite.py. Use the validated implementation frommanage.pyso all read, write, and management operations enforce the same policy. -
Validate names before storing them in metadata:
python def create_dataset(dataset_name, fields): try: validate_dataset_name(dataset_name) # Continue with metadata creation. -
Resolve and verify every final dataset path immediately before filesystem access. Prefer
os.path.commonpath()over string-prefix checks:python def safe_dataset_path(dataset_name): validate_dataset_name(dataset_name) root = os.path.realpath(get_data_path()) destination = os.path.realpath(os.path.join(root, dataset_name)) if os.path.commonpath([root, destination]) != root: raise ValueError("Dataset path escapes the data directory") return destination -
Use
safe_dataset_path()consistently in append, batch append, update, delete, backup, read, and drop operations. -
Reject duplicate dataset names and consider restricting names to a conservative allowlist such as letters, digits, underscores, and hyphens.
-
Add regression tests covering absolute paths,
../traversal, both path separator styles, symbolic-link path components, empty names, and valid names.
-
