Install
openclaw skills install @3mper0rr/dataset-creation-skillDataset Creation & Training
openclaw skills install @3mper0rr/dataset-creation-skillComplete guide for building high-quality datasets and training models. Quality data consistently outperforms model size and hyperparameter tuning — this skill covers the full lifecycle from raw collection to trained model.
1.1 Define the objective
Before collecting data, answer:
1.2 Identify data sources
| Source Type | Best For | Considerations |
|---|---|---|
| Existing public datasets | Baseline, benchmarking | Check license, verify quality |
| Web scraping | Domain-specific text | Legal review required, clean aggressively |
| Internal logs / databases | Production-aligned data | PII risk, access controls |
| Human annotation | High-precision labels | Expensive, needs guidelines |
| Synthetic generation | Augmentation, bootstrapping | Risk of model collapse, verify quality |
1.3 Collect raw data
import pandas as pd
# Example: Load from multiple sources
sources = [
pd.read_csv("source_a.csv"),
pd.read_json("source_b.jsonl", lines=True),
pd.read_parquet("source_c.parquet"),
]
raw = pd.concat(sources, ignore_index=True)
print(f"Collected {len(raw)} raw records")
print(f"Columns: {raw.columns.tolist()}")