Simulacrum Data Annotation Workflow
Complete end-to-end workflow for time series dataset preparation and annotation on the Data Annotation platform (data.smlcrm.com).
What This Skill Does
This skill captures the precise workflow for processing time series datasets (Energy, Manufacturing, Climate) from discovery to CLEAN status:
- Find Dataset: Search Kaggle for Energy/Manufacturing/Climate time series data
- Download: Get CSV files via browser or Kaggle CLI
- Clean: Run Python/pandas script to handle missing values, duplicates, formatting
- Upload RAW: Upload original CSV with metadata (name, domain, source URL, description)
- Configure Headers: Set column types (Time, Target, Covariate, Group) and units
- Assign Groups: Select ALL variables (target + covariates), apply ALL group tags
- Upload Cleaned: Final upload → CLEAN status
Supported Domains
- Energy: Power consumption, utilities, renewable energy, grid data
- Manufacturing: Industrial processes, steel production, emissions, equipment data
- Climate: CO2 emissions, environmental monitoring, weather correlation data
Quick Start
For the full pipeline from Kaggle to annotated dataset:
1. Find dataset on Kaggle
2. Download (browser or kaggle CLI)
3. Clean with scripts/clean_dataset.py
4. Upload RAW dataset to data.smlcrm.com (with metadata)
5. Click "Clean" and upload cleaned file
6. Configure column metadata (types, units)
7. Assign groups to variables
8. Upload cleaned dataset → CLEAN status
Workflow Steps
Step 1: Find and Download Dataset
From Kaggle (Browser Method):
- Navigate to kaggle.com/datasets
- Search for relevant dataset (e.g., "steel industry energy consumption", "manufacturing emissions", "climate CO2")
- Review data description, file list, and preview
- Click "Download" button
- Extract CSV file from downloaded zip
Alternative: Kaggle CLI
# Install if needed: pip install kaggle
# Configure: kaggle competitions list
scripts/download_kaggle.sh <dataset-name> [output-dir]
# Example: scripts/download_kaggle.sh csafrit2/steel-industry-energy-consumption
Step 2: Clean the Dataset
Always run the cleaning script before upload:
python3 scripts/clean_dataset.py <input.csv> [-o <output.csv>]
What the script does:
- Strips whitespace from column names
- Removes duplicate rows
- Fills missing numeric values with median
- Fills missing categorical values with mode or 'Unknown'
- Converts timestamp columns to datetime format
- Outputs column summary for metadata configuration
Output:
- Cleaned CSV file ready for upload
- Column summary printed to console (save this for metadata config)
Step 3: Upload Raw Dataset to Platform
- Navigate to data.smlcrm.com/dashboard
- Click "Upload Dataset" button
- Fill in metadata for the RAW dataset:
- Name: Descriptive dataset name
- Domain: Category (Energy, Manufacturing, Climate, etc.)
- Source URL: Kaggle or original source URL
- Description: Brief summary of the dataset
- Upload the original/raw CSV file (not cleaned yet)
- Click Upload
Result: Dataset appears in list with RAW status
Step 4: Upload Cleaned File & Configure Metadata
- Find the RAW dataset in the list
- Click "Clean" button
- Upload the cleaned CSV file (from Step 2)
- Configure headers for each column:
| Setting |
Description |
| Name |
Column name (editable) |
| Units |
Measurement units (kWh, °C, %, ratio, tCO2, etc.) |
| Type |
Time / Target / Covariate / Group |
Column Type Guide:
- Time: Timestamp/datetime columns (usually required)
- Target: Variable to predict (at least one required)
- Covariate: Input features/independent variables
- Group: Categorical segment variables (WeekStatus, Day_of_week, Load_Type, etc.)
Bulk Configuration:
- Select multiple rows via checkboxes
- Use "Apply" dropdown to set type for selected columns
- Set units individually or in bulk
Common Unit Patterns:
- Energy: kWh, MWh, MW
- Power: kVarh, kW
- Emissions: tCO2, kgCO2
- Ratios: ratio, %
- Time: seconds, minutes, hours
Step 5: Assign Groups to Variables
Purpose: Group variables define how data is segmented for analysis.
Exact Workflow:
Select ALL variables by checking their checkboxes:
- Target variable(s)
- ALL covariate variables
Apply ALL group tags to selected variables:
- Click first group tag (e.g., WeekStatus) → all selected get this group
- Click second group tag (e.g., Day_of_week) → all selected get this group
- Click third group tag (e.g., Load_Type) → all selected get this group
- Continue for all available group tags
Result: All variables have all groups assigned (e.g., "WeekStatus × Day_of_week × Load_Type")
Important: Assign groups to BOTH target variables AND all covariates.
Step 6: Final Upload
- Click "Upload Cleaned Dataset" button
- Wait for processing
- Dataset status changes from RAW → CLEAN
- Verify data points count is correct
Example: Steel Industry Energy Dataset
Source: https://www.kaggle.com/datasets/csafrit2/steel-industry-energy-consumption
Metadata:
- Name: Steel Industry Energy Consumption (South Korea)
- Domain: Energy
- Data Points: 350,400
Column Configuration:
| Column |
Type |
Units |
| Timestamps |
Time |
- |
| Usage_kWh |
Target |
kWh |
| Lagging_Current_Reactive.Power_kVarh |
Covariate |
kVarh |
| Leading_Current_Reactive_Power_kVarh |
Covariate |
kVarh |
| CO2(tCO2) |
Covariate |
tCO2 |
| Lagging_Current_Power_Factor |
Covariate |
ratio |
| Leading_Current_Power_Factor |
Covariate |
ratio |
| NSM |
Covariate |
seconds |
| WeekStatus |
Group |
- |
| Day_of_week |
Group |
- |
| Load_Type |
Group |
- |
Group Assignment:
- Select: Usage_kWh, Lagging_Current_Reactive.Power_kVarh, Leading_Current_Reactive_Power_kVarh, CO2(tCO2), Lagging_Current_Power_Factor, Leading_Current_Power_Factor, NSM
- Click: WeekStatus → all selected get WeekStatus
- Click: Day_of_week → all selected get Day_of_week
- Click: Load_Type → all selected get Load_Type
- Final: All variables show "WeekStatus × Day_of_week × Load_Type"
Reference Materials
For detailed platform configuration guidance, see references/platform_guide.md.
Troubleshooting
"Next" button disabled:
- Check at least one Time column is set
- Check at least one Target column is set
- Verify all columns have types assigned
Groups not appearing:
- Columns must be marked as "Group" type first
- Proceed to next step after setting Group types
Upload fails:
- Re-run cleaning script
- Check CSV format (comma-delimited)
- Verify no empty column names
Scripts
| Script |
Purpose |
scripts/clean_dataset.py |
Clean and prepare CSV for upload |
scripts/download_kaggle.sh |
Download datasets via Kaggle CLI |
Platform URL
Data Annotation Platform: https://data.smlcrm.com
1---2name: data-cleaning-annotation-workflow3description: Complete workflow for time series datasets (Energy, Manufacturing, Climate) on Kaggle to Data Annotation platform (data.smlcrm.com). Includes downloading, cleaning with pandas, uploading RAW with metadata, configuring columns (Time/Target/Covariate/Group), setting units (kWh, kVarh, tCO2, ratio, seconds), and assigning groups by selecting all variables and applying all group tags. Use when finding Kaggle datasets, cleaning for ML, uploading with metadata, configuring types/units, assigning groups to all variables, or complete pipeline to CLEAN status.4---5
6# Simulacrum Data Annotation Workflow
7
8Complete end-to-end workflow for time series dataset preparation and annotation on the Data Annotation platform (data.smlcrm.com).
9
10## What This Skill Does
11
12This skill captures the precise workflow for processing time series datasets (Energy, Manufacturing, Climate) from discovery to CLEAN status:
13
141. **Find Dataset**: Search Kaggle for Energy/Manufacturing/Climate time series data
152. **Download**: Get CSV files via browser or Kaggle CLI
163. **Clean**: Run Python/pandas script to handle missing values, duplicates, formatting
174. **Upload RAW**: Upload original CSV with metadata (name, domain, source URL, description)
185. **Configure Headers**: Set column types (Time, Target, Covariate, Group) and units
196. **Assign Groups**: Select ALL variables (target + covariates), apply ALL group tags
207. **Upload Cleaned**: Final upload → **CLEAN** status
21
22## Supported Domains
23
24- **Energy**: Power consumption, utilities, renewable energy, grid data
25- **Manufacturing**: Industrial processes, steel production, emissions, equipment data
26- **Climate**: CO2 emissions, environmental monitoring, weather correlation data
27
28## Quick Start
29
30For the full pipeline from Kaggle to annotated dataset:
31
32```
331. Find dataset on Kaggle
342. Download (browser or kaggle CLI)
353. Clean with scripts/clean_dataset.py
364. Upload RAW dataset to data.smlcrm.com (with metadata)
375. Click "Clean" and upload cleaned file
386. Configure column metadata (types, units)
397. Assign groups to variables
408. Upload cleaned dataset → CLEAN status
41```
42
43## Workflow Steps
44
45### Step 1: Find and Download Dataset
46
47**From Kaggle (Browser Method):**
481. Navigate to kaggle.com/datasets
492. Search for relevant dataset (e.g., "steel industry energy consumption", "manufacturing emissions", "climate CO2")
503. Review data description, file list, and preview
514. Click "Download" button
525. Extract CSV file from downloaded zip
53
54**Alternative: Kaggle CLI**
55```bash
56# Install if needed: pip install kaggle
57# Configure: kaggle competitions list
58
59scripts/download_kaggle.sh <dataset-name> [output-dir]
60# Example: scripts/download_kaggle.sh csafrit2/steel-industry-energy-consumption
61```
62
63### Step 2: Clean the Dataset
64
65**Always run the cleaning script before upload:**
66
67```bash
68python3 scripts/clean_dataset.py <input.csv> [-o <output.csv>]
69```
70
71**What the script does:**
72- Strips whitespace from column names
73- Removes duplicate rows
74- Fills missing numeric values with median
75- Fills missing categorical values with mode or 'Unknown'
76- Converts timestamp columns to datetime format
77- Outputs column summary for metadata configuration
78
79**Output:**
80- Cleaned CSV file ready for upload
81- Column summary printed to console (save this for metadata config)
82
83### Step 3: Upload Raw Dataset to Platform
84
851. Navigate to data.smlcrm.com/dashboard
862. Click **"Upload Dataset"** button
873. Fill in metadata for the RAW dataset:
88 - **Name**: Descriptive dataset name
89 - **Domain**: Category (Energy, Manufacturing, Climate, etc.)
90 - **Source URL**: Kaggle or original source URL
91 - **Description**: Brief summary of the dataset
924. Upload the **original/raw** CSV file (not cleaned yet)
935. Click **Upload**
94
95**Result:** Dataset appears in list with **RAW** status
96
97### Step 4: Upload Cleaned File & Configure Metadata
98
991. Find the RAW dataset in the list
1002. Click **"Clean"** button
1013. Upload the **cleaned** CSV file (from Step 2)
1024. Configure headers for each column:
103
104| Setting | Description |
105|---------|-------------|
106| **Name** | Column name (editable) |
107| **Units** | Measurement units (kWh, °C, %, ratio, tCO2, etc.) |
108| **Type** | Time / Target / Covariate / Group |
109
110**Column Type Guide:**
111- **Time**: Timestamp/datetime columns (usually required)
112- **Target**: Variable to predict (at least one required)
113- **Covariate**: Input features/independent variables
114- **Group**: Categorical segment variables (WeekStatus, Day_of_week, Load_Type, etc.)
115
116**Bulk Configuration:**
117- Select multiple rows via checkboxes
118- Use "Apply" dropdown to set type for selected columns
119- Set units individually or in bulk
120
121**Common Unit Patterns:**
122- Energy: kWh, MWh, MW
123- Power: kVarh, kW
124- Emissions: tCO2, kgCO2
125- Ratios: ratio, %
126- Time: seconds, minutes, hours
127
128### Step 5: Assign Groups to Variables
129
130**Purpose:** Group variables define how data is segmented for analysis.
131
132**Exact Workflow:**
1331. **Select ALL variables** by checking their checkboxes:
134 - Target variable(s)
135 - ALL covariate variables
136
1372. **Apply ALL group tags** to selected variables:
138 - Click first group tag (e.g., WeekStatus) → all selected get this group
139 - Click second group tag (e.g., Day_of_week) → all selected get this group
140 - Click third group tag (e.g., Load_Type) → all selected get this group
141 - Continue for all available group tags
142
1433. **Result:** All variables have all groups assigned (e.g., "WeekStatus × Day_of_week × Load_Type")
144
145**Important:** Assign groups to BOTH target variables AND all covariates.
146
147### Step 6: Final Upload
148
1491. Click **"Upload Cleaned Dataset"** button
1502. Wait for processing
1513. Dataset status changes from **RAW** → **CLEAN**
1524. Verify data points count is correct
153
154## Example: Steel Industry Energy Dataset
155
156**Source:** https://www.kaggle.com/datasets/csafrit2/steel-industry-energy-consumption
157
158**Metadata:**
159- **Name:** Steel Industry Energy Consumption (South Korea)
160- **Domain:** Energy
161- **Data Points:** 350,400
162
163**Column Configuration:**
164| Column | Type | Units |
165|--------|------|-------|
166| Timestamps | Time | - |
167| Usage_kWh | Target | kWh |
168| Lagging_Current_Reactive.Power_kVarh | Covariate | kVarh |
169| Leading_Current_Reactive_Power_kVarh | Covariate | kVarh |
170| CO2(tCO2) | Covariate | tCO2 |
171| Lagging_Current_Power_Factor | Covariate | ratio |
172| Leading_Current_Power_Factor | Covariate | ratio |
173| NSM | Covariate | seconds |
174| WeekStatus | Group | - |
175| Day_of_week | Group | - |
176| Load_Type | Group | - |
177
178**Group Assignment:**
1791. Select: Usage_kWh, Lagging_Current_Reactive.Power_kVarh, Leading_Current_Reactive_Power_kVarh, CO2(tCO2), Lagging_Current_Power_Factor, Leading_Current_Power_Factor, NSM
1802. Click: WeekStatus → all selected get WeekStatus
1813. Click: Day_of_week → all selected get Day_of_week
1824. Click: Load_Type → all selected get Load_Type
1835. Final: All variables show "WeekStatus × Day_of_week × Load_Type"
184
185## Reference Materials
186
187For detailed platform configuration guidance, see [references/platform_guide.md](references/platform_guide.md).
188
189## Troubleshooting
190
191**"Next" button disabled:**
192- Check at least one Time column is set
193- Check at least one Target column is set
194- Verify all columns have types assigned
195
196**Groups not appearing:**
197- Columns must be marked as "Group" type first
198- Proceed to next step after setting Group types
199
200**Upload fails:**
201- Re-run cleaning script
202- Check CSV format (comma-delimited)
203- Verify no empty column names
204
205## Scripts
206
207| Script | Purpose |
208|--------|---------|
209| `scripts/clean_dataset.py` | Clean and prepare CSV for upload |
210| `scripts/download_kaggle.sh` | Download datasets via Kaggle CLI |
211
212## Platform URL
213
214Data Annotation Platform: https://data.smlcrm.com