ML Visualizer
A data toolkit for ingesting, transforming, querying, and visualizing machine learning datasets. Manage your entire data pipeline — from raw ingestion through profiling and validation — all from the command line.
Commands
| Command |
Description |
ml-visualizer ingest <input> |
Ingest raw data or record a data source entry |
ml-visualizer transform <input> |
Log a data transformation step or operation |
ml-visualizer query <input> |
Record a query against your dataset |
ml-visualizer filter <input> |
Log a filter operation applied to data |
ml-visualizer aggregate <input> |
Record an aggregation or rollup operation |
ml-visualizer visualize <input> |
Log a visualization request or chart specification |
ml-visualizer export <input> |
Record an export operation or export all data |
ml-visualizer sample <input> |
Log a data sampling operation |
ml-visualizer schema <input> |
Record or describe a data schema |
ml-visualizer validate <input> |
Log a data validation check |
ml-visualizer pipeline <input> |
Record a full pipeline definition or step |
ml-visualizer profile <input> |
Log a data profiling run |
ml-visualizer stats |
Show summary statistics across all entry types |
ml-visualizer export <fmt> |
Export all data (formats: json, csv, txt) |
ml-visualizer search <term> |
Search across all entries by keyword |
ml-visualizer recent |
Show the 20 most recent activity log entries |
ml-visualizer status |
Health check — version, disk usage, last activity |
ml-visualizer help |
Show the built-in help message |
ml-visualizer version |
Print the current version (v2.0.0) |
Each data command (ingest, transform, query, etc.) works in two modes:
- Without arguments — displays the 20 most recent entries of that type
- With arguments — saves the input as a new timestamped entry
Data Storage
All data is stored as plain-text log files in ~/.local/share/ml-visualizer/:
- Each command type gets its own log file (e.g.,
ingest.log, transform.log, visualize.log)
- Entries are stored in
timestamp|value format for easy parsing
- A unified
history.log tracks all activity across command types
- Export to JSON, CSV, or TXT at any time with the
export command
Set the ML_VISUALIZER_DIR environment variable to override the default data directory.
Requirements
- Bash 4.0+ (uses
set -euo pipefail)
- Standard Unix utilities:
date, wc, du, tail, grep, sed, cat
- No external dependencies or API keys required
When to Use
- Building a data pipeline journal — use
ingest, transform, and pipeline to document each step of your ML data preparation workflow
- Tracking data quality — use
validate and profile to log validation checks and profiling runs, ensuring data integrity before model training
- Logging visualization requests — use
visualize to record what charts and plots you've generated for model diagnostics (confusion matrices, ROC curves, feature importance)
- Managing dataset schemas — use
schema to document the structure of your datasets, track schema changes over time, and share definitions with your team
- Auditing data operations — use
search, recent, and stats to review your complete data processing history and find specific operations
Examples
# Ingest a new data source
ml-visualizer ingest "Loaded training set from s3://ml-data/train.csv — 50,000 rows, 24 features"
# Record a transformation step
ml-visualizer transform "Applied StandardScaler to numeric columns, one-hot encoded categoricals"
# Log a visualization
ml-visualizer visualize "Generated confusion matrix for RandomForest classifier — 94% accuracy"
# Define a schema entry
ml-visualizer schema "users table: id(int), age(int), income(float), segment(str), churn(bool)"
# Search past operations
ml-visualizer search "StandardScaler"
Output
All commands print results to stdout. Redirect to a file if needed:
ml-visualizer stats > pipeline-report.txt
ml-visualizer export json
Powered by BytesAgain | bytesagain.com | hello@bytesagain.com
1---2name: yellowbrick3description: Visual analysis and diagnostic tools to help machine learning model selection. ml-visualizer, python, anaconda, estimator, machine-learning, matplotlib.4---5
6# ML Visualizer
7
8A data toolkit for ingesting, transforming, querying, and visualizing machine learning datasets. Manage your entire data pipeline — from raw ingestion through profiling and validation — all from the command line.
9
10## Commands
11
12| Command | Description |
13|---------|-------------|
14| `ml-visualizer ingest <input>` | Ingest raw data or record a data source entry |
15| `ml-visualizer transform <input>` | Log a data transformation step or operation |
16| `ml-visualizer query <input>` | Record a query against your dataset |
17| `ml-visualizer filter <input>` | Log a filter operation applied to data |
18| `ml-visualizer aggregate <input>` | Record an aggregation or rollup operation |
19| `ml-visualizer visualize <input>` | Log a visualization request or chart specification |
20| `ml-visualizer export <input>` | Record an export operation or export all data |
21| `ml-visualizer sample <input>` | Log a data sampling operation |
22| `ml-visualizer schema <input>` | Record or describe a data schema |
23| `ml-visualizer validate <input>` | Log a data validation check |
24| `ml-visualizer pipeline <input>` | Record a full pipeline definition or step |
25| `ml-visualizer profile <input>` | Log a data profiling run |
26| `ml-visualizer stats` | Show summary statistics across all entry types |
27| `ml-visualizer export <fmt>` | Export all data (formats: `json`, `csv`, `txt`) |
28| `ml-visualizer search <term>` | Search across all entries by keyword |
29| `ml-visualizer recent` | Show the 20 most recent activity log entries |
30| `ml-visualizer status` | Health check — version, disk usage, last activity |
31| `ml-visualizer help` | Show the built-in help message |
32| `ml-visualizer version` | Print the current version (v2.0.0) |
33
34Each data command (ingest, transform, query, etc.) works in two modes:
35- **Without arguments** — displays the 20 most recent entries of that type
36- **With arguments** — saves the input as a new timestamped entry
37
38## Data Storage
39
40All data is stored as plain-text log files in `~/.local/share/ml-visualizer/`:
41
42- Each command type gets its own log file (e.g., `ingest.log`, `transform.log`, `visualize.log`)
43- Entries are stored in `timestamp|value` format for easy parsing
44- A unified `history.log` tracks all activity across command types
45- Export to JSON, CSV, or TXT at any time with the `export` command
46
47Set the `ML_VISUALIZER_DIR` environment variable to override the default data directory.
48
49## Requirements
50
51- Bash 4.0+ (uses `set -euo pipefail`)
52- Standard Unix utilities: `date`, `wc`, `du`, `tail`, `grep`, `sed`, `cat`
53- No external dependencies or API keys required
54
55## When to Use
56
571. **Building a data pipeline journal** — use `ingest`, `transform`, and `pipeline` to document each step of your ML data preparation workflow
582. **Tracking data quality** — use `validate` and `profile` to log validation checks and profiling runs, ensuring data integrity before model training
593. **Logging visualization requests** — use `visualize` to record what charts and plots you've generated for model diagnostics (confusion matrices, ROC curves, feature importance)
604. **Managing dataset schemas** — use `schema` to document the structure of your datasets, track schema changes over time, and share definitions with your team
615. **Auditing data operations** — use `search`, `recent`, and `stats` to review your complete data processing history and find specific operations
62
63## Examples
64
65```bash
66# Ingest a new data source
67ml-visualizer ingest "Loaded training set from s3://ml-data/train.csv — 50,000 rows, 24 features"
68
69# Record a transformation step
70ml-visualizer transform "Applied StandardScaler to numeric columns, one-hot encoded categoricals"
71
72# Log a visualization
73ml-visualizer visualize "Generated confusion matrix for RandomForest classifier — 94% accuracy"
74
75# Define a schema entry
76ml-visualizer schema "users table: id(int), age(int), income(float), segment(str), churn(bool)"
77
78# Search past operations
79ml-visualizer search "StandardScaler"
80```
81
82## Output
83
84All commands print results to stdout. Redirect to a file if needed:
85
86```bash
87ml-visualizer stats > pipeline-report.txt
88ml-visualizer export json
89```
90
91---
92
93Powered by BytesAgain | bytesagain.com | hello@bytesagain.com