1---2name: influxdb3description: Store and query time-series data with proper schema design and retention.4---56## Version Differences78- InfluxDB 2.x uses Flux query language, 1.x uses InfluxQL—syntax completely different9- 2.x: buckets, organizations, tokens; 1.x: databases, retention policies, users10- Don't mix documentation—check version before copying queries1112## Tags vs Fields (Critical)1314- Tags are indexed, fields are not—filter on tags, aggregate on fields15- Tag values must be strings—numbers as tags work but waste index space16- Fields support numbers, strings, booleans—store metrics as fields17- Wrong choice kills query performance—can't change after data written1819## Cardinality Trap2021- High-cardinality tags destroy performance—unique user IDs as tags = disaster22- Cardinality = unique combinations of tag values—grows multiplicatively23- Check with `SHOW CARDINALITY` (1.x) or `influx bucket inspect` (2.x)24- Rule of thumb: <100K series per measurement; millions = problems2526## Line Protocol2728- Format: `measurement,tag1=v1,tag2=v2 field1=1,field2="str" timestamp`29- No spaces around `=` in tags—space separates tags from fields30- String fields need quotes, tag values don't—`field="text"` vs `tag=text`31- Timestamps in nanoseconds by default—specify precision to avoid mistakes3233## Timestamps3435- Default precision is nanoseconds—sending seconds without precision flag = year 2000 data36- Specify on write: `precision=s` for seconds, `precision=ms` for milliseconds37- Missing timestamp uses server time—usually fine for real-time ingestion38- Timestamps are UTC—client timezone doesn't matter3940## Retention and Downsampling4142- Set retention policy/bucket duration—data older than retention auto-deleted43- Raw data at 10s intervals for 7 days, downsample to 1min for 30 days, 1h for 1 year44- 2.x: Tasks for downsampling; 1.x: Continuous Queries45- Without downsampling, storage grows forever and queries slow down4647## Flux Query Patterns (2.x)4849- Always start with `from(bucket:)` then `|> range(start:)`—range is required50- `|> filter(fn: (r) => r._measurement == "cpu")` for filtering51- `|> aggregateWindow(every: 1h, fn: mean)` for time-based aggregation52- Chain transforms with `|>` pipe operator—order matters for performance5354## InfluxQL Patterns (1.x)5556- `SELECT mean("value") FROM "measurement" WHERE time > now() - 1h GROUP BY time(5m)`57- Double quotes for identifiers, single quotes for string literals58- `GROUP BY time()` for time-based aggregation—required for most dashboards59- `FILL(none)` to skip empty intervals, `FILL(previous)` to carry forward6061## Schema Design6263- Measurement name = table name—one per metric type (cpu, memory, requests)64- Tag for dimensions you filter/group by—host, region, service65- Field for values you aggregate—usage_percent, count, latency_ms66- Avoid encoding data in measurement names—`cpu.host1` wrong, `cpu` + `host=host1` right6768## Write Performance6970- Batch writes—individual points have HTTP overhead71- Telegraf for production ingestion—handles batching, buffering, retry72- Write to localhost if possible—network latency adds up at high throughput73- `async` writes in client libraries—don't block on each write7475## Query Performance7677- Always include time range—unbounded queries scan everything78- Filter on tags before fields—tags use index, fields scan data79- Limit results with `LIMIT` or `|> limit()`—dashboard doesn't need 1M points80- Use `GROUP BY` / `aggregateWindow` to reduce data before returning8182## Common Errors8384- "partial write: field type conflict"—same field with different types; fix at source85- "max-values-per-tag limit exceeded"—cardinality too high; redesign schema86- "database not found"—2.x uses buckets, not databases; check API version87- Query timeout—add narrower time range or aggregate more aggressively