1---2name: influxdb3description: Store and query time-series data with proper schema design and retention.4---5
6## Version Differences
7
8- InfluxDB 2.x uses Flux query language, 1.x uses InfluxQL—syntax completely different
9- 2.x: buckets, organizations, tokens; 1.x: databases, retention policies, users
10- Don't mix documentation—check version before copying queries
11
12## Tags vs Fields (Critical)
13
14- Tags are indexed, fields are not—filter on tags, aggregate on fields
15- Tag values must be strings—numbers as tags work but waste index space
16- Fields support numbers, strings, booleans—store metrics as fields
17- Wrong choice kills query performance—can't change after data written
18
19## Cardinality Trap
20
21- High-cardinality tags destroy performance—unique user IDs as tags = disaster
22- Cardinality = unique combinations of tag values—grows multiplicatively
23- Check with `SHOW CARDINALITY` (1.x) or `influx bucket inspect` (2.x)
24- Rule of thumb: <100K series per measurement; millions = problems
25
26## Line Protocol
27
28- Format: `measurement,tag1=v1,tag2=v2 field1=1,field2="str" timestamp`
29- No spaces around `=` in tags—space separates tags from fields
30- String fields need quotes, tag values don't—`field="text"` vs `tag=text`
31- Timestamps in nanoseconds by default—specify precision to avoid mistakes
32
33## Timestamps
34
35- Default precision is nanoseconds—sending seconds without precision flag = year 2000 data
36- Specify on write: `precision=s` for seconds, `precision=ms` for milliseconds
37- Missing timestamp uses server time—usually fine for real-time ingestion
38- Timestamps are UTC—client timezone doesn't matter
39
40## Retention and Downsampling
41
42- Set retention policy/bucket duration—data older than retention auto-deleted
43- Raw data at 10s intervals for 7 days, downsample to 1min for 30 days, 1h for 1 year
44- 2.x: Tasks for downsampling; 1.x: Continuous Queries
45- Without downsampling, storage grows forever and queries slow down
46
47## Flux Query Patterns (2.x)
48
49- Always start with `from(bucket:)` then `|> range(start:)`—range is required
50- `|> filter(fn: (r) => r._measurement == "cpu")` for filtering
51- `|> aggregateWindow(every: 1h, fn: mean)` for time-based aggregation
52- Chain transforms with `|>` pipe operator—order matters for performance
53
54## InfluxQL Patterns (1.x)
55
56- `SELECT mean("value") FROM "measurement" WHERE time > now() - 1h GROUP BY time(5m)`
57- Double quotes for identifiers, single quotes for string literals
58- `GROUP BY time()` for time-based aggregation—required for most dashboards
59- `FILL(none)` to skip empty intervals, `FILL(previous)` to carry forward
60
61## Schema Design
62
63- Measurement name = table name—one per metric type (cpu, memory, requests)
64- Tag for dimensions you filter/group by—host, region, service
65- Field for values you aggregate—usage_percent, count, latency_ms
66- Avoid encoding data in measurement names—`cpu.host1` wrong, `cpu` + `host=host1` right
67
68## Write Performance
69
70- Batch writes—individual points have HTTP overhead
71- Telegraf for production ingestion—handles batching, buffering, retry
72- Write to localhost if possible—network latency adds up at high throughput
73- `async` writes in client libraries—don't block on each write
74
75## Query Performance
76
77- Always include time range—unbounded queries scan everything
78- Filter on tags before fields—tags use index, fields scan data
79- Limit results with `LIMIT` or `|> limit()`—dashboard doesn't need 1M points
80- Use `GROUP BY` / `aggregateWindow` to reduce data before returning
81
82## Common Errors
83
84- "partial write: field type conflict"—same field with different types; fix at source
85- "max-values-per-tag limit exceeded"—cardinality too high; redesign schema
86- "database not found"—2.x uses buckets, not databases; check API version
87- Query timeout—add narrower time range or aggregate more aggressively