Entropy
When to Use
Use this skill when working on entropy problems in information theory.
Decision Tree
Shannon Entropy
- H(X) = -sum p(x) log2 p(x)
- Maximum for uniform distribution: H_max = log2(n)
- Minimum = 0 for deterministic (one outcome certain)
scipy.stats.entropy(p, base=2) for discrete
Entropy Properties
- Non-negative: H(X) >= 0
- Concave in p
- Chain rule: H(X,Y) = H(X) + H(Y|X)
z3_solve.py prove "entropy_nonnegative"
Joint and Conditional Entropy
- H(X,Y) = -sum sum p(x,y) log2 p(x,y)
- H(Y|X) = H(X,Y) - H(X)
- H(Y|X) <= H(Y) with equality iff independent
Differential Entropy (Continuous)
- h(X) = -integral f(x) log f(x) dx
- Can be negative!
- Gaussian: h(X) = 0.5 * log2(2pie*sigma^2)
sympy_compute.py integrate "-f(x)*log(f(x))" --var x
Maximum Entropy Principle
- Given constraints, max entropy distribution is least biased
- Uniform for no constraints
- Exponential for E[X] = mu constraint
- Gaussian for E[X], Var[X] constraints
Tool Commands
Scipy_Entropy
uv run python -c "from scipy.stats import entropy; p = [0.25, 0.25, 0.25, 0.25]; H = entropy(p, base=2); print('Entropy:', H, 'bits')"
Scipy_Kl_Div
uv run python -c "from scipy.stats import entropy; p = [0.5, 0.5]; q = [0.9, 0.1]; kl = entropy(p, q); print('KL divergence:', kl)"
Sympy_Entropy
uv run python -m runtime.harness scripts/sympy_compute.py simplify "-p*log(p, 2) - (1-p)*log(1-p, 2)"
Key Techniques
From indexed textbooks:
- [Elements of Information Theory] Elements of Information Theory -- Thomas M_ Cover & Joy A_ Thomas -- 2_, Auflage, New York, NY, 2012 -- Wiley-Interscience -- 9780470303153 -- 2fcfe3e8a16b3aeefeaf9429fcf9a513 -- Anna’s Archive. What is the channel capacity of this channel? This is the multiple-access channel solved by Liao and Ahlswede.
Cognitive Tools Reference
See .claude/skills/math-mode/SKILL.md for full tool documentation.
1---2name: entropy3description: Problem-solving strategies for entropy in information theory4---5
6# Entropy
7
8## When to Use
9
10Use this skill when working on entropy problems in information theory.
11
12## Decision Tree
13
14
151. **Shannon Entropy**
16 - H(X) = -sum p(x) log2 p(x)
17 - Maximum for uniform distribution: H_max = log2(n)
18 - Minimum = 0 for deterministic (one outcome certain)
19 - `scipy.stats.entropy(p, base=2)` for discrete
20
212. **Entropy Properties**
22 - Non-negative: H(X) >= 0
23 - Concave in p
24 - Chain rule: H(X,Y) = H(X) + H(Y|X)
25 - `z3_solve.py prove "entropy_nonnegative"`
26
273. **Joint and Conditional Entropy**
28 - H(X,Y) = -sum sum p(x,y) log2 p(x,y)
29 - H(Y|X) = H(X,Y) - H(X)
30 - H(Y|X) <= H(Y) with equality iff independent
31
324. **Differential Entropy (Continuous)**
33 - h(X) = -integral f(x) log f(x) dx
34 - Can be negative!
35 - Gaussian: h(X) = 0.5 * log2(2*pi*e*sigma^2)
36 - `sympy_compute.py integrate "-f(x)*log(f(x))" --var x`
37
385. **Maximum Entropy Principle**
39 - Given constraints, max entropy distribution is least biased
40 - Uniform for no constraints
41 - Exponential for E[X] = mu constraint
42 - Gaussian for E[X], Var[X] constraints
43
44
45## Tool Commands
46
47### Scipy_Entropy
48```bash
49uv run python -c "from scipy.stats import entropy; p = [0.25, 0.25, 0.25, 0.25]; H = entropy(p, base=2); print('Entropy:', H, 'bits')"
50```
51
52### Scipy_Kl_Div
53```bash
54uv run python -c "from scipy.stats import entropy; p = [0.5, 0.5]; q = [0.9, 0.1]; kl = entropy(p, q); print('KL divergence:', kl)"
55```
56
57### Sympy_Entropy
58```bash
59uv run python -m runtime.harness scripts/sympy_compute.py simplify "-p*log(p, 2) - (1-p)*log(1-p, 2)"
60```
61
62## Key Techniques
63
64*From indexed textbooks:*
65
66- [Elements of Information Theory] Elements of Information Theory -- Thomas M_ Cover & Joy A_ Thomas -- 2_, Auflage, New York, NY, 2012 -- Wiley-Interscience -- 9780470303153 -- 2fcfe3e8a16b3aeefeaf9429fcf9a513 -- Anna’s Archive. What is the channel capacity of this channel? This is the multiple\-access channel solved by Liao and Ahlswede.
67
68## Cognitive Tools Reference
69
70See `.claude/skills/math-mode/SKILL.md` for full tool documentation.