Estimation Patterns (Meta-Skill)
Systematic approaches for producing accurate, defensible software estimates.
Installation
OpenClaw / Moltbot / Clawbot
npx clawhub@latest install estimation-patterns
When to Use
- Estimating a feature, bug fix, or project timeline
- Breaking down work for sprint planning or roadmap forecasting
- Presenting estimates to stakeholders or product managers
- Reviewing historical accuracy to calibrate future estimates
- Noticing a pattern of missed deadlines or blown budgets
Estimation Methods
Choose the method that matches your context and audience.
| Method |
Best For |
Granularity |
Pros |
Cons |
| T-Shirt Sizing |
Roadmap planning, backlog grooming |
XS, S, M, L, XL |
Fast, low-friction, good for relative ranking |
Not actionable for scheduling |
| Story Points |
Sprint planning, team velocity |
Fibonacci (1-21) |
Abstracts away individual speed, tracks velocity |
Meaningless outside the team, gaming risk |
| Time-Based |
Client quotes, contractor work |
Hours / days |
Universally understood, maps to budgets |
Anchoring bias, implies false precision |
| Three-Point |
High-uncertainty tasks |
Min / likely / max |
Captures uncertainty range, enables PERT |
Requires discipline to set honest bounds |
| Reference Comparison |
Recurring task types |
Relative to past |
Grounded in real data, hard to argue with |
Requires historical records, breaks on novelty |
Three-point formula (PERT):
Expected = (Optimistic + 4 x Likely + Pessimistic) / 6
Standard Deviation = (Pessimistic - Optimistic) / 6
Use the standard deviation to express confidence ranges (e.g., "3-5 days at 68% confidence, 2-6 days at 95%").
Task Decomposition
Break work down until every sub-task is < 4 hours of effort. Anything larger hides unknowns.
| Level |
Example |
Target Size |
| Epic |
User authentication system |
2-6 weeks |
| Feature |
OAuth2 login with Google |
3-10 days |
| Task |
Implement callback handler |
1-3 days |
| Sub-task |
Parse and validate OAuth token |
1-4 hours |
| Atomic step |
Write token expiry check function |
30-90 minutes |
Decomposition checklist:
- Can I describe what "done" looks like in one sentence?
- Is there exactly one unknown, or zero?
- Could a teammate pick this up without a walkthrough?
- Is it under 4 hours? If no — split again.
If you cannot decompose a task, it signals a spike is needed. Timebox the spike (2-4 hours), then re-estimate.
Complexity Multipliers
Apply these multipliers to your base estimate when complexity factors are present. Multipliers stack multiplicatively.
| Factor |
Multiplier |
Rationale |
| New technology / stack |
1.5x |
Learning curve, unexpected gotchas, doc-hunting |
| Unclear requirements |
2.0x |
Discovery work, rework cycles, stakeholder alignment |
| Legacy code |
1.5x |
Undocumented behavior, fragile tests, hidden coupling |
| Cross-team dependency |
1.5x |
Coordination overhead, blocking, API negotiation |
| First-time task |
2.0x |
No reference point, unknown unknowns dominate |
| Regulatory / compliance |
1.5x |
Audit trails, review gates, documentation overhead |
Example: A 2-day base estimate on legacy code (1.5x) with unclear requirements (2.0x) becomes 2 x 1.5 x 2.0 = 6 days.
Rule: Never apply more than 3 multipliers — if that many factors converge, the task needs a spike or a scope reduction, not a bigger number.
Buffer Calculation
Raw estimates are point predictions. Reality is a distribution.
| Buffer Type |
Rule of Thumb |
When to Apply |
| Known unknowns |
+20% of total estimate |
Integration points, third-party APIs, minor gaps |
| Unknown unknowns |
+50% of total estimate |
New domain, first release, greenfield system |
| Team velocity factor |
/ focus ratio (e.g., 0.7) |
Account for meetings, reviews, context switching |
| Sequential dependency |
+10% per handoff |
Each team/person boundary adds coordination drag |
Effective estimate formula:
Effective = (Base Estimate x Multipliers) / Focus Ratio + Buffer
Focus ratio guidelines:
| Scenario |
Typical Focus Ratio |
| Dedicated to one project |
0.75-0.85 |
| Split across 2 projects |
0.50-0.60 |
| On-call rotation active |
0.60-0.70 |
| Heavy meeting load (> 3h/day) |
0.45-0.55 |
Historical Calibration
Track actual vs estimated to improve over time. This is the single most effective way to get better at estimation.
Tracking table:
| Task |
Estimated |
Actual |
Ratio (A/E) |
Notes |
| Auth flow |
3 days |
5 days |
1.67 |
OAuth docs were outdated |
| Dashboard charts |
5 days |
4 days |
0.80 |
Reused existing component |
| DB migration |
2 days |
6 days |
3.00 |
Discovered data quality issues |
Accuracy ratio: Calculate your rolling average of Actual / Estimated over the last 10-20 tasks.
- Ratio < 0.8 — you're overestimating (sandbagging or excessive buffers)
- Ratio 0.8-1.2 — well calibrated
- Ratio > 1.2 — you're underestimating (apply the ratio as a correction factor)
Calibration action: Multiply future estimates by your rolling accuracy ratio until it converges toward 1.0.
Common Estimation Biases
Recognize these cognitive traps — awareness alone reduces their effect.
| Bias |
Description |
Mitigation |
| Planning Fallacy |
Assuming best-case scenario despite past evidence |
Use historical data, not intuition |
| Anchoring |
First number heard dominates all subsequent estimates |
Estimate independently before discussing |
| Optimism Bias |
"It'll be simpler than last time" |
Apply the three-point method, honor the pessimistic |
| Scope Creep |
Estimate stays fixed while scope grows |
Re-estimate when scope changes, always |
| Hofstadter's Law |
"It always takes longer, even when you account for it" |
Add buffer, then add more buffer for novel work |
| Dunning-Kruger |
Novices underestimate; experts sometimes overestimate |
Cross-check with a second estimator |
| Sunk Cost Pressure |
Refusing to re-estimate because the original was "approved" |
Treat estimates as living artifacts, update often |
Estimation by Task Type
Use these ranges as starting heuristics, then adjust with multipliers and historical data.
| Task Type |
Typical Range |
Key Variables |
| Bug fix (isolated) |
2-8 hours |
Reproducibility, code familiarity, test coverage |
| Bug fix (systemic) |
1-3 days |
Root cause depth, blast radius, regression risk |
| Small feature |
1-3 days |
Spec clarity, UI complexity, number of endpoints |
| Medium feature |
3-10 days |
Cross-cutting concerns, data model changes |
| Large feature |
2-4 weeks |
Architecture decisions, team coordination |
| Refactor (local) |
1-3 days |
Test coverage, coupling, blast radius |
| Refactor (systemic) |
1-4 weeks |
Number of callers, migration strategy needed |
| Spike / research |
2-8 hours (timeboxed) |
Always timebox — output is knowledge, not code |
| DevOps / infra |
1-5 days |
Provider docs quality, IAM complexity, testing |
Communication
How you present an estimate matters as much as the number itself.
Always present as a range, never a single number:
- Bad: "It'll take 5 days."
- Good: "3-7 days, most likely 5. The range depends on the payment API response format — I'll know more after the spike."
Confidence levels:
| Confidence |
What It Means |
When to Use |
| High (+-15%) |
Well-understood scope, done similar before |
Familiar task, clear spec |
| Medium (+-30%) |
Some unknowns, reasonable decomposition |
Most sprint-level estimates |
| Low (+-50%+) |
Significant unknowns, rough order of magnitude |
Roadmap forecasts, presale quotes |
Stakeholder communication rules:
- State the range and the confidence level together
- Name the top 1-3 risks that could push toward the upper bound
- Offer to de-risk with a timeboxed spike before committing
- Explicitly state what is not included (e.g., "does not include QA, deployment, or docs")
- Update estimates proactively when new information surfaces — don't wait until the deadline
Anti-Patterns
| Anti-Pattern |
Why It's Harmful |
Better Approach |
| Padding silently |
Erodes trust when discovered; hides real uncertainty |
Use explicit buffers with stated rationale |
| Sandbagging |
Destroys velocity data; breeds complacency |
Track accuracy ratio, aim for calibration |
| Not decomposing |
Large estimates hide unknowns and compound errors |
Break to < 4-hour sub-tasks, estimate bottom-up |
| Single-point estimates |
Implies false certainty, no room for variance |
Always give a range with confidence level |
| Estimating under pressure |
Anchoring to what the stakeholder wants to hear |
Ask for time to decompose; never estimate on the spot |
| Copy-paste estimates |
Every task has different context and risk profile |
Estimate fresh, use references as starting points only |
| Ignoring rework cycles |
First pass is rarely final — reviews, feedback, QA |
Factor in at least one review-and-revise loop |
NEVER Do
- NEVER give a single-number estimate without a range — it communicates false precision and sets you up for failure
- NEVER estimate a task you haven't decomposed — large estimates are guesses wearing a suit
- NEVER let an old estimate stand after scope changes — estimates are invalidated the moment requirements shift
- NEVER estimate in someone else's units — your days are not their days; clarify assumptions about focus time and interrupts
- NEVER skip recording actuals — estimation without feedback is astrology, not engineering
- NEVER commit to an estimate made under pressure — say "let me break this down and get back to you in an hour"
- NEVER treat an estimate as a promise or a deadline — estimates are probabilistic forecasts, not contracts
1---2name: estimation-patterns3description: Practical estimation techniques for software tasks — methods comparison, decomposition, complexity multipliers, buffer calculation, bias awareness, and communication strategies. Use when estimating features, sprint planning, or presenting timelines to stakeholders.4---5
6# Estimation Patterns (Meta-Skill)
7
8Systematic approaches for producing accurate, defensible software estimates.
9
10
11## Installation
12
13### OpenClaw / Moltbot / Clawbot
14
15```bash
16npx clawhub@latest install estimation-patterns
17```
18
19
20---
21
22## When to Use
23
24- Estimating a feature, bug fix, or project timeline
25- Breaking down work for sprint planning or roadmap forecasting
26- Presenting estimates to stakeholders or product managers
27- Reviewing historical accuracy to calibrate future estimates
28- Noticing a pattern of missed deadlines or blown budgets
29
30---
31
32## Estimation Methods
33
34Choose the method that matches your context and audience.
35
36| Method | Best For | Granularity | Pros | Cons |
37|----------------------|-------------------------------|-----------------|-----------------------------------------------|-----------------------------------------------|
38| T-Shirt Sizing | Roadmap planning, backlog grooming | XS, S, M, L, XL | Fast, low-friction, good for relative ranking | Not actionable for scheduling |
39| Story Points | Sprint planning, team velocity | Fibonacci (1-21) | Abstracts away individual speed, tracks velocity | Meaningless outside the team, gaming risk |
40| Time-Based | Client quotes, contractor work | Hours / days | Universally understood, maps to budgets | Anchoring bias, implies false precision |
41| Three-Point | High-uncertainty tasks | Min / likely / max | Captures uncertainty range, enables PERT | Requires discipline to set honest bounds |
42| Reference Comparison | Recurring task types | Relative to past | Grounded in real data, hard to argue with | Requires historical records, breaks on novelty |
43
44**Three-point formula (PERT):**
45
46```
47Expected = (Optimistic + 4 x Likely + Pessimistic) / 6
48Standard Deviation = (Pessimistic - Optimistic) / 6
49```
50
51Use the standard deviation to express confidence ranges (e.g., "3-5 days at 68% confidence, 2-6 days at 95%").
52
53---
54
55## Task Decomposition
56
57Break work down until every sub-task is **< 4 hours** of effort. Anything larger hides unknowns.
58
59| Level | Example | Target Size |
60|----------------|-------------------------------------------|---------------|
61| Epic | User authentication system | 2-6 weeks |
62| Feature | OAuth2 login with Google | 3-10 days |
63| Task | Implement callback handler | 1-3 days |
64| Sub-task | Parse and validate OAuth token | 1-4 hours |
65| Atomic step | Write token expiry check function | 30-90 minutes |
66
67**Decomposition checklist:**
68
691. Can I describe what "done" looks like in one sentence?
702. Is there exactly one unknown, or zero?
713. Could a teammate pick this up without a walkthrough?
724. Is it under 4 hours? If no — split again.
73
74**If you cannot decompose a task**, it signals a spike is needed. Timebox the spike (2-4 hours), then re-estimate.
75
76---
77
78## Complexity Multipliers
79
80Apply these multipliers to your base estimate when complexity factors are present. Multipliers stack multiplicatively.
81
82| Factor | Multiplier | Rationale |
83|--------------------------|------------|----------------------------------------------------|
84| New technology / stack | 1.5x | Learning curve, unexpected gotchas, doc-hunting |
85| Unclear requirements | 2.0x | Discovery work, rework cycles, stakeholder alignment |
86| Legacy code | 1.5x | Undocumented behavior, fragile tests, hidden coupling |
87| Cross-team dependency | 1.5x | Coordination overhead, blocking, API negotiation |
88| First-time task | 2.0x | No reference point, unknown unknowns dominate |
89| Regulatory / compliance | 1.5x | Audit trails, review gates, documentation overhead |
90
91**Example:** A 2-day base estimate on legacy code (1.5x) with unclear requirements (2.0x) becomes `2 x 1.5 x 2.0 = 6 days`.
92
93**Rule:** Never apply more than 3 multipliers — if that many factors converge, the task needs a spike or a scope reduction, not a bigger number.
94
95---
96
97## Buffer Calculation
98
99Raw estimates are point predictions. Reality is a distribution.
100
101| Buffer Type | Rule of Thumb | When to Apply |
102|------------------------|-------------------------|-------------------------------------------------|
103| Known unknowns | +20% of total estimate | Integration points, third-party APIs, minor gaps |
104| Unknown unknowns | +50% of total estimate | New domain, first release, greenfield system |
105| Team velocity factor | / focus ratio (e.g., 0.7) | Account for meetings, reviews, context switching |
106| Sequential dependency | +10% per handoff | Each team/person boundary adds coordination drag |
107
108**Effective estimate formula:**
109
110```
111Effective = (Base Estimate x Multipliers) / Focus Ratio + Buffer
112```
113
114**Focus ratio guidelines:**
115
116| Scenario | Typical Focus Ratio |
117|-----------------------------------|---------------------|
118| Dedicated to one project | 0.75-0.85 |
119| Split across 2 projects | 0.50-0.60 |
120| On-call rotation active | 0.60-0.70 |
121| Heavy meeting load (> 3h/day) | 0.45-0.55 |
122
123---
124
125## Historical Calibration
126
127Track actual vs estimated to improve over time. This is the single most effective way to get better at estimation.
128
129**Tracking table:**
130
131| Task | Estimated | Actual | Ratio (A/E) | Notes |
132|---------------------|-----------|--------|-------------|--------------------------|
133| Auth flow | 3 days | 5 days | 1.67 | OAuth docs were outdated |
134| Dashboard charts | 5 days | 4 days | 0.80 | Reused existing component |
135| DB migration | 2 days | 6 days | 3.00 | Discovered data quality issues |
136
137**Accuracy ratio:** Calculate your rolling average of `Actual / Estimated` over the last 10-20 tasks.
138
139- Ratio **< 0.8** — you're overestimating (sandbagging or excessive buffers)
140- Ratio **0.8-1.2** — well calibrated
141- Ratio **> 1.2** — you're underestimating (apply the ratio as a correction factor)
142
143**Calibration action:** Multiply future estimates by your rolling accuracy ratio until it converges toward 1.0.
144
145---
146
147## Common Estimation Biases
148
149Recognize these cognitive traps — awareness alone reduces their effect.
150
151| Bias | Description | Mitigation |
152|---------------------|----------------------------------------------------------|---------------------------------------------------|
153| Planning Fallacy | Assuming best-case scenario despite past evidence | Use historical data, not intuition |
154| Anchoring | First number heard dominates all subsequent estimates | Estimate independently before discussing |
155| Optimism Bias | "It'll be simpler than last time" | Apply the three-point method, honor the pessimistic |
156| Scope Creep | Estimate stays fixed while scope grows | Re-estimate when scope changes, always |
157| Hofstadter's Law | "It always takes longer, even when you account for it" | Add buffer, then add more buffer for novel work |
158| Dunning-Kruger | Novices underestimate; experts sometimes overestimate | Cross-check with a second estimator |
159| Sunk Cost Pressure | Refusing to re-estimate because the original was "approved" | Treat estimates as living artifacts, update often |
160
161---
162
163## Estimation by Task Type
164
165Use these ranges as starting heuristics, then adjust with multipliers and historical data.
166
167| Task Type | Typical Range | Key Variables |
168|---------------------|------------------|------------------------------------------------|
169| Bug fix (isolated) | 2-8 hours | Reproducibility, code familiarity, test coverage |
170| Bug fix (systemic) | 1-3 days | Root cause depth, blast radius, regression risk |
171| Small feature | 1-3 days | Spec clarity, UI complexity, number of endpoints |
172| Medium feature | 3-10 days | Cross-cutting concerns, data model changes |
173| Large feature | 2-4 weeks | Architecture decisions, team coordination |
174| Refactor (local) | 1-3 days | Test coverage, coupling, blast radius |
175| Refactor (systemic) | 1-4 weeks | Number of callers, migration strategy needed |
176| Spike / research | 2-8 hours (timeboxed) | Always timebox — output is knowledge, not code |
177| DevOps / infra | 1-5 days | Provider docs quality, IAM complexity, testing |
178
179---
180
181## Communication
182
183How you present an estimate matters as much as the number itself.
184
185**Always present as a range, never a single number:**
186
187- Bad: "It'll take 5 days."
188- Good: "3-7 days, most likely 5. The range depends on the payment API response format — I'll know more after the spike."
189
190**Confidence levels:**
191
192| Confidence | What It Means | When to Use |
193|------------|--------------------------------------------|------------------------------------|
194| High (+-15%) | Well-understood scope, done similar before | Familiar task, clear spec |
195| Medium (+-30%) | Some unknowns, reasonable decomposition | Most sprint-level estimates |
196| Low (+-50%+) | Significant unknowns, rough order of magnitude | Roadmap forecasts, presale quotes |
197
198**Stakeholder communication rules:**
199
2001. State the range and the confidence level together
2012. Name the top 1-3 risks that could push toward the upper bound
2023. Offer to de-risk with a timeboxed spike before committing
2034. Explicitly state what is **not** included (e.g., "does not include QA, deployment, or docs")
2045. Update estimates proactively when new information surfaces — don't wait until the deadline
205
206---
207
208## Anti-Patterns
209
210| Anti-Pattern | Why It's Harmful | Better Approach |
211|------------------------|-----------------------------------------------------|----------------------------------------------|
212| Padding silently | Erodes trust when discovered; hides real uncertainty | Use explicit buffers with stated rationale |
213| Sandbagging | Destroys velocity data; breeds complacency | Track accuracy ratio, aim for calibration |
214| Not decomposing | Large estimates hide unknowns and compound errors | Break to < 4-hour sub-tasks, estimate bottom-up |
215| Single-point estimates | Implies false certainty, no room for variance | Always give a range with confidence level |
216| Estimating under pressure | Anchoring to what the stakeholder wants to hear | Ask for time to decompose; never estimate on the spot |
217| Copy-paste estimates | Every task has different context and risk profile | Estimate fresh, use references as starting points only |
218| Ignoring rework cycles | First pass is rarely final — reviews, feedback, QA | Factor in at least one review-and-revise loop |
219
220---
221
222## NEVER Do
223
2241. **NEVER give a single-number estimate without a range** — it communicates false precision and sets you up for failure
2252. **NEVER estimate a task you haven't decomposed** — large estimates are guesses wearing a suit
2263. **NEVER let an old estimate stand after scope changes** — estimates are invalidated the moment requirements shift
2274. **NEVER estimate in someone else's units** — your days are not their days; clarify assumptions about focus time and interrupts
2285. **NEVER skip recording actuals** — estimation without feedback is astrology, not engineering
2296. **NEVER commit to an estimate made under pressure** — say "let me break this down and get back to you in an hour"
2307. **NEVER treat an estimate as a promise or a deadline** — estimates are probabilistic forecasts, not contracts