Tool Procurement Eval Skill
Tools enter stacks backwards: someone sees a demo, gets excited, and the "evaluation" becomes a justification ritual (vendor-comparison-matrix fights this at the compare stage; this skill fights it at the door). The forward order: the need stated first (which problem, whose, costing what — purchase-justification arithmetic), the stack-fit check before the trial (does something we own already do this? — the overlap audit that kills half of tool requests honestly), the trial designed with success criteria written before day one (or the trial's warm feelings decide), and the security/data review sized to what the tool touches — because the fun tool that ingests customer data is a compliance decision wearing a productivity costume.
What This Skill Produces
- The need statement — the problem, its owner, its cost, and the requirements it implies (must vs. nice — pre-demo)
- The stack-fit audit — the overlap check against owned tools, the integration reality, and the sprawl tax named
- The trial design — duration, participants, the pre-written success criteria, and the decision date
- The verdict — adopt (with owner and rollout) / decline (with the reason logged) — plus the security-review gate where data warrants
Required Inputs
Ask for these if not provided:
- The problem, not the tool — what's broken/slow/manual today, for whom, costing what; "I saw this cool tool" gets reverse-engineered into its implied need, which sometimes evaporates on contact
- The current stack — what's owned that's adjacent (the overlap audit needs the inventory — contract-renewal-tracker's list is the source); most orgs own 30% more capability than they use
- What the tool would touch — customer data? Credentials? Just public content? The security review's depth follows (skill-vetting blast-radius thinking, applied to SaaS)
- The trial population — who'd actually test it (the enthusiast and a skeptic — enthusiast-only trials always pass)
Framework: The Eval Rules
- Need before tool: the statement — problem, owner, frequency, cost of the status quo — written before any vendor contact; requirements derive from it (must-haves gate, nice-to-haves score). Tools without a need statement are solutions shopping for problems on your budget.
- The overlap audit runs early and honestly: does an owned tool cover 80% of the need? (The honest answer kills the request — cheaply and correctly.) Half-used owned tools get their config/training gap named instead ("we own this in Notion; nobody set it up") — the sprawl tax (another login, another admin, another renewal row, another data silo) is a real cost the shiny demo never quotes.
- Trials have pre-written criteria and a skeptic: "success = the team's weekly report time drops below 2 hours, and 4 of 6 pilots choose to keep it" — written before day one (survey-design-basics pre-commitment), tested with the enthusiast and the skeptic (the enthusiast finds the ceiling; the skeptic finds the floor), timeboxed with a decision date. Trials without criteria are extended demos that always end in purchase.
- The security gate scales to the data: public-content tools get the light pass (vendor's security page, the data-processing basics); anything touching customer data, credentials, or internal documents gets the real review (where's the data stored, who can access, the deletion story, the tos-decoder read of their terms) — before the trial pipes real data in, not after. "It's just a trial" is how customer data ends up in un-reviewed vendors.
- The verdict gets logged either way: adopt → owner named, rollout planned, the renewal row created at signature ([the intake rule]) · decline → the reason in the decision-log ("evaluated [tool] July 2026 — declined: 80% covered by owned stack") — because the same tool returns with a new champion every eighteen months, and the log converts the rematch into a lookup.
Output Format
Tool Eval: [tool] — need: [the problem statement]
The Need + Requirements
[Problem/owner/cost · must-haves · nice-to-haves — dated pre-demo]
Stack-Fit Audit
[Overlap: (owned tools × coverage %) · the config-gap finding if applicable · integration reality · the sprawl tax lines]
Trial Design
[Duration · pilots (enthusiast + skeptic named) · the pre-written criteria · decision date · the security gate status before real data]
The Verdict
[Adopt: owner/rollout/renewal-row · Decline: the logged reason · either way: in the decision log]
Quality Checks
Anti-Patterns
1---2name: tool-procurement-eval3description: Evaluate a new tool before it joins the stack — the problem-first framing (tools answer needs, not demos), the trial designed with success criteria upfront, the stack-fit check (integration, overlap, the tool-sprawl tax), and the security/data review sized to the stakes. Use when asked should we buy this tool, evaluate this software for the team, we have three tools that do this already, or run a proper trial before committing. Produces the need statement, the trial design with pre-set criteria, the stack-fit audit, and the adopt/decline verdict with its reasoning.4---5
6# Tool Procurement Eval Skill
7
8Tools enter stacks backwards: someone sees a demo, gets excited, and the "evaluation" becomes a justification ritual ([vendor-comparison-matrix](../vendor-comparison-matrix/SKILL.md) fights this at the compare stage; this skill fights it at the door). The forward order: the *need* stated first (which problem, whose, costing what — [purchase-justification](../purchase-justification/SKILL.md) arithmetic), the *stack-fit* check before the trial (does something we own already do this? — the overlap audit that kills half of tool requests honestly), the *trial designed* with success criteria written before day one (or the trial's warm feelings decide), and the security/data review sized to what the tool touches — because the fun tool that ingests customer data is a compliance decision wearing a productivity costume.
9
10## What This Skill Produces
11
12- **The need statement** — the problem, its owner, its cost, and the requirements it implies (must vs. nice — pre-demo)
13- **The stack-fit audit** — the overlap check against owned tools, the integration reality, and the sprawl tax named
14- **The trial design** — duration, participants, the pre-written success criteria, and the decision date
15- **The verdict** — adopt (with owner and rollout) / decline (with the reason logged) — plus the security-review gate where data warrants
16
17## Required Inputs
18
19Ask for these if not provided:
20- **The problem, not the tool** — what's broken/slow/manual today, for whom, costing what; "I saw this cool tool" gets reverse-engineered into its implied need, which sometimes evaporates on contact
21- **The current stack** — what's owned that's adjacent (the overlap audit needs the inventory — [contract-renewal-tracker](../contract-renewal-tracker/SKILL.md)'s list is the source); most orgs own 30% more capability than they use
22- **What the tool would touch** — customer data? Credentials? Just public content? The security review's depth follows ([skill-vetting](../skill-vetting/SKILL.md) blast-radius thinking, applied to SaaS)
23- **The trial population** — who'd actually test it (the enthusiast *and* a skeptic — enthusiast-only trials always pass)
24
25## Framework: The Eval Rules
26
271. **Need before tool:** the statement — problem, owner, frequency, cost of the status quo — written before any vendor contact; requirements derive from it (must-haves gate, nice-to-haves score). Tools without a need statement are solutions shopping for problems on your budget.
282. **The overlap audit runs early and honestly:** does an owned tool cover 80% of the need? (The honest answer kills the request — cheaply and correctly.) Half-used owned tools get their config/training gap named instead ("we own this in Notion; nobody set it up") — the sprawl tax (another login, another admin, another [renewal](../contract-renewal-tracker/SKILL.md) row, another data silo) is a real cost the shiny demo never quotes.
293. **Trials have pre-written criteria and a skeptic:** "success = the team's weekly report time drops below 2 hours, and 4 of 6 pilots choose to keep it" — written before day one ([survey-design-basics](../survey-design-basics/SKILL.md) pre-commitment), tested with the enthusiast *and* the skeptic (the enthusiast finds the ceiling; the skeptic finds the floor), timeboxed with a decision date. Trials without criteria are extended demos that always end in purchase.
304. **The security gate scales to the data:** public-content tools get the light pass (vendor's security page, the data-processing basics); anything touching customer data, credentials, or internal documents gets the real review (where's the data stored, who can access, the deletion story, the [tos-decoder](../tos-decoder/SKILL.md) read of their terms) — *before* the trial pipes real data in, not after. "It's just a trial" is how customer data ends up in un-reviewed vendors.
315. **The verdict gets logged either way:** adopt → owner named, rollout planned, the renewal row created at signature ([the intake rule]) · decline → the reason in the [decision-log](../decision-log-setup/SKILL.md) ("evaluated [tool] July 2026 — declined: 80% covered by owned stack") — because the same tool returns with a new champion every eighteen months, and the log converts the rematch into a lookup.
32
33## Output Format
34
35# Tool Eval: [tool] — need: [the problem statement]
36
37## The Need + Requirements
38[Problem/owner/cost · must-haves · nice-to-haves — dated pre-demo]
39
40## Stack-Fit Audit
41[Overlap: (owned tools × coverage %) · the config-gap finding if applicable · integration reality · the sprawl tax lines]
42
43## Trial Design
44[Duration · pilots (enthusiast + skeptic named) · the pre-written criteria · decision date · the security gate status before real data]
45
46## The Verdict
47[Adopt: owner/rollout/renewal-row · Decline: the logged reason · either way: in the decision log]
48
49## Quality Checks
50
51- [ ] The need statement predates vendor contact
52- [ ] The overlap audit ran against the real inventory with honest coverage
53- [ ] Trial criteria were written before day one and include a skeptic
54- [ ] The security review preceded real data entering the trial
55- [ ] The verdict is logged with reasons, adopt or decline
56
57## Anti-Patterns
58
59- [ ] Do not evaluate backwards from the demo — the need statement is the eval's spine
60- [ ] Do not skip the overlap audit — the cheapest tool is the one already owned and unconfigured
61- [ ] Do not run criteria-free trials — warm feelings always vote adopt
62- [ ] Do not pipe customer data into "just a trial" — the gate runs first at exactly that moment
63- [ ] Do not decline silently — the unlogged rejection is next year's rematch, at full cost