Evaluate the company the evidence supports, not the story it tells. Distinguish business health, venture suitability, and financing readiness; they are different decisions.
Usage Template
Provide: company, stage, startup type, evaluation decision, customer, problem, product, traction, team, economics, runway, round terms, and known risks. Use the rubric in references/evaluation-rubric.md when scoring is requested.
Workflow
Classify stage, type (SME, innovation-driven, venture-scale, hard-tech, AI-native), lens (founder, diligence, fundraising, pivot), and evidence state. Define the decision and time horizon before calculating a score.
Separate facts, assumptions, self-reported claims, and missing evidence. If the target decision or company identity is unclear, return NEEDS_INPUT. Continue with missing metrics only when the output is explicitly provisional and each gap has a probe.
Rank demand evidence from belief and interviews through behavior, payment, retention, expansion, and referral.
Score eight dimensions using stage-adjusted weights: pain/beachhead, market/timing, value step-change, PMF/traction, business model/economics, team/governance, capital/runway, and moat/risk.
For AI-native or hard-tech cases, test what remains defensible as components cheapen and identify physical, regulatory, deployment, or supply-chain bottlenecks. For AI value capture, separate usage, productivity, customer ROI, and vendor profit; do not infer durable economics from token volume or revenue growth alone.
Separate the spending engine (CapEx, inference, integration, service labor,
energy, and deployment cost) from the earning engine (retention, expansion,
pricing power, gross margin, and free cash flow). Test who owns institutional
learning: workflow exceptions, context, permissions, and feedback write-back.
Diagnose runway and whether spend buys evidence for the next milestone.
Name the single constraint most likely to invalidate or unlock the company.
Specify the cheapest test, threshold, owner, budget, and stop condition.
Do not average away fatal risk. A healthy cash-flow business may still be a poor venture investment; a large market cannot rescue absent demand evidence.
Trace every score and verdict to the evidence ledger. Stress-test the conclusion against churn, paid acquisition dependence, founder conflict, financing timing, platform dependency, falling model prices, rising inference/service cost, and customer ROI that fails to become vendor margin. Calibrate confidence to the weakest decision-critical claim.
Failure Protocol
NEEDS_INPUT: the evaluation decision, stage, or company boundary is unclear.
INSUFFICIENT_EVIDENCE: the requested verdict depends on unavailable demand, retention, economics, or terms data.
VERIFY_FAILED: a score lacks evidence or contradicts the ledger; rescore or mark unknown.
BUDGET_STOP: return the provisional constraint and highest-value evidence request.
Output Contract
Return status, result (verdict, scorecard, top constraint, fatal risks, and next test), evidence (fact/assumption ledger), unknowns, and next_action with owner and threshold.
Edge Cases
Pre-revenue company has no retention data: do not assign traction maturity; score the available behavioral test and make payment/usage the next gate.
Bootstrapped company has strong cash flow but a small market: rate business health separately from venture-scale suitability.
AI usage grows while gross margin and customer retention fall: record adoption
without calling it value capture; test pricing, service labor, and inference
economics separately.
Success Metrics
Stage, type, decision lens, and evidence maturity are explicit.
Verdict confidence follows observed behavior rather than narrative polish.
AI-native verdicts distinguish adoption, customer value, and supplier profit.
One top constraint and one cheap falsifiable test govern the recommendation.
Quality Gates
Facts, assumptions, self-reports, and missing evidence are separated.
Score weights fit stage and business type.
Fatal risks are not hidden by averages.
AI cases separate spending, usage, productivity, customer ROI, and vendor profit.
Runway, milestone, test threshold, owner, and stop condition are explicit.
1---2name: startup-evaluation3description: Use when a startup needs an evidence-weighted health check, investor lens, runway diagnosis, top constraint, or cheapest next validation test.4---56# Startup Evaluation78<skill_contract>9 <input>Company, stage, evaluation decision, customer evidence, traction, economics, team, runway, terms, and risks.</input>10 <output>An evidence-weighted health, venture-suitability, and financing assessment with one top constraint and cheapest test.</output>11 <done>Every score and verdict traces to evidence, fatal risks remain visible, and the next test has owner, threshold, budget, and stop.</done>12 <non_goals>Investment advice by narrative, averaging away fatal risk, or treating market size, conviction, or pitch quality as demand.</non_goals>1314Evaluate the company the evidence supports, not the story it tells. Distinguish business health, venture suitability, and financing readiness; they are different decisions.1516## Usage Template1718Provide: company, stage, startup type, evaluation decision, customer, problem, product, traction, team, economics, runway, round terms, and known risks. Use the rubric in `references/evaluation-rubric.md` when scoring is requested.1920## Workflow2122<intake>2324Classify stage, type (`SME`, innovation-driven, venture-scale, hard-tech, AI-native), lens (founder, diligence, fundraising, pivot), and evidence state. Define the decision and time horizon before calculating a score.2526</intake>2728<unknowns_gate>2930Separate facts, assumptions, self-reported claims, and missing evidence. If the target decision or company identity is unclear, return `NEEDS_INPUT`. Continue with missing metrics only when the output is explicitly provisional and each gap has a probe.3132</unknowns_gate>3334<execute>35361. Rank demand evidence from belief and interviews through behavior, payment, retention, expansion, and referral.372. Score eight dimensions using stage-adjusted weights: pain/beachhead, market/timing, value step-change, PMF/traction, business model/economics, team/governance, capital/runway, and moat/risk.383. For investor work, cross-check 5T: Team, Target Market, Tech/Product, Traction, Terms.394. For AI-native or hard-tech cases, test what remains defensible as components cheapen and identify physical, regulatory, deployment, or supply-chain bottlenecks. For AI value capture, separate **usage**, **productivity**, **customer ROI**, and **vendor profit**; do not infer durable economics from token volume or revenue growth alone.405. Separate the spending engine (CapEx, inference, integration, service labor,41 energy, and deployment cost) from the earning engine (retention, expansion,42 pricing power, gross margin, and free cash flow). Test who owns institutional43 learning: workflow exceptions, context, permissions, and feedback write-back.446. Diagnose runway and whether spend buys evidence for the next milestone.457. Name the single constraint most likely to invalidate or unlock the company.468. Specify the cheapest test, threshold, owner, budget, and stop condition.4748Do not average away fatal risk. A healthy cash-flow business may still be a poor venture investment; a large market cannot rescue absent demand evidence.4950</execute>5152<evaluate>5354Trace every score and verdict to the evidence ledger. Stress-test the conclusion against churn, paid acquisition dependence, founder conflict, financing timing, platform dependency, falling model prices, rising inference/service cost, and customer ROI that fails to become vendor margin. Calibrate confidence to the weakest decision-critical claim.5556</evaluate>5758## Failure Protocol5960- `NEEDS_INPUT`: the evaluation decision, stage, or company boundary is unclear.61- `INSUFFICIENT_EVIDENCE`: the requested verdict depends on unavailable demand, retention, economics, or terms data.62- `VERIFY_FAILED`: a score lacks evidence or contradicts the ledger; rescore or mark unknown.63- `BUDGET_STOP`: return the provisional constraint and highest-value evidence request.6465## Output Contract6667Return `status`, `result` (verdict, scorecard, top constraint, fatal risks, and next test), `evidence` (fact/assumption ledger), `unknowns`, and `next_action` with owner and threshold.6869## Edge Cases7071- Pre-revenue company has no retention data: do not assign traction maturity; score the available behavioral test and make payment/usage the next gate.72- Bootstrapped company has strong cash flow but a small market: rate business health separately from venture-scale suitability.73- AI usage grows while gross margin and customer retention fall: record adoption74 without calling it value capture; test pricing, service labor, and inference75 economics separately.7677## Success Metrics7879- Stage, type, decision lens, and evidence maturity are explicit.80- Verdict confidence follows observed behavior rather than narrative polish.81- AI-native verdicts distinguish adoption, customer value, and supplier profit.82- One top constraint and one cheap falsifiable test govern the recommendation.8384## Quality Gates8586- [ ] Facts, assumptions, self-reports, and missing evidence are separated.87- [ ] Score weights fit stage and business type.88- [ ] Fatal risks are not hidden by averages.89- [ ] AI cases separate spending, usage, productivity, customer ROI, and vendor profit.90- [ ] Runway, milestone, test threshold, owner, and stop condition are explicit.9192</skill_contract>
Run npx skillmds@latest add mark393295827/startup-evaluation in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.
Use when a startup needs an evidence-weighted health check, investor lens, runway diagnosis, top constraint, or cheapest next validation test. It is listed under Coding & Dev Tools on SkillMD.
This skill has not completed SkillMD's automated safety review yet. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.
This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.
Yes. Installing skills from SkillMD is free, and the skill stays under its author's original license.
Mark393295827 (@mark393295827) published this skill. Their other Agent Skills are listed on their SkillMD profile.