Results for “case-evaluation”

59 skills
memento-teams
Skill Creator
Create new skills, modify and improve existing skills, and measure skill performance through iterative evaluation and benchmarking.
1.5k · bundle
lord1egypt
Sue
Evaluates whether a lawsuit is worth pursuing, explains the litigation process from filing to resolution, and guides case preparation, settlement negotiations, and small claims alternatives.
2
dvy1987
Eval Rubric Design
Design structured evaluation rubrics for scoring LLM and agent outputs — defining quality dimensions, scoring scales, hard gates, score descriptions, and edge cases. Load when the user asks to create an eval rubric, define evaluation criteria, design scoring dimensions, write an eval spec, or says "what should I evaluate", "design a rubric", "create eval criteria", "define quality dimensions", "evaluation rubric for", "how do I measure quality of". Sub-skill of eval-output orchestrator.
3 · bundle
casemark
Case Summary
Produces an attorney-ready memo from a corpus of legal documents supplied by the user. Use when a user shows up with a folder, zip, or vault of case documents and asks for a case summary, case evaluation, litigation package, intake memo, matter overview, or "can you summarize this case for me." The skill ingests the corpus into a searchable index, OCRs anything non-searchable, inventories and diagnoses the practice area, loads the appropriate practice-area playbook module(s) (PI/tort, commercial litigation, IP infringement, or user-authored extensions), iteratively searches the corpus across eight core dimensions plus any module-specific dimensions, defers specialized document clusters (depositions, medical records, discovery, liens) to dedicated sibling skills, and synthesizes a cited memo.
34 · bundle
More results
qhjqhj00
Arc Eval
Benchmarks systems on the Abstraction and Reasoning Corpus (ARC) by requiring inference of abstract transformation rules from few input-output grid demonstrations and application to novel test cases, reporting the fraction of tasks solved.
3
machenjie
Use Case Modeling
`analysis-agent`: use when actors, goals, preconditions, triggers, paths, guarantees, postconditions, or acceptance traces need modeling; skip when no use-case decision exists.
4 · bundle
klotzkette
Case Dpia Drift
Ordnet Akteninhalt, Belege, Lücken und Nachforderungen zu Tatbestandsmerkmalen, Beweisfragen und Beleglage; liefert eine Beweislast- und Substantiierungsmatrix.
1.5k
matrixx0070
Cs Escalation
Package a customer escalation for engineering, product, or leadership as a decision-ready handoff — impact, timeline, facts vs. hypotheses, and a specific ask with an owner.
0
machenjie
Task Handoff Context
Select evidence-bearing context for a downstream consumer after work when claims, artifacts, validation, unresolved decisions, or omissions must cross a boundary.
4 · bundle
rulebase-co
Cx Case Timeline
Use to reconstruct what actually happened to one customer across every conversation, channel and handoff, for an escalation, complaint, post-mortem or goodwill decision. Trigger for "build a timeline for this customer", "what happened on this case", "summarize the back-and-forth before I decide", "when were they first told", escalation summaries, complaint investigations, or a case spanning several tickets.
1
mmehdi0606
Critique
Evaluate design from a UX perspective, assessing visual hierarchy, information architecture, emotional resonance, cognitive load, and overall quality with quantitative scoring, persona-based testing, automated anti-pattern detection, and actionable feedback. Use when the user asks to review, critique, evaluate, or give feedback on a design or component.
2 · bundle
auto-skiller
Analyze Context
Breaks down current state, context, user input, and data into a structured assessment before taking action.
1 · bundle
lambenthan
Exp Eval
实验判决门:Review LLM 独立评判实验结果 → 4 种判决路径 → 自动更新 claims confidence、ideas status、graph edges
77
alirezarezvani
Experiment Designer
Design, prioritize, and evaluate product experiments with clear hypotheses and defensible decisions, including A/B testing, sample size estimation, and statistical interpretation.
20.4k · bundle
michaelschecht
Model Selection
Recommend model families and validation strategy based on data, constraints, and objective. Use when: (1) choosing algorithms, (2) balancing bias/variance, (3) planning benchmark baselines. NOT for: final legal/compliance sign-off.
0
ekatasingh1107
Case Study Builder
Turns client wins into structured case studies in multiple formats. Gathers before/after metrics, structures as Problem-Solution-Results, and outputs landing page copy, one-pager, social snippets, and email insert.
2 · bundle
thedixitjain
Ce Pov
Give a decisive, project-grounded point of view in the subject's own shape: a graded verdict on an external-adoption question, a holistic take on a document, or a position on a user-supplied approach set. Use for a solo POV, a mid-session second opinion, a named-peer cross-check, any request to consult other models or reconcile their opinions, an `oracle` panel, or a correction-cost-gated proactive cross-check offer. Not for findings review (use ce-doc-review), neutral explainers, or generating options (use ce-ideate or ce-brainstorm).
2 · bundle
kursku
Gemini Ask
Deliberative consultation with Gemini CLI. Ask, evaluate, critique, iterate until workable agreement.
55 · bundle
construct-ai-primary
Edge Case Analysis
Use when designing solutions to identify and test boundary conditions, unusual inputs, and uncommon scenarios that could cause failures. This skill provides a systematic approach to finding edge cases before they become bugs in production.
0
akillness
Grill Me
Systematic plan stress-testing through relentless one-question-at-a-time decision-tree interviewing
42
rulebase-co
Cx QA Appeal Process
Use to design or audit a QA dispute and appeal workflow with timeboxes, adjudication standards, and second-level consistency so appeals improve trust instead of rewriting scores without rules. Trigger for "QA appeal process", "agents disputing scores", "who adjudicates QA disputes", "overturn rate too high", second-level review standards, or calibration erosion from ad-hoc score changes.
1
diegojcn
Codex Ask
Deliberative consultation with Codex CLI. Ask, evaluate, critique, iterate until workable agreement.
1 · bundle
kursku
Kimi Ask
Deliberative consultation with Kimi API. Ask, evaluate, critique, iterate until workable agreement.
55 · bundle
muratcankoylan
Advanced Evaluation
Provides production-grade techniques for evaluating LLM outputs using LLMs as judges, covering direct scoring, pairwise comparison, bias mitigation, rubric generation, and confidence calibration.
16.9k · bundle
joshuashepherd
Resume Coach
Reviews resumes for formatting, design, content, and job-specific fit; scores job postings for fit and interview odds; develops rubrics; extracts content from conversations; and audits the resume generation pipeline.
1
intelli-verse-x
Ivx Cf Evaluation
Design and implement evaluation harnesses for models, agents, and code. Use when creating benchmarks, designing eval metrics, or comparing system outputs.
0 · bundle
machenjie
Scenario Decomposition
`analysis-agent`: use when a request needs normal, failure, edge, abuse, recovery, or operational scenarios; skip when no scenario-decomposition decision exists.
4 · bundle
michaelschecht
Model Evaluation
Evaluate model quality with task-appropriate metrics and systematic error analysis. Use when: (1) comparing models, (2) analyzing failures, (3) setting go/no-go thresholds. NOT for: production monitoring implementation.
0
kursku
Claude Ask
Deliberative consultation with another Claude model via Task tool. Opus consults Sonnet, Sonnet consults Opus. Ask, evaluate, critique, iterate until workable agreement.
55 · bundle
owl-listener
Case Study
Craft portfolio-ready case studies that tell the story of a design project, demonstrating process, thinking, and impact.
1.7k
klotzkette
Status
Erstellt zielgruppengerechte Fallstatus-Zusammenfassungen für eine studentische Rechtsberatungsstelle, mit Modi für Mandant, intern und Gericht/Behörde.
1.5k
dvy1987
Setup Evaluation
Validate process decomposition and architecture design quality before execution begins. Load when the setup-evaluator agent fires (automatic for agent-chain tasks), or when user says "evaluate this setup", "check the decomposition", "validate the architecture", "is this plan sound", "review the agent design". Catches structural errors, missing knowledge, unrealistic step ordering, and topology mismatches. Does NOT modify — only evaluates.
3 · bundle
winbda
Case Brief
Write case briefs with facts, issues, and holdings. TRIGGERS - Use when user needs help with case-brief related tasks.
3
machenjie
Failure Diagnosis
`analysis-agent`/`task-agent`/`review-agent`: use when symptoms, logs, metrics, regressions, or incidents need cause analysis; skip when no diagnosis decision exists.
4 · bundle
sbroggioadv
Estrategia Processual
Avalia riscos e oportunidades do caso civel e sugere os proximos passos processuais. Pondera probabilidade de exito, onus da prova, custos/beneficios de acordo x litigio, momento de tutela de urgencia, escolha de pedidos e de via (conhecimento x cumprimento), e calibra a postura. Use quando o operador disser qual a estrategia, vale a pena recorrer, devo fazer acordo, quais os riscos, proximo passo, ou pedir analise estrategica do caso.
6
brycewang-stanford
B2
VS-Enhanced Evidence Quality Appraiser - Prevents Mode Collapse with context-adaptive quality assessment Enhanced VS 3-Phase process: Avoids automatic tool application, delivers research-specific evaluation strategies Use when: appraising study quality, assessing risk of bias, grading evidence Triggers: quality appraisal, RoB, GRADE, Newcastle-Ottawa, risk of bias, methodological quality
1k