Results for “accuracy”

95 skills
More results
curiositech
Helm Liang 2022
Holistic evaluation framework for language models measuring accuracy, calibration, robustness, and fairness
10 · bundle
qhjqhj00
Bbh Eval
Benchmarks zero-shot in-context learning on BIG-Bench Hard multiple-choice tasks, comparing self-generated demonstrations against direct prompting and chain-of-thought baselines, and reports accuracy.
3
qhjqhj00
Bss Eval
Evaluates speech language models on beyond-semantic speech attributes such as dialect comprehension, multi-turn context memory, emotion perception, age-aware response generation, and non-verbal cue handling, reporting accuracy and judge-based scores.
3
ecnu-icalk
Fasttext
编写评估FastText文本分类模型的Python函数,计算accuracy、F1、recall和precision指标,并处理特定格式的标签文本分割。
559
eli-yu-first
Data Quality Checker
Validates data quality with completeness, consistency, accuracy checks and generates quality reports
6 · bundle
matrixx0070
Data Validate
QA a completed analysis for methodology, accuracy, and bias before it is shared, shipped, or acted on.
0
tianhao909
Awq Quantization
Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner.
1 · bundle
qcmuu
Awq Quantization
Activation-aware weight quantization for 4-bit LLM compression with 3x speedup and minimal accuracy loss. Use when deploying large models (7B-70B) on limited GPU memory, when you need faster inference than GPTQ with better accuracy preservation, or for instruction-tuned and multimodal models. MLSys 2024 Best Paper Award winner.
0 · bundle
lionelsimai
Coding Guide
Create medical coding guides for billing accuracy. TRIGGERS - Use when user needs help with coding-guide related tasks.
22
winbda
Coding Guide
Create medical coding guides for billing accuracy. TRIGGERS - Use when user needs help with coding-guide related tasks.
3
galyarderlabs
Copy Editing
Tightens existing copy for clarity, accuracy, tone consistency, and executive readability without altering the underlying message.
20
auto-skiller
Write Docs
Produces high-quality written knowledge artifacts such as READMEs, guides, API specs, tutorials, runbooks, and wikis, with audience-appropriate structure and accuracy.
1 · bundle
oyi77
Zhive
Registers a trading agent on zHive, fetches crypto signals, posts predictions with conviction, and competes for accuracy rewards.
10
johnalbertini14-glitch
Zhive
Registers as a trading agent on zHive, fetches crypto signals, posts predictions with conviction, and competes for accuracy rewards.
1 · bundle
trailofbits
Dimensional Analysis
Orchestrates a dimensional-analysis pipeline to annotate codebases with unit/dimension comments, discover dimensional vocabulary, and detect arithmetic bugs from unit mismatches or precision loss.
6k · bundle
schattenspiegel
Mpmath Python
Use for writing, reviewing, debugging, testing, or validating Python mpmath arbitrary-precision numerical code. Trigger on mpf, mpc, mp.dps, workdps, interval arithmetic, high-precision quadrature, root finding, special functions, matrices, inverse transforms, or precision/convergence failures. Do not use for ordinary NumPy vectorization, SymPy symbolic manipulation, decimal currency arithmetic, or machine-float code with no precision requirement.
0 · bundle
mcollina
Skill Optimizer
Improves AI skills for activation, clarity, and cross-model reliability through benchmarking, salience tuning, and regression triage.
1.9k · bundle
qhjqhj00
Bleurt
Evaluates the correlation between automatic text generation scores and human quality ratings, including robustness to domain and quality drift, using metrics like Kendall's Tau and Pearson correlation.
3
qhjqhj00
Polos
Scores generated image captions against reference captions and source images using the Polos metric, which is trained to align with human judgments and probes hallucination robustness and open-vocabulary evaluation.
3
qhjqhj00
Mdad
Quantifies the minimum accuracy gap needed between two models for a sampled micro-benchmark to reliably preserve their ranking, using the MDAD metric from Yauney et al. (2025).
3
curiositech
Alphago Deep Rl
Strategic patterns for solving intractable problems through cascading approximation, self-improvement, and heterogeneous evaluation from DeepMind's AlphaGo system
10 · bundle
anantha-236
Skill Comply
Visualize whether skills, rules, and agent definitions are actually followed — auto-generates scenarios at 3 prompt strictness levels, runs agents, classifies behavioral sequences, and reports compliance rates with full tool call timelines
1 · bundle
projectious-work
Data Quality
Data quality framework — completeness, accuracy, consistency, validation, and contracts. Use when implementing data validation, setting up quality checks for a pipeline, defining data contracts between teams, or investigating data anomalies.
0
vvieira010-pixel
Confidence Calibration Check
Capture confidence ratings before and after a learning attempt to identify overconfidence and underconfidence patterns. Use when a student wants to understand how well they actually know something versus how well they think they know it.
0
mesteriis
Performance
Improves measured performance while preserving correctness, reliability, and maintainability.
0
projectious-work
Infographics
Creates data-driven infographics and charts as accessible SVG. Use when visualizing data, choosing a chart type, generating an SVG chart or infographic, or reviewing a visualization for clarity and accuracy.
0 · bundle
qhjqhj00
Squad
Computes the SQuAD metric using torchmetrics, given predictions and ground truth. Use when evaluating question-answering outputs with exact match and F1 scores.
3
salacoste
Deep Interview
Socratic deep interview with mathematical ambiguity gating before explicit execution approval
1
qhjqhj00
Geco
Evaluates geometric consistency in text-to-video generation by measuring structural and motion coherence across camera trajectories, detecting deformation and occlusion artifacts in static scenes.
3
qhjqhj00
Cider
Computes CIDEr and related metrics to score how well generated image descriptions align with human consensus, using reference sentences and triplet annotations.
3