all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 61 of 76

  1. ▌
    Mixqg Qg Eval · qhjqhj00
    Evaluates a neural question generation model's ability to produce fluent, relevant questions conditioned on a target answer and context. It probes both automatic n-gram/embedding-based similarity metrics and human-rated utility for educational quiz creation. Use when the user wants to benchmark on SQuAD, NQ, Quoref, DROP, or asks about evaluating this task. Reports question approval rate.
    3 repo stars
  2. ▌
    Mlissard Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform simple sequential reasoning tasks (e.g., counting, copying, list intersection) while extrapolating to longer input sequences. It specifically probes length generalization and the impact of multilingual in-context examples on reasoning robustness across different languages. Use when the user wants to benchmark on MLissard, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  3. ▌
    Mm Graph Eval · qhjqhj00
    Evaluates multimodal graph learning models on node classification, link prediction, and knowledge graph completion tasks. It probes the ability of models to integrate high-resolution visual and textual node features with graph structure to make accurate predictions across diverse real-world domains. Use when the user wants to benchmark on Amazon-Sports, Amazon-Cloth, Goodreads-LP, Goodreads-NC, Ele-fashion, MM-CoDEx-s, MM-CoDEx-m, or asks about evaluating this task. Reports MRR, accuracy.
    3 repo stars
  4. ▌
    Mm Scale Eval · qhjqhj00
    Evaluates vision-language models' ability to perform fine-grained moral reasoning and safety alignment on multimodal scenarios. It probes how well models rank, calibrate, and separate safe from unsafe situations when provided with text, image, or combined modalities. Use when the user wants to benchmark on MM-Scale, or asks about evaluating this task. Reports NDCG@5.
    3 repo stars
  5. ▌
    Mmmu Pro Eval · qhjqhj00
    This benchmark evaluates multimodal models' ability to perform robust, multi-discipline reasoning by forcing them to integrate visual and textual information without relying on shortcuts. It specifically probes resistance to guessing strategies through augmented multiple-choice options and tests true vision-text integration by embedding questions directly within images. Use when the user wants to benchmark on MMMU-Pro, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  6. ▌
    Mobiface Eval · qhjqhj00
    This benchmark evaluates the capability of visual tracking algorithms to maintain robust face localization in unconstrained, mobile-captured video sequences. It specifically probes resilience to challenging real-world conditions such as rapid camera motion, out-of-plane rotations, scale changes, and partial occlusions. Use when the user wants to benchmark on iBUG MobiFace, or asks about evaluating this task. Reports AUC.
    3 repo stars
  7. ▌
    Mobile O Eval · qhjqhj00
    Evaluates a compact on-device unified vision-language-diffusion model's capabilities in multimodal understanding, text-to-image generation, and image editing. It probes the model's ability to align cross-modal representations and generate high-fidelity images while maintaining real-time inference speeds on edge hardware. Use when the user wants to benchmark on GenEval, MMMU, MM-Vet, SEED, TextVQA, ChartQA, POPE, GQA, ImageEdit, or asks about evaluating this task. Reports GenEval overall score.
    3 repo stars
  8. ▌
    Moisesdb Eval · qhjqhj00
    Evaluates the ability of audio source separation models to isolate individual instrument stems from mixed stereo recordings. It probes fine-grained separation capabilities across a hierarchical taxonomy of up to 11 stems, testing robustness to stem imbalance and un-mastered audio characteristics. Use when the user wants to benchmark on MoisesDB, or asks about evaluating this task. Reports SDR.
    3 repo stars
  9. ▌
    Molbench Eval · qhjqhj00
    Evaluates autonomous AI agents' ability to execute complex, multi-step drug discovery workflows, including molecular screening (property filtering, binding affinity comparison, docking) and molecular optimization (structural editing, physicochemical property improvement). Use when the user wants to benchmark on MolBench, or asks about evaluating this task. Reports optimization success rate.
    3 repo stars
  10. ▌
    Molmoweb Eval · qhjqhj00
    Evaluates the capability of vision-language web agents to navigate live websites and complete complex, multi-step tasks using only screenshot inputs. It probes GUI perception, action grounding, and long-horizon planning under real-world web constraints. Use when the user wants to benchmark on WebVoyager, Online-Mind2Web, DeepShop, WebTailBench, ScreenSpot, ScreenSpot v2, or asks about evaluating this task. Reports pass@k accuracy.
    3 repo stars
  11. ▌
    Mosaicml Eval · qhjqhj00
    Evaluates downstream language model capabilities across 33 question-answering tasks. It measures how effectively data pruning strategies improve general performance compared to unpruned baselines, using a normalized accuracy metric that accounts for random guessing baselines. Use when the user wants to benchmark on MosaicML evaluation gauntlet, or asks about evaluating this task. Reports average normalized accuracy.
    3 repo stars
  12. ▌
    Moviesum Eval · qhjqhj00
    Evaluates the ability of abstractive summarization models to generate concise, coherent summaries of long, dispersed movie screenplay narratives. It probes long-document understanding, narrative coherence, and the model's capacity to synthesize information across thousands of tokens. Use when the user wants to benchmark on MovieSum, or asks about evaluating this task. Reports ROUGE F1 (1/2/L).
    3 repo stars
  13. ▌
    Ms Marco Eval · qhjqhj00
    Evaluates machine reading comprehension models on real-world search queries across multiple answer types (numeric, yes/no, descriptive) and tasks (answer generation, span extraction, passage ranking). Probes a model's ability to extract or generate accurate answers from noisy, multi-document web contexts and handle unanswerable questions. Use when the user wants to benchmark on MS MARCO, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  14. ▌
    Msvbench Eval · qhjqhj00
    Evaluates multi-shot video generation models on narrative coherence, cross-shot consistency, visual fidelity, and motion quality. It probes whether models can maintain character and scene identity across sequential shots and adhere to physical laws, rather than merely generating isolated visual interpolations. Use when the user wants to benchmark on MSVBench, or asks about evaluating this task. Reports Spearman’s ρ.
    3 repo stars
  15. ▌
    Mt Bench Eval · qhjqhj00
    Evaluates the conversational quality and instruction-following capability of aligned language models across multiple knowledge domains. It also measures whether alignment fine-tuning causes regression in base reasoning, truthfulness, and commonsense capabilities. Use when the user wants to benchmark on MT-Bench, Open LLM Leaderboard Benchmarks, or asks about evaluating this task. Reports MT-Bench Average Score.
    3 repo stars
  16. ▌
    Mtbbench Eval · qhjqhj00
    Evaluates AI agents' ability to perform longitudinal, multimodal clinical decision-making in oncology. Agents must integrate evolving patient data across pathology, genomics, hematology, and imaging over multiple turns to answer diagnostic and prognostic questions, simulating molecular tumor board workflows. Use when the user wants to benchmark on MTBBench-Multimodal (HANCOCK subset), MTBBench-Longitudinal (MSK-CHORD subset), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    Mtrag Un Eval · qhjqhj00
    Evaluates multi-turn RAG systems on handling unanswerable, underspecified, and non-standalone questions. It probes both retrieval ranking quality and generation quality, including the model's ability to correctly refuse or request clarification when context is insufficient. Use when the user wants to benchmark on MTRAG-UN, or asks about evaluating this task. Reports RB_llm.
    3 repo stars
  18. ▌
    Muben Uq Eval · qhjqhj00
    This benchmark evaluates the uncertainty quantification (UQ) capabilities of molecular representation models across diverse backbone architectures and input modalities. It probes how accurately models predict molecular properties (binary classification and continuous regression) while simultaneously estimating their own predictive uncertainty under both in-distribution and out-of-distribution conditions. The evaluation specifically tests robustness to molecular scaffold shifts, which better simulates real-world drug discovery and materials design scenarios. Use when the user wants to benchmark on MoleculeNet, or asks about evaluating this task. Reports ROC-AUC, RMSE.
    3 repo stars
  19. ▌
    Mulberry Eval · qhjqhj00
    Evaluates the multimodal reasoning and understanding capabilities of MLLMs across diverse domains including mathematics, chart interpretation, scientific/medical images, and hallucination detection. It measures how well models generate step-by-step reasoning paths and reflect on errors to produce correct answers. Use when the user wants to benchmark on MathVista, MMStar, MMMU, ChartQA, DynaMath, HallBench, MM-Math, MMEsum, or asks about evaluating this task. Reports Average Benchmark Score.
    3 repo stars
  20. ▌
    Multiclasseer · qhjqhj00
    Compute the MulticlassEER metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassEER, or asks how to score with MulticlassEER.
    3 repo stars
  21. ▌
    Multiclassroc · qhjqhj00
    Compute the MulticlassROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassROC, or asks how to score with MulticlassROC.
    3 repo stars
  22. ▌
    Multilabeleer · qhjqhj00
    Compute the MultilabelEER metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelEER, or asks how to score with MultilabelEER.
    3 repo stars
  23. ▌
    Multilabelroc · qhjqhj00
    Compute the MultilabelROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelROC, or asks how to score with MultilabelROC.
    3 repo stars
  24. ▌
    Multimed Eval · qhjqhj00
    Evaluates multimodal medical understanding across 11 diverse tasks including disease classification, imaging analysis, genomics, proteomics, and medical VQA. Probes cross-modal integration, generalization to out-of-distribution organs and cell types, and robustness to few-shot and zero-shot scenarios. Use when the user wants to benchmark on MultiMed, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Multinrc Eval · qhjqhj00
    This benchmark evaluates LLMs' ability to perform multi-step reasoning in native non-English languages (French, Spanish, Chinese) across linguistic, wordplay, cultural/tradition, and culturally-grounded math categories. It specifically probes whether models rely on translation bias or possess deep cultural and linguistic contextual knowledge required for accurate problem-solving. Use when the user wants to benchmark on MultiNRC, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  26. ▌
    Multivox Eval · qhjqhj00
    Evaluates voice assistants' ability to jointly ground visual and paralinguistic speech cues (e.g., pitch, emotion, volume, background sounds) in context-aware responses. It specifically tests robustness against confounding samples that flip speech properties to prevent overreliance on unimodal priors. Use when the user wants to benchmark on MultiVox, or asks about evaluating this task. Reports visual grounding and non-verbal speech signals.
    3 repo stars
  27. ▌
    Multiwoz Eval · qhjqhj00
    Evaluates end-to-end task-oriented dialogue systems on their ability to track user goals, fulfill multi-domain requests, and generate contextually appropriate responses. It specifically probes how well models maintain conversation state and achieve user objectives without relying on full historical dialogue context. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports inform rate, success rate.
    3 repo stars
  28. ▌
    Muri 101 Eval · qhjqhj00
    Evaluates multilingual instruction-following capabilities across Natural Language Understanding (NLU) and open-ended generation (NLG) tasks, specifically probing performance on low-resource and multilingual settings using translated and native benchmarks. Use when the user wants to benchmark on Multilingual MMLU, TranslatedDolly, Taxi1500, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  29. ▌
    Musicgen Eval · qhjqhj00
    Evaluates the capability of text-to-music generation models to produce high-fidelity, controllable audio that aligns with textual descriptions and matches human perceptual quality standards. Use when the user wants to benchmark on MusicCaps, or asks about evaluating this task. Reports FAD.
    3 repo stars
  30. ▌
    Musicsem Eval · qhjqhj00
    Evaluates multimodal models on their ability to understand, generate, and retrieve music based on semantically rich, context-aware natural language descriptions. It probes fine-grained musical semantics beyond technical attributes, including atmospheric, situational, and contextual cues. Use when the user wants to benchmark on MusicSem, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  31. ▌
    Must RAG Eval · qhjqhj00
    This evaluation probes a model's ability to answer music-specific factual and contextual questions using retrieval-augmented generation. It measures accuracy on both in-domain artist metadata and out-of-domain music knowledge across multiple-choice formats. Use when the user wants to benchmark on ArtistMus, TrustMus, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  32. ▌
    Mvtec Ad Eval · qhjqhj00
    Evaluates an image classification and segmentation model's ability to detect and localize defects in industrial products without seeing anomalous examples during training. It probes the model's capacity to learn nominal feature distributions and identify deviations at both image and pixel levels. Use when the user wants to benchmark on MVTec AD, Magnetic Tile Defects (MTD), Mini Shanghai Tech Campus (mSTC), or asks about evaluating this task. Reports AUROC.
    3 repo stars
  33. ▌
    Myriadal Eval · qhjqhj00
    Evaluates active few-shot learning frameworks for histopathology image classification under extremely tight annotation budgets (1, 5, and 10 labeled samples). It probes how well uncertainty-based diversity sampling and self-supervised contrastive pretraining can reduce sample redundancy and improve classification performance compared to standard few-shot and active learning baselines. Use when the user wants to benchmark on NCT-CRC-HE-100K, BreaKHis, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  34. ▌
    Nanoknow Eval · qhjqhj00
    Evaluates how pre-training data exposure and external context influence closed-book and open-book question answering accuracy. It probes the model's reliance on parametric knowledge versus retrieved evidence, and measures the impact of answer frequency and distractors. Use when the user wants to benchmark on Natural Questions, SQuAD, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  35. ▌
    Tpc 268 Eval · qhjqhj00
    Evaluates class-agnostic counting (CAC) capabilities on fine-grained, taxonomically diverse plant species under few-shot exemplar guidance. It probes models' ability to handle dense occlusion, multi-scale observation, and cross-domain generalization from generic objects to biological imagery. Use when the user wants to benchmark on TPC–268, or asks about evaluating this task. Reports MAE.
    3 repo stars
  36. ▌
    Tracsum Eval · qhjqhj00
    Evaluates a model's ability to generate aspect-specific summaries from clinical abstracts and accurately cite the supporting source sentences. It probes factual recall, conciseness, and traceability in a medical domain setting. Use when the user wants to benchmark on TracSum, or asks about evaluating this task. Reports Claim Recall.
    3 repo stars
  37. ▌
    Trec Dl Eval · qhjqhj00
    Evaluates document retrieval models on a large-scale corpus by ranking millions of documents against a set of queries. It probes the model's ability to perform full-document retrieval and produce accurate ranked lists using explicit and latent matching signals. Use when the user wants to benchmark on TREC Deep Learning track (MS MARCO), or asks about evaluating this task. Reports MRR.
    3 repo stars
  38. ▌
    Trec Ir Eval · qhjqhj00
    Evaluates information retrieval ranking methods on ad hoc and routing tasks across diverse text collections. It probes how well term weighting schemes capture semantic relevance and handle difficult or verbose queries compared to classic baselines like BM25 and TF-IDF. Use when the user wants to benchmark on TREC Collections and Topics, or asks about evaluating this task. Reports MAP.
    3 repo stars
  39. ▌
    Tsmd F1 Eval · qhjqhj00
    Evaluates the quality of discovered motif sets in time series by measuring alignment with ground truth segments. It accounts for variable-length patterns and time warping while penalizing both false discoveries and missed patterns. Use when the user wants to benchmark on TSMD benchmark datasets, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  40. ▌
    Twnertc Eval · qhjqhj00
    Evaluates the annotation quality of the TWNERTC corpus for Turkish named entity recognition (NER) and text categorization (TC) by comparing automated labels against human-verified ground truths. It measures how well coarse-grained and fine-grained entity types, as well as domain categories, align with human judgment across different noise-reduction post-processing variants. Use when the user wants to benchmark on TWNERTC, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  41. ▌
    Ubb Ntp Eval · qhjqhj00
    Evaluates the predictive accuracy and computational efficiency of a user-behavior-based network traffic forecasting method against statistical and neural network baselines on real-world SMS traffic data. Use when the user wants to benchmark on Guangzhou and Milan SMS datasets, or asks about evaluating this task. Reports R2.
    3 repo stars
  42. ▌
    Ultrahr Eval · qhjqhj00
    Evaluates text-to-image diffusion models on ultra-high-resolution generation, probing semantic alignment with prompts, fine-grained texture preservation, and overall perceptual quality at resolutions ≥4096px. Use when the user wants to benchmark on UltraHR-eval4K, Aesthetic-Eval@4096, or asks about evaluating this task. Reports FID.
    3 repo stars
  43. ▌
    Unieval Eval · qhjqhj00
    Evaluates natural language generation models across multiple quality dimensions (e.g., coherence, fluency, consistency, relevance) by reframing assessment as a Boolean QA task. Measures how well automated scores align with human judgments using correlation metrics. Use when the user wants to benchmark on SummEval, Topical-Chat, SFRES, SFHOT, QAGS, or asks about evaluating this task. Reports Spearman correlation.
    3 repo stars
  44. ▌
    Unifact Eval · qhjqhj00
    Evaluates large language models' factual correctness by dynamically generating responses to factual questions and measuring how well hallucination detection and fact verification methods can predict the ground-truth factuality label of those responses. It probes a model's susceptibility to hallucination and the effectiveness of external evidence retrieval in verifying generated claims. Use when the user wants to benchmark on TriviaQA, NQ-Open, PopQA, 2WikiMultihopQA, HotpotQA, or asks about evaluating this task. Reports factuality prediction accuracy.
    3 repo stars
  45. ▌
    Usc Asr Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) systems on Uzbek language audio by measuring character and word error rates against manually transcribed ground truth. It probes the model's ability to accurately transcribe low-resource speech data without relying on external linguistic resources or pronunciation dictionaries. Use when the user wants to benchmark on USC, or asks about evaluating this task. Reports WER.
    3 repo stars
  46. ▌
    Vabench Eval · qhjqhj00
    Evaluates the quality, cross-modal consistency, synchronization, and spatial audio rendering of text-to-audio-video and image-to-audio-video generation models. It probes physical plausibility, emotional expressiveness, and stereo separation across seven real-world sound categories. Use when the user wants to benchmark on VABench, or asks about evaluating this task. Reports Audio-Visual Align.
    3 repo stars
  47. ▌
    Validitysoft · qhjqhj00
    Evaluates the faithfulness and semantic plausibility of model-agnostic XAI techniques by generating soft counterfactuals via token-level perturbations. It measures whether perturbations actually change model predictions and whether the generated explanations align with the true causal impact of those changes. Use when the user has predictions and gold and needs to compute Validitysoft, Csoft.
    3 repo stars
  48. ▌
    Vc Inspector · qhjqhj00
    Evaluates the factual accuracy and overall quality of video captions in a reference-free setting. It measures how well a model's predicted quality scores and explanations align with human judgments across diverse video and image domains. Use when the user has predictions and gold and needs to compute Kendall's correlation ($\tau_b$).
    3 repo stars
  49. ▌
    Vcbench Eval · qhjqhj00
    Evaluates multimodal mathematical reasoning capabilities of vision-language models, specifically focusing on vision-centric elementary math problems that require explicit visual dependencies across multiple images. It probes spatial, temporal, geometric, logical, and pattern recognition skills to measure how well models integrate cross-modal information for compositional reasoning. Use when the user wants to benchmark on VCBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  50. ▌
    Vcc2018 Eval · qhjqhj00
    Evaluates voice conversion systems on speech naturalness and target speaker similarity using crowdsourced perceptual tests, and assesses linguistic consistency via automatic speech recognition word error rates. It covers both parallel (Hub) and non-parallel (Spoke) conversion tasks. Use when the user wants to benchmark on VCC2018, or asks about evaluating this task. Reports Naturalness (MOS).
    3 repo stars
  51. ▌
    Vilbias Eval · qhjqhj00
    This benchmark probes a model's ability to detect framing bias in multimodal news content (text-image pairs) and generate grounded, correct rationales for its decisions. It evaluates both closed-ended classification accuracy and open-ended reasoning quality using an LLM-as-judge protocol. Use when the user wants to benchmark on ViLBias, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  52. ▌
    Visdial Eval · qhjqhj00
    Evaluates an AI agent's ability to maintain conversational context, resolve co-references, and ground follow-up questions in visual content. The task requires ranking a set of candidate answers based on an image and dialog history. Use when the user wants to benchmark on VisDial v0.9, or asks about evaluating this task. Reports MRR.
    3 repo stars
  53. ▌
    Vlguard Eval · qhjqhj00
    Evaluates the safety alignment and helpfulness of vision-language models (VLLMs) by measuring their ability to reject harmful image-text prompts while maintaining performance on benign queries. Use when the user wants to benchmark on VLGuard, or asks about evaluating this task. Reports ASR.
    3 repo stars
  54. ▌
    Vn Mteb Eval · qhjqhj00
    Evaluates the quality of Vietnamese text embeddings across six standard information retrieval and NLP tasks. It probes a model's ability to capture semantic similarity, perform document retrieval, classify text, cluster documents, and rank relevant passages in Vietnamese. Use when the user wants to benchmark on VN-MTEB, or asks about evaluating this task. Reports Average Task Score.
    3 repo stars
  55. ▌
    Voxeval Eval · qhjqhj00
    VoxEval probes the knowledge understanding and mathematical reasoning capabilities of end-to-end spoken language models (SLMs). It specifically evaluates how well these models comprehend spoken questions and generate accurate spoken answers under diverse audio conditions, including different speakers, speaking styles, and audio qualities. Use when the user wants to benchmark on VoxEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  56. ▌
    Vqa Gen Eval · qhjqhj00
    This benchmark evaluates a model's ability to generalize in Visual Question Answering under coordinated visual and textual distribution shifts. It probes robustness to image corruptions, style transfers, and linguistic variations by measuring in-domain and cross-domain accuracy. Use when the user wants to benchmark on VQA-GEN, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  57. ▌
    We Math Eval · qhjqhj00
    Evaluates the visual mathematical reasoning capabilities of Large Multimodal Models (LMMs). It probes their ability to decompose composite problems, apply hierarchical knowledge concepts, and reason through multi-step visual math tasks without relying on rote memorization. Use when the user wants to benchmark on We-Math testmini, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  58. ▌
    Wearvox Eval · qhjqhj00
    Evaluates speech-language models on egocentric, multi-channel wearable audio tasks, probing their ability to handle noisy real-world acoustic conditions, reject side-talk, execute tool calls, answer questions with or without context, and translate speech in conversational settings. Use when the user wants to benchmark on WearVox, or asks about evaluating this task. Reports Turn-basedMicro-avg.
    3 repo stars
  59. ▌
    Wikiasp Eval · qhjqhj00
    Evaluates multi-domain aspect-based summarization, requiring models to first discover relevant aspects (Wikipedia section titles) from cited references and then generate domain-specific summaries. It probes content selection, cross-document pronoun resolution, and temporal ordering in multi-source generation. Use when the user wants to benchmark on WikiAsp, or asks about evaluating this task. Reports R-2.
    3 repo stars
  60. ▌
    Wikihow Eval · qhjqhj00
    Evaluates text summarization systems on procedural, step-by-step articles written by non-journalists. It probes the model's ability to handle long sequences, non-inverted-pyramid structures, and high-abstraction content compared to standard news datasets. Use when the user wants to benchmark on WikiHow, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  61. ▌
    Wikisql Eval · qhjqhj00
    Evaluates a model's ability to translate natural language questions into executable SQL queries against a given database schema. It probes semantic parsing, table understanding, and the generation of syntactically and semantically correct structured queries. Use when the user wants to benchmark on WikiSQL, or asks about evaluating this task. Reports Acc_ex.
    3 repo stars
  62. ▌
    Wildasr Eval · qhjqhj00
    This benchmark probes the robustness of automatic speech recognition (ASR) systems under realistic, out-of-distribution conditions. It specifically evaluates performance degradation across environmental noise, demographic shifts (accent, age, child speech), and linguistic diversity (short, incomplete, code-switched utterances), while also measuring semantic hallucination rates beyond standard lexical error metrics. Use when the user wants to benchmark on WildASR, or asks about evaluating this task. Reports WER/CER.
    3 repo stars
  63. ▌
    Wildsci Eval · qhjqhj00
    Evaluates language models' scientific reasoning capabilities by testing their ability to answer domain-specific multiple-choice questions derived from peer-reviewed literature and established scientific benchmarks. Use when the user wants to benchmark on WildSci-Val, GPQA-Aug, SuperGPQA, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  64. ▌
    Wmt Ape Eval · qhjqhj00
    Evaluates a model's ability to perform Automatic Post-Editing (APE) by correcting machine-translated German sentences using the original source text. It probes error detection, grammatical correction, and word-copying capabilities in a multi-source sequence-to-sequence setting. Use when the user wants to benchmark on WMT APE, or asks about evaluating this task. Reports case-sensitive BLEU.
    3 repo stars
  65. ▌
    Wmt24 Eval · qhjqhj00
    Evaluates machine translation systems across 55 languages and dialects using automatic metrics and significance testing to compare translation quality across different domains and language pairs. Use when the user wants to benchmark on WMT24++, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  66. ▌
    Wordinfolost · qhjqhj00
    Compute the WordInfoLost metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute WordInfoLost, or asks how to score with WordInfoLost.
    3 repo stars
  67. ▌
    Worldqa Eval · qhjqhj00
    Evaluates multimodal video understanding and long-chain reasoning by requiring models to integrate visual, auditory, and external world knowledge to answer open-ended and multiple-choice questions. Use when the user wants to benchmark on WorldQA, or asks about evaluating this task. Reports GPT-4 open-ended score.
    3 repo stars
  68. ▌
    X Topic Eval · qhjqhj00
    This benchmark evaluates multilingual topic classification on social media tweets across four languages (English, Spanish, Japanese, Greek). It probes models' ability to generalize across languages and training regimes, including zero-shot, few-shot, monolingual, cross-lingual, and multilingual fine-tuning settings. Use when the user wants to benchmark on X-Topic, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  69. ▌
    Xrbench Eval · qhjqhj00
    Evaluates ML inference accelerators under realistic Extended Reality (XR) workloads. It probes the system's ability to handle real-time, multi-task, multi-model (MTMM) pipelines with dynamic dependencies while meeting strict latency, energy, and quality-of-experience (QoE) constraints. Use when the user wants to benchmark on XRBench Scenarios, or asks about evaluating this task. Reports overall score.
    3 repo stars
  70. ▌
    Xtragpt Eval · qhjqhj00
    Evaluates an LLM's ability to perform context-aware, instruction-guided revisions of academic paper sections. It probes controllable editing capabilities, specifically measuring adherence to revision instructions, clarity, conciseness, and alignment with scientific writing standards through automated pairwise comparisons and human scoring. Use when the user wants to benchmark on XtraQA, or asks about evaluating this task. Reports Length-controlled (LC) win rate.
    3 repo stars
  71. ▌
    Paperqa Local RAG · qhjqhj00 bundle
    Run PaperQA2 locally on a folder of scientific PDFs to get high-accuracy, fully-cited answers. Self-hosted, open-source RAG (Apache-2.0) — needs only an LLM key (OpenAI/Anthropic/local), no FutureHouse credits. Use when the user has a local corpus of papers and wants grounded answers, or wants to avoid the hosted Crow/Falcon for privacy / cost reasons.
    3 repo stars
  72. ▌
    Phoenix Chemistry · qhjqhj00 bundle
    Cheminformatics-grounded chemistry agent (Phoenix, the successor to ChemCrow) via the FutureHouse Platform. Use for retrosynthesis, reaction planning, molecular property prediction, SMILES manipulation, and proposing new molecules with chemistry tools backing the reasoning. Trigger on chemistry / drug-design / synthesis / molecule questions.
    3 repo stars
  73. ▌
    3mdbench Eval · qhjqhj00
    Evaluates Large Vision-Language Models in realistic telemedicine consultations by simulating multi-agent dialogues between a doctor and a temperament-based patient. It probes diagnostic accuracy from multimodal inputs (images + text) and assesses clinical competence and dialogue quality. Use when the user wants to benchmark on 3MDBench, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  74. ▌
    4seasons Eval · qhjqhj00
    Evaluates visual SLAM and long-term localization for autonomous driving under challenging cross-season, multi-weather, and long-term environmental changes. Specifically probes visual odometry, global place recognition, and map-based visual localization capabilities. Use when the user wants to benchmark on 4Seasons, or asks about evaluating this task. Reports horizontal RMSE.
    3 repo stars
  75. ▌
    Abhinaw Score · qhjqhj00
    Evaluates text fidelity and typography accuracy in AI-generated images by measuring spelling, case sensitivity, repetition, and structural inconsistencies against reference prompts. Use when the user has predictions and gold and needs to compute ABHINAW Score.
    3 repo stars
  76. ▌
    Adte Tta Eval · qhjqhj00
    Evaluates test-time adaptation (TTA) capabilities of vision-language models under distribution shift and class imbalance. It probes how well a model can adapt to out-of-distribution and cross-domain image classification tasks without training, using adaptive entropy-based uncertainty estimation to select confident augmented views. Use when the user wants to benchmark on ImageNet & Cross-Domain Benchmarks, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  77. ▌
    Afrimteb Eval · qhjqhj00
    Evaluates text embedding models on a wide range of African language tasks, including classification, retrieval, semantic similarity, clustering, and bitext mining. It probes cross-lingual transfer, language coverage, and the ability of embeddings to capture semantic and discriminative signals across 59 African languages. Use when the user wants to benchmark on AfriMTEB, AfriMTEB-Lite, or asks about evaluating this task. Reports macro average score.
    3 repo stars
  78. ▌
    Afro Nmt Eval · qhjqhj00
    Evaluates neural machine translation performance across five low-resource African languages (Swahili, Amharic, Tigrigna, Oromo, Somali) paired with English. It probes model robustness to domain shifts and compares single-language, semi-supervised, transfer-learning, and multilingual training strategies. Use when the user wants to benchmark on AfroNMT, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  79. ▌
    Agentehr Eval · qhjqhj00
    Evaluates autonomous clinical decision-making agents on Electronic Health Record (EHR) data. It probes multi-step reasoning, long-context dependency preservation, and robustness to distribution shifts across different hospital databases and clinical event types. Use when the user wants to benchmark on MIMIC-IV / MIMIC-III, or asks about evaluating this task. Reports average score.
    3 repo stars
  80. ▌
    Aidovecl Eval · qhjqhj00
    Evaluates the effectiveness of AI-generated outpainted vehicle images as data augmentation for training object detection models. It probes the model's ability to generalize to real-world vehicle classification and bounding box localization when trained on synthetically augmented data. Use when the user wants to benchmark on AIDOVECL augmented dataset, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  81. ▌
    Align Vl Eval · qhjqhj00
    Evaluates a vision-language model's cross-modal retrieval capabilities (matching images to text and vice versa) and zero-shot image classification performance without task-specific fine-tuning. It also measures transfer learning effectiveness on downstream visual benchmarks via linear probing and full fine-tuning. Use when the user wants to benchmark on Flickr30K, MSCOCO, or asks about evaluating this task. Reports R@10.
    3 repo stars
  82. ▌
    Almanacs Eval · qhjqhj00
    Evaluates whether language model explanations (e.g., weights, qualitative descriptions) enable a second predictor model to accurately simulate and predict the behavior of a synthetic linear model across safety-relevant scenarios. The benchmark specifically probes simulatability and robustness to distributional shift between training and test variable values. Use when the user wants to benchmark on ALMANACS Synthetic Dataset, or asks about evaluating this task. Reports probability.
    3 repo stars
  83. ▌
    Alope Qe Eval · qhjqhj00
    This evaluation probes the ability of LLM-based frameworks to estimate translation quality in a reference-free setting by predicting continuous quality scores for source-target sentence pairs across multiple low-resource language directions. It specifically tests how intermediate Transformer layer representations and adaptive regression heads improve cross-lingual alignment and quality prediction compared to standard fine-tuning or zero-shot prompting. Use when the user wants to benchmark on Low-resource QE language pairs (En-Gu, En-Hi, En-Mr, En-Ta, En-Te, Et-En, Ne-En, Si-En), or asks about evaluating this task. Reports Spearman correlation.
    3 repo stars
  84. ▌
    Alpbench Eval · qhjqhj00
    Evaluates active learning pipelines by comparing query strategies paired with tabular classifiers across multiple datasets. It measures how efficiently pipelines improve test performance as the labeled data budget increases, highlighting the interplay between learner choice and query strategy. Use when the user wants to benchmark on OpenML-CC18 and TabZilla Benchmark Suite, or asks about evaluating this task. Reports AUBC (Area Under the Budget Curve).
    3 repo stars
  85. ▌
    Animal3d Eval · qhjqhj00
    This benchmark evaluates the ability of deep learning models to estimate 3D pose and shape of diverse mammal species from single images. It probes cross-species generalization, synthetic-to-real transfer, and the adaptation of human-centric pose estimation architectures to non-human anatomies. Use when the user wants to benchmark on Animal3D, or asks about evaluating this task. Reports S-MPJPE.
    3 repo stars
  86. ▌
    Apex Mem Eval · qhjqhj00
    Evaluates long-term conversational memory, temporal reasoning, and factual consistency across multi-session dialogues and noisy search-augmented contexts. Probes an agent's ability to retrieve, resolve temporal conflicts, and answer complex queries over extended interaction histories. Use when the user wants to benchmark on LOCOMO, LongMemEval, SealQA-Hard, or asks about evaluating this task. Reports LOCOMO Overall Accuracy.
    3 repo stars
  87. ▌
    Ar Bench Eval · qhjqhj00
    Evaluates large language models on appellate review tasks for criminal judgments, specifically detecting, classifying, and correcting legal errors in finalized court decisions. It probes fine-grained legal reasoning, diagnostic accuracy, and the ability to generate legally valid corrections. Use when the user wants to benchmark on AR-Bench, or asks about evaluating this task. Reports Accuracy (Acc), Macro F1 (MaF1).
    3 repo stars
  88. ▌
    Artefact Eval · qhjqhj00
    Evaluates the ability of segmentation models (CNNs, Transformers, diffusion models, and vision foundation models) to detect and classify diverse damage types on analogue media across different material and content categories. It probes cross-media generalization using a leave-one-out protocol and tests the effectiveness of zero-shot, supervised, and text-guided prompting strategies for pixel-level damage localization. Use when the user wants to benchmark on ARTeFACT, or asks about evaluating this task. Reports macro-averaged F1 Score.
    3 repo stars
  89. ▌
    Arxivcap Eval · qhjqhj00
    Evaluates large vision-language models' ability to comprehend and generate text for scientific figures. It probes capabilities in single and multi-figure captioning, contextualized captioning using in-context examples, and inferring paper titles from figure-caption sequences. Use when the user wants to benchmark on ArXivCap, or asks about evaluating this task. Reports BLEU-2.
    3 repo stars
  90. ▌
    Askbench Eval · qhjqhj00
    Evaluates LLMs' ability to detect intent deficiencies or overconfidence in user queries and request targeted clarification during multi-turn interactive QA. It measures how well models balance asking clarifying questions versus providing final answers, using rubric-based checkpoints to score clarification quality and final answer accuracy. Use when the user wants to benchmark on AskBench, HealthBench, or asks about evaluating this task. Reports single-turn accuracy.
    3 repo stars
  91. ▌
    Asnm Tun Eval · qhjqhj00
    Evaluates classifier robustness against tunneling and non-payload-based adversarial obfuscations in network traffic. It tests whether models trained on direct attacks can detect obfuscated variants and how training data augmentation with obfuscated samples improves detection. Use when the user wants to benchmark on ASNM-TUN, or asks about evaluating this task. Reports F1-measure.
    3 repo stars
  92. ▌
    Asr4real Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) models on their robustness and fairness across diverse real-world conditions, including accented speech, rehearsed speech, and spontaneous conversational speech. It specifically probes performance disparities related to speaker accent, gender, and socio-economic background. Use when the user wants to benchmark on ALLSSTAR, NISP, VoxPopuli, Buckeye, CORAAL, or asks about evaluating this task. Reports WER.
    3 repo stars
  93. ▌
    Atrbench Eval · qhjqhj00
    Evaluates SAR automatic target recognition (ATR) capabilities under realistic, wild conditions. It probes fine-grained vehicle classification and detection robustness across varying imaging geometries, scene complexities, and domain shifts (SOC vs EOC settings). Use when the user wants to benchmark on ATRBench, or asks about evaluating this task. Reports overall accuracy (%).
    3 repo stars
  94. ▌
    Audeeter Eval · qhjqhj00
    This benchmark probes the ability of deepfake audio detectors to generalise to novel synthesis systems and diverse human voice corpora in open-world settings. It specifically evaluates robustness against domain shifts in both speech synthesis methods and real audio sources, revealing how well models handle unseen acoustic patterns and distribution shifts. Use when the user wants to benchmark on AUDETER, or asks about evaluating this task. Reports Equal Error Rate (EER).
    3 repo stars
  95. ▌
    Autofish Eval · qhjqhj00
    Fine-grained instance segmentation and length estimation of visually similar fish species under realistic conveyor-belt conditions. The benchmark evaluates model robustness across separated, touching, and occluded fish configurations, using group-based splits to prevent data cross-contamination. Use when the user wants to benchmark on AutoFish, or asks about evaluating this task. Reports mAP.
    3 repo stars
  96. ▌
    Averitec Eval · qhjqhj00
    Evaluates a model's ability to verify real-world claims by retrieving web evidence, generating supporting questions, predicting veracity stance, and producing textual justifications. It probes retrieval quality, stance detection, and justification generation under realistic conditions with temporal and context constraints. Use when the user wants to benchmark on AVeriTeC, or asks about evaluating this task. Reports Macro-F1.
    3 repo stars
  97. ▌
    Avsd Dst Eval · qhjqhj00
    Evaluates a model's ability to perform dialogue state tracking in an open-domain, multimodal setting by framing it as a question-answering task. It measures how well the system tracks conversation state and generates accurate answers based on video/audio context and dialogue history. Use when the user wants to benchmark on AVSD, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  98. ▌
    Badrobot Eval · qhjqhj00
    Evaluates the safety and alignment of embodied LLMs by measuring their susceptibility to voice-based and text-based adversarial prompts that induce harmful physical actions, privacy violations, or fraud. It probes cascading jailbreaks, safety misalignment between language and action, and gaps in physical world knowledge. Use when the user wants to benchmark on BadRobot Physical Action Benchmark, or asks about evaluating this task. Reports MSR (Manipulate Success Rate).
    3 repo stars
  99. ▌
    Bench360 Eval · qhjqhj00
    Evaluates local LLM inference across multiple dimensions, including task-specific quality (e.g., accuracy, F1, ROUGE) and system-level performance (latency, throughput, energy, memory, cold-start) under simulated workloads (single-stream, batch, server). Use when the user wants to benchmark on mmlu, squad_v2, cnn_dailymail, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Benchecg Eval · qhjqhj00
    Evaluates ECG foundation models on diverse clinical tasks including classification, regression, detection, and survival analysis across multiple populations and signal lengths. It probes the model's ability to generalize across datasets, modalities (ECG vs PPG), and long-context temporal dependencies. Use when the user wants to benchmark on CODE-15%, Sleep-Apnea-ECG, MIT-BIH Arrhythmia, PTB-XL, CPSC2018, MIMIC-IV-ECG, Exercise-ECG, or asks about evaluating this task. Reports BenchECG score.
    3 repo stars