all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 72 of 76

  1. ▌
    Famteb Eval · qhjqhj00
    Evaluates the effectiveness of text embedding models across seven diverse tasks (classification, clustering, pair classification, reranking, retrieval, semantic textual similarity, and summary retrieval) specifically for the Persian language. Use when the user wants to benchmark on FaMTEB, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  2. ▌
    Fbeta Score · qhjqhj00
    Compute the fbeta_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute fbeta_score, or asks how to score with fbeta_score.
    3 repo stars
  3. ▌
    Fg Cxr Eval · qhjqhj00
    Evaluates a model's ability to generate accurate, clinically correct chest X-ray reports while aligning its visual focus with radiologist gaze patterns. It probes both natural language generation quality and interpretable attention prediction to ensure diagnostic reasoning is visually grounded. Use when the user wants to benchmark on FG-CXR, or asks about evaluating this task. Reports C, F1_ex, fwIoU.
    3 repo stars
  4. ▌
    Fig QA Eval · qhjqhj00
    Evaluates language models' ability to interpret nonliteral, creative metaphors by selecting the correct literal meaning from two opposing options, and generating sensible interpretations for novel metaphors. It probes commonsense grounding and contextual understanding beyond literal paraphrase tasks. Use when the user wants to benchmark on Fig-QA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  5. ▌
    Fingen Eval · qhjqhj00
    Evaluates forward-looking argument generation in finance across text-to-claim, chart-to-argument, and news-to-argument tasks. Probes a model's ability to generate plausible, structured future scenarios and claims based on financial inputs while maintaining factual consistency and handling financial terminology and numerals. Use when the user wants to benchmark on FinGen, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  6. ▌
    Finmme Eval · qhjqhj00
    Evaluates financial multi-modal reasoning capabilities of models on chart-based analysis and domain-specific knowledge. It probes perception, analysis, and reasoning across 18 financial domains and 6 asset classes using multiple-choice and computational problems. Use when the user wants to benchmark on FinMME, or asks about evaluating this task. Reports FinScore.
    3 repo stars
  7. ▌
    Finn R Eval · qhjqhj00
    Evaluates the performance, power, and resource efficiency of quantized neural networks deployed on various FPGA platforms using the FINN-R framework. It probes the trade-offs between network precision, hardware resource usage, throughput, and classification accuracy across embedded and datacenter-scale hardware. Use when the user wants to benchmark on MNIST, CIFAR-10, GTSRB, SVHN, VOC 2007, ImageNet, or asks about evaluating this task. Reports Top-1 Accuracy.
    3 repo stars
  8. ▌
    Finset Eval · qhjqhj00
    Evaluates financial LLMs across seven text-based tasks (sentiment analysis, NER, number understanding, summarization, stock movement prediction, credit scoring, firm disclosure) and three multimodal/hallucination tasks (ChartQA, FinVQA, FinTerms). It measures domain-specific reasoning, instruction following, and hallucination mitigation in financial contexts. Use when the user wants to benchmark on FinSet, ChartQA, FinVQA, FinTerms-MCQ, FinTerms-Gen, Finance Bench, or asks about evaluating this task. Reports Task Accuracy/F1.
    3 repo stars
  9. ▌
    Fleisskappa · qhjqhj00
    Compute the FleissKappa metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute FleissKappa, or asks how to score with FleissKappa.
    3 repo stars
  10. ▌
    Fleurs Eval · qhjqhj00
    Evaluates universal speech representations across 102 languages using few-shot learning on parallel speech data. Probes capabilities in automatic speech recognition (ASR), speech language identification, and retrieval tasks. Use when the user wants to benchmark on FLEURS, or asks about evaluating this task. Reports character level error rate.
    3 repo stars
  11. ▌
    Fotbcd Eval · qhjqhj00
    Evaluates cross-dataset generalization and geographic domain shift in building change detection models. It measures how well models trained on one geographic region or dataset perform when tested on entirely different datasets, highlighting the impact of training data diversity on remote sensing model transferability. Use when the user wants to benchmark on FOTBCD-Binary, LEVIR-CD+, WHU-CD, or asks about evaluating this task. Reports IoU.
    3 repo stars
  12. ▌
    Fs Mol Eval · qhjqhj00
    Few-shot molecular property prediction and regression on a large-scale benchmark with thousands of tasks. Probes generalization across diverse protein targets and varying support set sizes. Use when the user wants to benchmark on FS-Mol, or asks about evaluating this task. Reports ΔAUPRC.
    3 repo stars
  13. ▌
    Fysics Eval · qhjqhj00
    Evaluates multimodal large language models' ability to perceive, reason about, and generate physical attributes and laws from images, videos, and audio. It probes causal physical reasoning, material property mapping, and cross-modal consistency rather than superficial pattern matching. Use when the user wants to benchmark on FysicsEval, or asks about evaluating this task. Reports average score.
    3 repo stars
  14. ▌
    Gaoyao Eval · qhjqhj00
    Evaluates the multilingual and multicultural capabilities of large language models across 26 languages and 51 cultures. It probes cognitive abilities (e.g., reasoning, reading, translation) and cultural understanding (monocultural and cross-cultural contexts) to identify geographical performance disparities and benchmark saturation. Use when the user wants to benchmark on GaoYao, or asks about evaluating this task. Reports accuracy, win_rate.
    3 repo stars
  15. ▌
    Gendeg Eval · qhjqhj00
    Evaluates the out-of-distribution (OoD) generalization and within-distribution performance of All-In-One Image Restoration (AIOR) models across six degradation types (haze, rain, snow, motion blur, raindrop, low-light) when trained with synthetic degradation data. Use when the user wants to benchmark on O-HAZE, LHP, RainDS, RSVD, GoPro, or asks about evaluating this task. Reports LPIPS.
    3 repo stars
  16. ▌
    Geneol Eval · qhjqhj00
    Evaluates training-free sentence embedding quality by aggregating LLM-generated semantic variations. Probes semantic similarity preservation and cross-task robustness without model fine-tuning. Use when the user wants to benchmark on STS benchmark, MTEB, or asks about evaluating this task. Reports Spearman rank correlation (cosine similarity).
    3 repo stars
  17. ▌
    Geneval T2i · qhjqhj00
    Evaluates text-to-image generation by measuring how accurately models follow complex prompts with multiple objects, attributes, and spatial constraints. Use when the user wants to benchmark on GenEval, or asks about evaluating this task. Reports GenEval Overall.
    3 repo stars
  18. ▌
    Gnn Ak Eval · qhjqhj00
    Evaluates the expressiveness and practical performance of GNN-AK, a framework that replaces star-shaped neighbor aggregation with subgraph-based encoding in Message Passing Neural Networks. It probes the model's ability to distinguish complex graph structures (e.g., strongly regular graphs, substructures) and predict graph-level properties on standard benchmarks. Use when the user wants to benchmark on ZINC-12K, CIFAR10, PATTERN, MolHIV, MolPCBA, EXP, SR25, or asks about evaluating this task. Reports accuracy (ACC), mean absolute error (MAE).
    3 repo stars
  19. ▌
    Got10k Eval · qhjqhj00
    Evaluates tracking performance on a large-scale dataset of 10,000 videos with diverse object categories, testing generalization to real-world scenarios with language guidance. Use when the user wants to benchmark on GOT-10k, or asks about evaluating this task. Reports AO.
    3 repo stars
  20. ▌
    Gwlans Eval · qhjqhj00
    Predicts target words in computer-aided translation based on source sentences, translation context (prefix, suffix, zero, bidirectional), and human-typed characters. It probes the model's ability to handle discontinuous context and weak positional information in real-world CAT scenarios. Use when the user wants to benchmark on GWLAN Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Hansel Eval · qhjqhj00
    Evaluates Chinese entity linking models on few-shot and zero-shot scenarios, specifically probing their ability to link mentions to tail and emerging Wikidata entities without relying on head entity popularity or dataset-specific fine-tuning. Use when the user wants to benchmark on Hansel, TAC-KBP2015, or asks about evaluating this task. Reports R@1.
    3 repo stars
  22. ▌
    Hasper Eval · qhjqhj00
    Evaluates the ability of computer vision models to classify hand shadow puppet silhouettes into one of 15 distinct categories. It probes feature extraction robustness, particularly for rotationally asymmetric and visually similar silhouettes under varying lighting and motion dynamics. Use when the user wants to benchmark on HaSPeR, or asks about evaluating this task. Reports top-1 accuracy.
    3 repo stars
  23. ▌
    Hazard Eval · qhjqhj00
    Evaluates embodied agents' decision-making and planning capabilities in dynamically changing disaster environments (fire, flood, wind). It probes the ability to reason about evolving object states, environmental propagation dynamics, and spatial-temporal trade-offs to successfully rescue valuable items. Use when the user wants to benchmark on HAZARD, or asks about evaluating this task. Reports rescued value rate (Value).
    3 repo stars
  24. ▌
    Hkmmlu Eval · qhjqhj00
    Evaluates large language models' multilingual comprehension of Hong Kong-specific knowledge, Cantonese linguistic capabilities, and reasoning across STEM, social sciences, and humanities in both Traditional and Simplified Chinese. Use when the user wants to benchmark on HKMMLU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Hm Eqa Eval · qhjqhj00
    This benchmark evaluates a robot's ability to autonomously explore an unseen indoor environment and answer multiple-choice questions requiring object identification, counting, spatial reasoning, and multi-goal navigation. It probes the agent's multimodal perception, iterative reasoning, and navigation efficiency in a dynamic, tool-invoking workflow. Use when the user wants to benchmark on HM-EQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  26. ▌
    I Star Eval · qhjqhj00
    This evaluation probes how anisotropic regularization (I-STAR) affects the downstream performance of fine-tuned language models across standard NLP benchmarks. It also measures the geometric properties of the resulting embedding spaces, specifically isotropy and intrinsic dimensionality, to correlate representation structure with task accuracy. Use when the user wants to benchmark on SST-2, QNLI, RTE, MRPC, QQP, COLA, STS-B, SST-5, SQUAD, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  27. ▌
    Iclerb Eval · qhjqhj00
    Evaluates embedding models and rerankers on their ability to retrieve contextually useful documents for In-Context Learning (ICL) tasks. It measures retrieval effectiveness by ranking candidate documents based on their utility in improving downstream LLM accuracy, rather than relying solely on semantic similarity. Use when the user wants to benchmark on TruthfulQA, Emotion, ProductER, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  28. ▌
    Iclr Points · qhjqhj00
    Quantifies the average research effort required to produce one publication at top-tier conferences across 27 computer science subfields. It enables cross-area comparisons of faculty productivity and publication effort by normalizing faculty headcounts against publication counts. Use when the user has predictions and gold and needs to compute ICLR points.
    3 repo stars
  29. ▌
    Idd Aw Eval · qhjqhj00
    Evaluates the robustness and safety of semantic segmentation models for autonomous driving in unstructured traffic and adverse weather. It specifically probes whether models can correctly identify critical road elements and traffic participants when visual quality degrades due to rain, fog, snow, or low light. Use when the user wants to benchmark on IDD-AW, or asks about evaluating this task. Reports Safe mIoU (SmIoU).
    3 repo stars
  30. ▌
    Ifeval Eval · qhjqhj00
    Evaluates large language models' ability to follow explicit, verifiable instructions embedded in prompts, such as length constraints, keyword inclusion, formatting rules, and language requirements. It measures both strict and loose compliance across individual instructions and entire prompts to assess structural and syntactic robustness. Use when the user wants to benchmark on IFEval, or asks about evaluating this task. Reports Inst-level strict-accuracy.
    3 repo stars
  31. ▌
    Im Iad Eval · qhjqhj00
    Evaluates industrial image anomaly detection algorithms across seven manufacturing datasets under unsupervised, few-shot, continual, and fully supervised settings. It probes both image-level classification and pixel-level localization capabilities, while also measuring computational efficiency like inference speed and GPU memory. Use when the user wants to benchmark on MVTec AD, MVTec LOCO-AD, MPDD, BTAD, MTD, VisA, DAGM, or asks about evaluating this task. Reports Image AUC.
    3 repo stars
  32. ▌
    Imdrug Eval · qhjqhj00
    Evaluates deep learning models for imbalanced and long-tailed classification and regression in AI-aided drug discovery. It probes model robustness to severe class imbalance, open long-tailed distributions, and out-of-distribution chemical splits across graph, sequence, and fingerprint molecular representations. Use when the user wants to benchmark on HIV, SBAP, USPTO-50K, DrugBank, or asks about evaluating this task. Reports Balanced-Acc.
    3 repo stars
  33. ▌
    Inatag Eval · qhjqhj00
    Evaluates multi-class image classification capabilities for agricultural species, genus, family, and crop/weed distinction. Probes fine-grained visual recognition and taxonomic hierarchy understanding in plant identification. Use when the user wants to benchmark on iNatAg, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  34. ▌
    Jaquad Eval · qhjqhj00
    Extractive machine reading comprehension in Japanese. It probes a model's ability to locate exact answer spans in Japanese Wikipedia text given a question, evaluating performance across different answer types, question reasoning types, and answer lengths. Use when the user wants to benchmark on JaQuAD, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  35. ▌
    Jat Rl Eval · qhjqhj00
    Evaluates a multi-modal transformer agent's ability to perform sequential decision-making across diverse reinforcement learning domains, including Atari games, grid-world navigation, and continuous control tasks, without task-specific fine-tuning. Use when the user wants to benchmark on Atari 57, BabyAI, MuJoCo, Meta-World, or asks about evaluating this task. Reports expert normalized score.
    3 repo stars
  36. ▌
    Emmo Eval · qhjqhj00
    This benchmark evaluates embodied mobile manipulation agents in open environments, testing their ability to execute long-horizon, language-conditioned tasks that require interleaved high-level planning and low-level continuous navigation/manipulation. It specifically probes reasoning fidelity, execution success, adaptability to failures, and path efficiency compared to expert trajectories. Use when the user wants to benchmark on EMMOE-100, or asks about evaluating this task. Reports PLWSR.
    3 repo stars
  37. ▌
    Ersb Eval · qhjqhj00
    This benchmark evaluates the environmental resilience of discrete speech codecs by measuring how reconstruction quality and downstream task performance degrade under varying signal-to-noise ratios, loudness levels, and real-world acoustic conditions. It probes both signal fidelity and semantic/intelligibility consistency after codec compression and subsequent speech enhancement or recognition. Use when the user wants to benchmark on Environment-Resilient Speech Codec Benchmark (ERSB), or asks about evaluating this task. Reports PESQ, STOI.
    3 repo stars
  38. ▌
    Fabl Eval · qhjqhj00
    Evaluates a joint learning framework (FABL) for real-time human behavior recognition using 3D skeletal data from depth sensors. It tests the method's ability to simultaneously select discriminative body parts and features for action classification across public benchmarks and a custom robot-interaction task. Use when the user wants to benchmark on MSR Action3D Dataset, Cornell Activity Dataset 60 (CAD-60), Baxter Robot Serving Drinks Task, or asks about evaluating this task. Reports average recognition accuracy.
    3 repo stars
  39. ▌
    Fife Eval · qhjqhj00
    Evaluates large language models' ability to follow complex financial instructions, with a strong emphasis on precise adherence to formatting, structural constraints, and conditional styling requirements. It probes whether models can maintain procedural compliance rather than just semantic correctness. Use when the user wants to benchmark on FIFE, or asks about evaluating this task. Reports Strict compliance.
    3 repo stars
  40. ▌
    Find Eval · qhjqhj00
    Evaluates automated interpretability methods on their ability to describe black-box functions across numeric, string, and semantic domains. It probes both static language model capabilities and interactive agent reasoning, including handling complexities like composition, noise, bias, and approximation. Use when the user wants to benchmark on FIND, or asks about evaluating this task. Reports adequately described rate.
    3 repo stars
  41. ▌
    Flue Eval · qhjqhj00
    Probes French sequence classification capabilities across sentiment analysis, paraphrase identification, and natural language inference. Use when the user wants to benchmark on CLS, PAWSX, XNLI, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  42. ▌
    Forb Eval · qhjqhj00
    Evaluates the quality of universal image embeddings for flat object retrieval across diverse 2D domains (e.g., logos, paintings, currency) under varying visual distortions. It probes both candidate rank accuracy and the matching score margin to assess out-of-distribution generalization. Use when the user wants to benchmark on FORB, or asks about evaluating this task. Reports mAP@5.
    3 repo stars
  43. ▌
    Fowm Eval · qhjqhj00
    Evaluates offline-to-online finetuning of model-based reinforcement learning world models on continuous control and visuomotor tasks. It probes the model's ability to adapt to seen and unseen task variations with limited online interactions while mitigating extrapolation errors via uncertainty regularization. Use when the user wants to benchmark on D4RL, xArm, Quadruped Locomotion, Real xArm, or asks about evaluating this task. Reports Success rate (%).
    3 repo stars
  44. ▌
    Gaia Eval · qhjqhj00
    GAIA probes the ability of AI assistants to perform real-world, conceptually simple tasks that require multi-step reasoning, tool use, and multi-modal processing. It measures robustness in practical everyday reasoning and factual validation on questions explicitly designed to be outside the model's training data. Use when the user wants to benchmark on GAIA, or asks about evaluating this task. Reports score.
    3 repo stars
  45. ▌
    Game Eval · qhjqhj00
    Evaluates a vision-action model's ability to predict actions and scale ratios from video game footage, testing in-distribution and out-of-distribution generalization across 2D and 3D games. Use when the user wants to benchmark on Video Games (In-Distribution & OOD), or asks about evaluating this task. Reports Pearson correlation.
    3 repo stars
  46. ▌
    Gaps Eval · qhjqhj00
    Evaluates AI clinicians on clinical reasoning depth, answer completeness, robustness to input perturbations, and safety risk mitigation. It probes how well models retrieve, synthesize, and apply evidence-based guidelines under varying cognitive loads and adversarial conditions. Use when the user wants to benchmark on GAPS-NCCN-NSCLC-preview, or asks about evaluating this task. Reports GAPS score (normalized rubric-based).
    3 repo stars
  47. ▌
    Glue Eval · qhjqhj00
    Evaluates the transfer learning capability of pre-trained language models across a diverse suite of natural language understanding tasks, including sentence classification, semantic textual similarity, and natural language inference. Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports GLUE Average.
    3 repo stars
  48. ▌
    Gpro Eval · qhjqhj00
    Evaluates large vision-language models on complex mathematical and visual reasoning tasks, measuring both correctness and computational efficiency. It specifically probes the model's ability to avoid excessive chain-of-thought generation (overthinking) by dynamically routing computation between fast perception, slow perception, and slow reasoning paths. Use when the user wants to benchmark on MathVision, MathVerse, MathVista, DynaMath, MM-Vet, or asks about evaluating this task. Reports accuracy (%).
    3 repo stars
  49. ▌
    Gres Eval · qhjqhj00
    Evaluates a model's ability to segment arbitrary numbers of target objects (including zero) in an image based on a natural language expression. It probes multi-target localization, no-target rejection, and robustness to complex linguistic structures like counting and compound relations. Use when the user wants to benchmark on gRefCOCO, or asks about evaluating this task. Reports generalized IoU (gIoU).
    3 repo stars
  50. ▌
    H3wb Eval · qhjqhj00
    Evaluates 3D whole-body human pose estimation and lifting capabilities. It probes a model's ability to reconstruct 133-keypoint 3D skeletons from complete 2D poses, occluded/incomplete 2D poses, or monocular RGB images, with specific focus on body, face, and hand regions. Use when the user wants to benchmark on H3WB, or asks about evaluating this task. Reports MPJPE.
    3 repo stars
  51. ▌
    Hans Eval · qhjqhj00
    Probes whether neural NLI models rely on superficial syntactic heuristics (e.g., lexical overlap, subsequence matching) rather than genuine logical reasoning by presenting structurally similar counterexamples where heuristics lead to incorrect predictions. Use when the user wants to benchmark on HANS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  52. ▌
    Hear Eval · qhjqhj00
    Evaluates the zero-shot generalization and transferability of pre-trained audio representations across 19 diverse downstream tasks spanning speech, environmental sounds, and music. The benchmark requires models to perform without fine-tuning, emphasizing robustness and cross-domain adaptability. Use when the user wants to benchmark on FSD50K, ESC-50, GTZAN, Vocal Imitations, LibriCount, CREMA-D, VoxLingua107, Speech Commands, DCASE 2016 Task 2, Gunshot Triangulation, Beijing Opera, Mridingham Stroke and Tonic, NSynth, Maestro, or asks about evaluating this task. Reports normalized score.
    3 repo stars
  53. ▌
    Helm Eval · qhjqhj00
    Evaluates long-horizon vision-language-action (VLA) manipulation capabilities, specifically testing a model's ability to maintain cross-phase context, predict action failures before execution, and recover from perturbations via rollback or replanning. Use when the user wants to benchmark on LIBERO-LONG, CALVIN ABC→D, LIBERO-Recovery, or asks about evaluating this task. Reports TSR.
    3 repo stars
  54. ▌
    Hingeloss · qhjqhj00
    Compute the HingeLoss metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute HingeLoss, or asks how to score with HingeLoss.
    3 repo stars
  55. ▌
    Holl Eval · qhjqhj00
    Evaluates the security resilience and hardware overhead of Higher-Order Logic Locking (HOLL) against a counterexample-guided inductive synthesis (CEGIS) attack on combinational circuits. It measures how long an attacker takes to recover the secret key relation and the area penalty incurred by the locking mechanism. Use when the user wants to benchmark on ISCAS'85 and MCNC benchmarks, or asks about evaluating this task. Reports attack_time.
    3 repo stars
  56. ▌
    Hsad Eval · qhjqhj00
    Evaluates audio spoof detection models on a newly constructed hybrid spoofing benchmark. It probes robustness against complex, real-world adversarial conditions including mixed-source speech, environmental noise, channel filtering, and compression artifacts. Use when the user wants to benchmark on Hybrid Spoofed Audio Dataset (HSAD), or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  57. ▌
    Hulk Eval · qhjqhj00
    Evaluates the computational efficiency and energy cost of NLP models across pretraining, fine-tuning, and inference phases. It measures the time and monetary cost required to reach predefined performance thresholds on standard NLP tasks, normalized against a BERT-Large baseline. Use when the user wants to benchmark on CoNLL 2003, MNLI, SST-2, or asks about evaluating this task. Reports efficiency score.
    3 repo stars
  58. ▌
    Hume Eval · qhjqhj00
    Evaluates text embedding models against human baselines across 16 MTEB datasets, probing semantic similarity, classification, clustering, and reranking capabilities. It specifically measures cross-lingual performance and identifies task ambiguities where model scores may reflect label pattern reproduction rather than genuine understanding. Use when the user wants to benchmark on MTEB (16 datasets, 26 task-language pairs), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Hupd Eval · qhjqhj00
    Evaluates NLP models on patent-related tasks including binary classification of patent acceptance, multi-class subject area classification using IPC codes, and abstractive summarization of patent claims or descriptions into abstracts. Use when the user wants to benchmark on Harvard USPTO Patent Dataset (HUPD), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  60. ▌
    Hvqr Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform high-order, multistep visual question answering by integrating visual scene graphs with external commonsense knowledge. It explicitly probes the model's reasoning process by requiring it to predict intermediate knowledge triplets alongside the final answer, enforcing explainability and self-diagnosis capabilities. Use when the user wants to benchmark on HVQR, or asks about evaluating this task. Reports triplet recall.
    3 repo stars
  61. ▌
    Idsr Eval · qhjqhj00
    Evaluates the accuracy and diversity of end-to-end sequential recommendation models by testing their ability to predict the next item in a user's behavior sequence while maintaining item diversity in the recommendation list. Use when the user wants to benchmark on ML100K, ML1M, or asks about evaluating this task. Reports Recall.
    3 repo stars
  62. ▌
    Ielm Eval · qhjqhj00
    Evaluates the zero-shot open information extraction (OIE) capability of pre-trained language models by measuring their ability to extract subject-predicate-object triples from text without task-specific training or fine-tuning. It probes whether LMs inherently store rich, open-world relational knowledge that can be accessed via attention mechanisms. Use when the user wants to benchmark on CaRB, Re-OIE2016, TAC KBP-OIE, Wikidata-OIE, or asks about evaluating this task. Reports F1.
    3 repo stars
  63. ▌
    Ifir Eval · qhjqhj00
    This benchmark evaluates an information retrieval system's ability to follow complex, domain-specific instructions when retrieving relevant passages. It probes whether models can interpret nuanced constraints (e.g., patient demographics, legal case details, financial goals) rather than just matching keyword semantics. Use when the user wants to benchmark on IfIR, or asks about evaluating this task. Reports InstFol@20.
    3 repo stars
  64. ▌
    Iirc Eval · qhjqhj00
    Evaluates lifelong learning algorithms on incremental label refinement, requiring models to predict both coarse (superclass) and fine-grained (subclass) labels over time without forgetting prior knowledge, while operating under incomplete information constraints. Use when the user wants to benchmark on IIRC-CIFAR, IIRC-ImageNet, or asks about evaluating this task. Reports pw-JS.
    3 repo stars
  65. ▌
    Imis Eval · qhjqhj00
    Evaluates the ability of vision models to perform interactive medical image segmentation using user prompts like clicks, bounding boxes, or text. It probes how well models generalize across different imaging modalities, anatomical structures, and interaction strategies (single vs. multi-turn). Use when the user wants to benchmark on IMed-361M, TotalSegmentator MRI dataset, ISLES dataset, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  66. ▌
    Ipds Eval · qhjqhj00
    Evaluates large language models' ability to support inpatient clinical decision-making by classifying patient cases into appropriate triage, diagnosis, and treatment pathways. It probes the models' clinical reasoning, diagnostic accuracy, and alignment with real-world physician judgments. Use when the user wants to benchmark on IPDS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  67. ▌
    Ipqa Eval · qhjqhj00
    This benchmark evaluates a model's ability to identify core user intents in personalized question answering. It probes whether systems can infer prioritized motivations from a user's historical Q&A interactions and a target question's narrative, rather than relying on explicit user statements. Use when the user wants to benchmark on IPQA, or asks about evaluating this task. Reports IPQA-Eval F1.
    3 repo stars
  68. ▌
    Irsc Eval · qhjqhj00
    Evaluates embedding models on multilingual information retrieval tasks across five query types (query, title, part-of-paragraph, keyword, summary). It probes semantic comprehension and cross-lingual retrieval alignment in Retrieval-Augmented Generation (RAG) scenarios. Use when the user wants to benchmark on IRSC Benchmark, or asks about evaluating this task. Reports r@10.
    3 repo stars
  69. ▌
    Irt2 Eval · qhjqhj00
    Evaluates neural and baseline models on inductive link prediction and ranking tasks across knowledge graphs of varying scales. It probes the models' ability to map textual entity mentions to graph vertices and rank candidate entities based on combined textual and structural signals, particularly under data scarcity conditions. Use when the user wants to benchmark on IRT2, or asks about evaluating this task. Reports MRR.
    3 repo stars
  70. ▌
    Kacc Eval · qhjqhj00
    Evaluates models' capabilities in knowledge abstraction, concretization, and completion within entity-concept knowledge graphs. It specifically probes multi-hop reasoning through hierarchical relations and cross-view knowledge transfer between entities and concepts. Use when the user wants to benchmark on KACC, or asks about evaluating this task. Reports Hits@10.
    3 repo stars
  71. ▌
    Kale Eval · qhjqhj00
    Evaluates large language models' ability to manipulate and apply stored knowledge across logical reasoning, reading comprehension, and natural language understanding tasks. It specifically probes the 'known & incorrect' phenomenon where models possess relevant facts but fail to apply them correctly during inference. Use when the user wants to benchmark on AbsR, Commonsense (Common), Big Bench Hard (BBH), RACE-H, RACE-M, MMLU, ARC-c, ARC-e, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  72. ▌
    Kgqa Eval · qhjqhj00
    This benchmark evaluates the ability of conversational AI models and traditional knowledge graph question-answering systems to accurately answer natural language questions over structured knowledge graphs. It probes factual grounding, recall on exhaustive lists, robustness to linguistic variations, and determinism across general and academic domains. Use when the user wants to benchmark on QALD-9, YAGO, DBLP, MAG, or asks about evaluating this task. Reports Micro F1 score.
    3 repo stars
  73. ▌
    Kilt Eval · qhjqhj00
    Evaluates a model's ability to perform knowledge-intensive language tasks by jointly assessing output generation accuracy and evidence retrieval from a fixed Wikipedia snapshot. It measures how well models can produce correct answers while providing verifiable text-span provenance to justify predictions. Use when the user wants to benchmark on KILT, or asks about evaluating this task. Reports KILT scores.
    3 repo stars
  74. ▌
    Kstt Eval · qhjqhj00
    Evaluates session-based recommendation models on predicting the next item in a user's click sequence by integrating knowledge graph attributes and temporal dynamics between clicks. Use when the user wants to benchmark on Yoochoose, Diginetica, Last-fm, or asks about evaluating this task. Reports Recall@20.
    3 repo stars
  75. ▌
    Lasq Eval · qhjqhj00
    Probes the ability to extract aspect-based sentiment quadruples (target, aspect, opinion, sentiment) from text in low-resource agglutinative languages. It evaluates exact-match performance across entity detection, relation linking, and full quadruple composition. Use when the user wants to benchmark on LASQ, or asks about evaluating this task. Reports F1.
    3 repo stars
  76. ▌
    Liar Eval · qhjqhj00
    This evaluation probes a model's ability to perform binary fact-checking on short political claims by mapping multi-class truthfulness labels to positive/negative categories. It specifically tests how well the system handles compositional reasoning and uncertainty, requiring it to output definitive verdicts or abstain. Use when the user wants to benchmark on LIAR, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  77. ▌
    Load Time · qhjqhj00
    Evaluates the data loading speed and rendering interactivity of the encube visual analytics framework on a tiled display system under varying data volumes and GPU memory constraints. Use when the user has predictions and gold and needs to compute Load time ($T_{\mathrm{Load}}$).
    3 repo stars
  78. ▌
    Lovr Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant long-form videos or fine-grained clips based on rich, narrative-driven text queries. It probes temporal reasoning, semantic alignment across extended durations, and robustness to long-context inputs and varying frame sampling strategies. Use when the user wants to benchmark on LoVR, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  79. ▌
    Lrrg Eval · qhjqhj00
    Evaluates a model's ability to generate accurate radiology reports from single-view X-ray images under varying quality conditions, from standard to severely degraded. It probes robustness to clinical acquisition artifacts and tests whether the model can extract quality-invariant diagnostic features without relying on historical patient data. Use when the user wants to benchmark on MIMIC-CXR LRRG Benchmarks, or asks about evaluating this task. Reports CheXbert F1.
    3 repo stars
  80. ▌
    Ltdr Eval · qhjqhj00
    Evaluates the vision-language understanding and domain generalization capabilities of Mixture-of-Experts (MoE) models. It tests how well a long-tailed distribution-aware router preserves routing variance for vision tokens while maintaining load balancing for language tokens, impacting both accuracy and inference efficiency. Use when the user wants to benchmark on GQA, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, MM-Vet, PACS, VLCS, Office-Home, DomainNet, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Luss Eval · qhjqhj00
    Evaluates the ability of models to perform pixel-level semantic segmentation on large-scale, diverse image collections without human annotations. It probes unsupervised representation learning, category discovery, and fine-grained mask prediction capabilities. Use when the user wants to benchmark on ImageNet-S, ImageNet-S50, ImageNet-S300, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  82. ▌
    M3av Eval · qhjqhj00
    Evaluates multimodal academic lecture understanding across speech recognition, speech synthesis, and slide/script generation. It probes models' ability to handle complex academic language, rare words, multimodal alignment, and knowledge comprehension. Use when the user wants to benchmark on M3AV, or asks about evaluating this task. Reports BWER, ROUGE-1/2/L.
    3 repo stars
  83. ▌
    M3da Eval · qhjqhj00
    This benchmark evaluates unsupervised domain adaptation (UDA) methods for 3D medical image segmentation across eight practical domain shifts, including inter-modality changes (MRI-CT), scanner parameters, contrast presence, and radiation dose. It measures how well models trained on a source domain can segment target domain volumes without target labels, highlighting the robustness of adaptation techniques to real-world imaging variability. Use when the user wants to benchmark on AMOS, LIDC, BraTS, CC359, or asks about evaluating this task. Reports multiclass Dice score.
    3 repo stars
  84. ▌
    M3it Eval · qhjqhj00
    Evaluates a vision-language model's ability to follow multi-modal instructions, answer knowledge-based visual questions, and generalize to unseen languages and video tasks. It probes cross-modal alignment, cross-lingual transfer, and the model's conversational response quality. Use when the user wants to benchmark on M^3IT, OK-VQA, A-OKVQA, ViQuAE, Flickr-8k-CN, FM-IQA, Chinese-FoodNet, MSRVTT, iVQA, ActivityNet-QA, MSRVTT-QA, MSVD-QA, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  85. ▌
    Made Eval · qhjqhj00
    Evaluates end-to-end autonomous materials discovery pipelines by measuring how effectively different policies (planners, generators, selectors, and agentic orchestrators) can find thermodynamically stable compounds under constrained oracle query budgets. It probes the trade-offs between discovery efficiency, structural diversity, and adaptivity as chemical complexity and stability thresholds increase. Use when the user wants to benchmark on MADE Benchmark Environments, or asks about evaluating this task. Reports AF (Acceleration Factor).
    3 repo stars
  86. ▌
    Maeb Eval · qhjqhj00
    Evaluates audio embedding models across 30 tasks spanning speech, music, environmental sounds, bioacoustics, emotion recognition, and cross-modal audio-text reasoning in over 100 languages. It probes the ability of models to generalize across acoustic domains, handle multilingual alignment, and perform both supervised and unsupervised audio understanding tasks. Use when the user wants to benchmark on MAEB, or asks about evaluating this task. Reports Average Score.
    3 repo stars
  87. ▌
    Mair Eval · qhjqhj00
    Evaluates retrieval models' ability to follow complex, task-specific instructions across diverse domains and long-tail tasks. It measures how instruction tuning impacts generalization and performance on heterogeneous query-document relevance tasks compared to non-instruction-tuned baselines. Use when the user wants to benchmark on MAIR, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  88. ▌
    Matq Eval · qhjqhj00
    Evaluates the accuracy of a charge-equilibrated equivariant foundation potential in predicting atomic energies, forces, partial charges, and bulk mechanical/thermal properties across diverse crystalline, molecular, and ionic systems. Use when the user wants to benchmark on MatQ, Custom Model Systems (C10H2/C10H3+, Ag3+/−, Na8/9Cl8+, Au2-MgO(001)), or asks about evaluating this task. Reports RMSE, MAE.
    3 repo stars
  89. ▌
    Maud Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform legal reading comprehension on merger agreements by answering specialized deal point questions. It probes the model's capacity to interpret complex contractual clauses and handle imbalanced classification tasks across various legal categories. Use when the user wants to benchmark on MAUD, or asks about evaluating this task. Reports AUPR.
    3 repo stars
  90. ▌
    Max Error · qhjqhj00
    Compute the max_error metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute max_error, or asks how to score with max_error.
    3 repo stars
  91. ▌
    Maxm Eval · qhjqhj00
    This benchmark evaluates multilingual visual question answering (mVQA) by testing models on images paired with questions in seven languages. It probes a model's ability to perform cross-lingual visual reasoning and generate accurate text answers without relying on costly human annotation. Use when the user wants to benchmark on MaXM, or asks about evaluating this task. Reports Exact Match Accuracy.
    3 repo stars
  92. ▌
    Maxmetric · qhjqhj00
    Compute the MaxMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MaxMetric, or asks how to score with MaxMetric.
    3 repo stars
  93. ▌
    Mbib Eval · qhjqhj00
    This benchmark evaluates a model's ability to identify various forms of media bias, including linguistic, cognitive, political, racial, gender, and hate speech bias, across diverse text sources like news articles, tweets, and social media comments. It probes whether models can generalize across different bias types and dataset sizes without being skewed by larger datasets. Use when the user wants to benchmark on MBIB, or asks about evaluating this task. Reports macro F1-score.
    3 repo stars
  94. ▌
    Mbpp Eval · qhjqhj00
    Evaluates a model's ability to generate correct, self-contained Python functions from natural language problem descriptions. It probes basic programming logic, standard library usage, and semantic grounding of simple algorithmic tasks. Use when the user wants to benchmark on Mostly Basic Programming Problems (MBPP), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  95. ▌
    Mcif Eval · qhjqhj00
    Evaluates multimodal models' ability to follow crosslingual instructions on scientific talks, testing speech recognition, translation, question answering, and summarization across short and long contexts in English, German, Italian, and Chinese. Use when the user wants to benchmark on MCIF, or asks about evaluating this task. Reports BERTScore.
    3 repo stars
  96. ▌
    Mcwq Eval · qhjqhj00
    Evaluates a model's ability to generalize compositionally to unseen syntactic structures and cross-lingual settings in semantic parsing. It measures how well models translate natural language questions into correct SPARQL queries across monolingual and zero-shot cross-lingual scenarios. Use when the user wants to benchmark on MCWQ, or asks about evaluating this task. Reports Exact Match (%).
    3 repo stars
  97. ▌
    Mdbi Mtbi · qhjqhj00
    Evaluates the robustness and safety of autonomous driving systems by quantifying how far and how long the vehicle operates between human interventions (disengagements). It enables unbiased comparison across different AV platforms and road environments by normalizing disengagement frequency with spatial and temporal data. Use when the user has predictions and gold and needs to compute MDBI, MTBI.
    3 repo stars
  98. ▌
    Mdec Eval · qhjqhj00
    Evaluates self-supervised monocular depth estimation models by measuring image-based accuracy and 3D pointcloud reconstruction quality. It probes the models' ability to generalize across diverse environments (urban, natural, agricultural, indoor) and highlights the impact of scale ambiguity and oversmoothing on relative object positioning. Use when the user wants to benchmark on SYNS-Patches, or asks about evaluating this task. Reports F-Score (Edges).
    3 repo stars
  99. ▌
    Mdia Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate coherent, contextually appropriate, and lexically diverse dialogue responses across 46 languages. It specifically probes cross-lingual transfer capabilities and measures the performance gap between high-resource and low-resource languages in open-domain conversation. Use when the user wants to benchmark on MDIA, or asks about evaluating this task. Reports sacreBLEU.
    3 repo stars
  100. ▌
    Mdpe Eval · qhjqhj00
    Evaluates multimodal deception detection across video, audio, and text modalities, while also probing how individual differences—specifically personality traits and emotional expressivity—influence deceptive behavior and detection accuracy. Use when the user wants to benchmark on MDPE, or asks about evaluating this task. Reports Accuracy.
    3 repo stars