all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 70 of 76

  1. ▌
    Mbgvc Eval · qhjqhj00
    Evaluates a model's ability to generate entity-aware, fine-grained text descriptions of basketball videos, specifically requiring accurate prediction of player names and precise action recognition. Use when the user wants to benchmark on MbgVC, or asks about evaluating this task. Reports GDS.
    3 repo stars
  2. ▌
    Meanmetric · qhjqhj00
    Compute the MeanMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanMetric, or asks how to score with MeanMetric.
    3 repo stars
  3. ▌
    Mecat Eval · qhjqhj00
    This benchmark evaluates fine-grained audio understanding by testing models on generating detailed, multi-perspective captions and answering probing questions across diverse acoustic domains. It specifically probes a model's ability to distinguish between speech, music, and sound events, reason about acoustic scenes, and assess technical audio quality without relying on generic descriptions. Use when the user wants to benchmark on MECAT, or asks about evaluating this task. Reports DATE.
    3 repo stars
  4. ▌
    Medec Eval · qhjqhj00
    Evaluates large language models' ability to detect and correct medical errors in clinical text. It probes the model's sensitivity to clinical inaccuracies, its precision in localizing erroneous sentences, and its capacity to generate semantically and lexically accurate corrections using different prompting strategies. Use when the user wants to benchmark on MEDEC, or asks about evaluating this task. Reports AggScore.
    3 repo stars
  5. ▌
    Midog Eval · qhjqhj00
    Evaluates deep learning models for detecting mitotic figures in histopathology whole-slide images, specifically probing their ability to generalize across different scanner-induced domain shifts such as color distribution, contrast, and depth-of-field variations. Use when the user wants to benchmark on MIDOG, or asks about evaluating this task. Reports F_1 score.
    3 repo stars
  6. ▌
    Mlgym Eval · qhjqhj00
    Evaluates LLM agents on open-ended AI research tasks across 13 diverse benchmarks. It measures the agent's ability to navigate codebases, run experiments, and improve model performance, assessing capabilities from reproducing existing research to achieving state-of-the-art results. Use when the user wants to benchmark on MLGym Benchmarks, or asks about evaluating this task. Reports AutoML-inspired optimization metric.
    3 repo stars
  7. ▌
    Mlsum Eval · qhjqhj00
    Evaluates abstractive text summarization models across multiple languages (French, German, Spanish, Russian, Turkish) to measure generation quality and investigate cross-lingual performance gaps and model biases. Use when the user wants to benchmark on MLSUM, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  8. ▌
    Mm Iq Eval · qhjqhj00
    Probes human-like abstraction and visual reasoning capabilities in multimodal models across eight fine-grained paradigms, including logical operations, geometry, and spatial relationships. It measures how well models generalize to novel abstract patterns without relying on memorized visual features. Use when the user wants to benchmark on MM-IQ, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  9. ▌
    Mmmeb Eval · qhjqhj00
    Evaluates cross-lingual and cross-modal embedding alignment for image-text retrieval, classification, visual question answering, and visual grounding. Probes whether multilingual adaptation preserves semantic consistency across languages and modalities without degrading English performance. Use when the user wants to benchmark on MMMEB, or asks about evaluating this task. Reports P@1.
    3 repo stars
  10. ▌
    Mmteb Eval · qhjqhj00
    Evaluates the quality of multilingual text embeddings across diverse tasks and languages. It probes capabilities like semantic similarity, classification, retrieval, and multilingual alignment. Use when the user wants to benchmark on MTEB(Multilingual), MTEB(Europe), MTEB(Indic), or asks about evaluating this task. Reports Borda count.
    3 repo stars
  11. ▌
    Modad Eval · qhjqhj00
    Evaluates a model's ability to mitigate spurious correlations (bias) in image classification by measuring performance on both overall test sets and specifically on bias-conflicting samples where the spurious attribute contradicts the true label. Use when the user wants to benchmark on Corrupted CIFAR-10, BAR, BFFHQ, Waterbirds, or asks about evaluating this task. Reports Average Accuracy.
    3 repo stars
  12. ▌
    Mokb6 Eval · qhjqhj00
    This benchmark evaluates multilingual knowledge graph embedding models on the task of completing missing facts across six languages. It specifically probes the model's ability to leverage cross-lingual information flow, benefit from translated training triples, and retain facts when queried in different scripts. Use when the user wants to benchmark on mOKB6, or asks about evaluating this task. Reports H@10.
    3 repo stars
  13. ▌
    Mot16 Eval · qhjqhj00
    Evaluates multi-object tracking algorithms on video sequences by measuring detection accuracy, identity consistency, and localization precision. It assesses how well trackers maintain object identities over time while correctly handling occlusions, distractors, and varying crowd densities. Use when the user wants to benchmark on MOT16, or asks about evaluating this task. Reports MOTA.
    3 repo stars
  14. ▌
    Motif Eval · qhjqhj00
    This benchmark evaluates the accuracy of malware family classification models and antivirus-based labeling tools on a large, expert-verified dataset. It probes a model's ability to correctly assign ground-truth family labels to malware samples, including handling open-set noise and alias resolution. Use when the user wants to benchmark on MOTIF, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  15. ▌
    Moverscore · qhjqhj00
    Automated evaluation of text generation quality by computing semantic distance between system outputs and human references using contextualized embeddings and Earth Mover's Distance (EMD). It probes a model's ability to capture meaning-based similarity rather than surface-level n-gram overlaps across machine translation, summarization, dialogue, and image captioning tasks. Use when the user has predictions and gold and needs to compute Pearson r.
    3 repo stars
  16. ▌
    Mtvqa Eval · qhjqhj00
    This benchmark evaluates the multilingual visual-textual alignment and comprehension capabilities of multimodal large language models (MLLMs). It specifically probes whether models can accurately perceive, extract, and reason about text embedded within images across nine different languages without relying on translation. Use when the user wants to benchmark on MTVQA, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  17. ▌
    Mucue Eval · qhjqhj00
    Evaluates a model's ability to understand music across a spectrum of tasks, ranging from low-level acoustic perception (e.g., pitch, chord, rhythm) to high-level cognitive reasoning (e.g., genre, mood, structure, lyrical comprehension). It probes whether foundation models can process long-context audio and lyrics jointly to answer standardized multiple-choice questions. Use when the user wants to benchmark on MuCUE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  18. ▌
    Musdb Eval · qhjqhj00
    Evaluates a model's ability to isolate individual musical stems (vocals, drums, bass, other) from mixed audio recordings, testing long-range context modeling and cross-domain attention capabilities in source separation. Use when the user wants to benchmark on MUSDB, or asks about evaluating this task. Reports SDR.
    3 repo stars
  19. ▌
    Mvtec Eval · qhjqhj00
    Unsupervised anomaly detection and pixel-level localization on industrial defect data. It probes the model's ability to distinguish normal from defective samples and precisely segment defect regions without using labeled anomalies during training. Use when the user wants to benchmark on MVTec, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  20. ▌
    Nacsp Eval · qhjqhj00
    Evaluates a model's ability to predict discrete neural audio codec parameters (quantizers, sampling rate, bits per second) from audio samples, enabling fine-grained source attribution of AI-generated speech. The protocol frames open-set attribution as a multi-task regression problem rather than binary classification, requiring the model to generalize across both seen and unseen codec configurations. Use when the user wants to benchmark on ST-Codecfake, CodecFake, or asks about evaluating this task. Reports MSE.
    3 repo stars
  21. ▌
    Ndcg Score · qhjqhj00
    Compute the ndcg_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute ndcg_score, or asks how to score with ndcg_score.
    3 repo stars
  22. ▌
    Nerel Eval · qhjqhj00
    Evaluates models on nested named entity recognition and relation extraction in Russian. It probes the ability to identify overlapping/contained entities and classify semantic relations between them, including cross-sentence and nested relations. Use when the user wants to benchmark on NEREL, or asks about evaluating this task. Reports F1.
    3 repo stars
  23. ▌
    Nerrf Eval · qhjqhj00
    Evaluates the ability to reconstruct 3D geometry and synthesize novel views of transparent and specular objects from monocular RGB images and silhouettes. It probes physically accurate light path simulation, including refraction, reflection, and Fresnel effects, using a differentiable rendering framework. Use when the user wants to benchmark on Blender Synthetic Dataset, or asks about evaluating this task. Reports Chamfer Distance (CD).
    3 repo stars
  24. ▌
    Nl2sh Eval · qhjqhj00
    This benchmark evaluates the ability of large language models to translate natural language instructions into executable Bash commands. It probes functional correctness by comparing model-generated commands against ground-truth commands using a functional equivalence heuristic that combines command execution with LLM-based output analysis. Use when the user wants to benchmark on NL2SH, InterCode-ALFA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Nlb21 Eval · qhjqhj00
    Evaluates latent variable models for their ability to infer neural population dynamics and rates from spiking data. It probes model fidelity in predicting held-out spiking activity, decoding behavioral variables, and capturing autonomous forward dynamics without relying on external labels. Use when the user wants to benchmark on MC_Maze, MC_Maze-L, MC_Maze-M, MC_Maze-S, MC_RTT, Area2_Bump, DMFC_RSG, or asks about evaluating this task. Reports Co-smoothing bps.
    3 repo stars
  26. ▌
    Norne Eval · qhjqhj00
    Evaluates Named Entity Recognition (NER) performance on Norwegian text, testing the model's ability to identify and classify entity boundaries and types (PER, ORG, LOC, GPE, PROD, EVT, DRV) across Bokmål and Nynorsk variants. Use when the user wants to benchmark on NorNE, or asks about evaluating this task. Reports F1 (strict).
    3 repo stars
  27. ▌
    Nubia Eval · qhjqhj00
    Evaluates a learned neural metric's ability to correlate with human judgments of text generation quality. It probes semantic similarity, logical inference, and sentence likelihood capabilities across machine translation and image captioning domains. Use when the user wants to benchmark on WMT (Machine Translation), Flickr 8K, or asks about evaluating this task. Reports Pearson correlation.
    3 repo stars
  28. ▌
    Nubot Eval · qhjqhj00
    Evaluates the ability of unbalanced optimal transport models to predict distributional shifts, mass creation (proliferation), and mass destruction (cell death) in heterogeneous populations under drug perturbation. Use when the user wants to benchmark on Synthetic Gaussian Mixture, Single-Cell Perturbation Response (Melanoma), or asks about evaluating this task. Reports weighted kernel MMD.
    3 repo stars
  29. ▌
    Nusax Eval · qhjqhj00
    Evaluates sentiment classification and machine translation capabilities across 10 low-resource Indonesian local languages, Indonesian, and English. It probes cross-lingual transferability, multilingual training benefits, and data efficiency for underrepresented Austronesian languages. Use when the user wants to benchmark on NusaX, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  30. ▌
    Nvtts Eval · qhjqhj00
    This evaluation protocol assesses the capability of zero-shot text-to-speech models to synthesize nonverbal vocalizations (NVs) like breathing, laughter, coughing, and sighs alongside emotional speech. It measures speech intelligibility, speaker and emotion fidelity, acoustic quality, and the precise alignment of generated NVs with reference audio. Use when the user wants to benchmark on NVTTS, or asks about evaluating this task. Reports WER.
    3 repo stars
  31. ▌
    Occam Eval · qhjqhj00
    Evaluates whether object-centric representations derived from zero-shot segmentation masks enable robust zero-shot classification under spurious background correlations, and compares them against slot-based OCL methods on unsupervised object discovery. Use when the user wants to benchmark on Movi-C, Movi-E, UrbanCars, ImageNet-D, ImageNet-9, Waterbirds, CounterAnimals, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  32. ▌
    Oodis Eval · qhjqhj00
    Evaluates a model's ability to detect and segment anomalous objects (out-of-distribution instances) in real-world driving scenes. It measures instance-level segmentation and object detection performance on rare, unpredictable obstacles that are not part of the standard in-distribution classes. Use when the user wants to benchmark on OoDIS Benchmark, or asks about evaluating this task. Reports AP.
    3 repo stars
  33. ▌
    Orbit Eval · qhjqhj00
    Evaluates recommendation models on candidate item ranking across multiple public sequential recommendation datasets and a large-scale synthetic hidden test (ClueWeb-Reco) to assess generalization to unseen item pools and real-world browsing scenarios. Use when the user wants to benchmark on ML-1M, Amazon Beauty, Amazon Toys, Amazon Sports, Amazon Books, ClueWeb-Reco, or asks about evaluating this task. Reports Recall@10, NDCG@10.
    3 repo stars
  34. ▌
    Osbad Eval · qhjqhj00
    Evaluates the ability of statistical and machine learning models to detect anomalies in battery discharge capacity profiles across different chemistries. It probes cross-chemistry generalization and model robustness on imbalanced, rare-anomaly datasets typical of electrochemical systems. Use when the user wants to benchmark on MIT/Stanford (Severson), Tohoku, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  35. ▌
    Otb99 Eval · qhjqhj00
    Evaluates short-term vision-language tracking performance on a curated subset of OTB100 with added textual annotations, testing robustness to appearance changes and scale variations. Use when the user wants to benchmark on OTB99, or asks about evaluating this task. Reports PR.
    3 repo stars
  36. ▌
    Ov Vg Eval · qhjqhj00
    Evaluates a model's ability to localize objects in images based on natural language descriptions without prior exposure to those specific categories. It probes visual-linguistic alignment, handling of novel vocabulary, and robustness to varying object scales and complex scenes. Use when the user wants to benchmark on OV-VG, or asks about evaluating this task. Reports Acc50.
    3 repo stars
  37. ▌
    Ozone Eval · qhjqhj00
    Evaluates a unified platform for standardizing heterogeneous transportation trajectory data, automating cross-dataset conversion, and benchmarking safety and behavior models across multiple cities and datasets. Use when the user wants to benchmark on Ozone Standardized Trajectory Suite (NGSIM, highD, CitySim, UTE), or asks about evaluating this task. Reports cross-city F1 score.
    3 repo stars
  38. ▌
    Pando Eval · qhjqhj00
    Evaluates whether mechanistic interpretability methods can recover decision-relevant signals from black-box models that lack faithful explanations. It probes the ability of gradient-based, representation-based, and black-box elicitation agents to predict held-out outcomes and identify correct decision-rule fields across varying explanation qualities and model complexities. Use when the user wants to benchmark on Pando, or asks about evaluating this task. Reports Held-out accuracy (%).
    3 repo stars
  39. ▌
    Pdfqa Eval · qhjqhj00
    Evaluates end-to-end question answering over PDF documents, probing parsing, retrieval, and reasoning capabilities across diverse document types, modalities, and complexity dimensions. Use when the user wants to benchmark on pdfQA, or asks about evaluating this task. Reports G-Eval correctness.
    3 repo stars
  40. ▌
    Pearl Eval · qhjqhj00
    Evaluates large vision-language models' ability to understand and generate culturally-aware Arabic content across multiple reasoning-centric question types. It probes hypothesis formation, comparative analysis, chronological reasoning, and explicit cultural grounding in both closed-form and open-ended multimodal tasks. Use when the user wants to benchmark on PeARL, or asks about evaluating this task. Reports relaxed-match accuracy (ACC).
    3 repo stars
  41. ▌
    Perplexity · qhjqhj00
    This protocol evaluates language model memorisation and training data contamination by measuring how well the model predicts benchmark text compared to out-of-distribution baselines. Lower perplexity on benchmark passages relative to a clean baseline indicates the model has likely seen the text during training. Use when the user has predictions and gold and needs to compute perplexity.
    3 repo stars
  42. ▌
    Phomt Eval · qhjqhj00
    This benchmark evaluates Vietnamese-English machine translation quality by comparing neural baselines and commercial engines. It probes translation accuracy across multiple domains and sentence lengths using both automatic metrics and human preference judgments. Use when the user wants to benchmark on PhoMT, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  43. ▌
    Phuma Eval · qhjqhj00
    Evaluates a humanoid robot's ability to imitate human motion and follow pelvis trajectories using physically-grounded retargeting. It probes full-body tracking accuracy and partial-state path-following control across diverse locomotion categories on Unitree G1 and H1-2 robots. Use when the user wants to benchmark on PHUMA, Unseen Video, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  44. ▌
    Phyre Eval · qhjqhj00
    Evaluates an agent's ability to reason about 2D Newtonian physics to solve goal-driven puzzles by placing dynamic objects. It probes sample-efficient learning and generalization across unseen task templates and action spaces. Use when the user wants to benchmark on PHYRE, or asks about evaluating this task. Reports AUCCESSION.
    3 repo stars
  45. ▌
    Polqa Eval · qhjqhj00
    Evaluates open-domain question answering in Polish by measuring both passage retrieval accuracy and answer generation quality. It probes a model's ability to retrieve relevant evidence from a large corpus and accurately extract or generate answers from those passages. Use when the user wants to benchmark on PolQA, or asks about evaluating this task. Reports fuzzy_match.
    3 repo stars
  46. ▌
    Powrl Eval · qhjqhj00
    Evaluates reinforcement learning agents for real-time power grid topology control under adversarial attacks and dynamic loads. It probes the agent's ability to maintain grid stability, minimize operational costs, and avoid blackouts across multiple challenging scenarios. Use when the user wants to benchmark on L2RPN NeurIPS 2020 (Robustness track) Offline, L2RPN NeurIPS 2020 (Robustness track) Online, L2RPN WCCI 2020 Offline, or asks about evaluating this task. Reports survival steps, scenario score.
    3 repo stars
  47. ▌
    Prism Eval · qhjqhj00
    Evaluates fine-grained, multi-aspect-aware paper-to-paper retrieval by decomposing long-form query papers into aspect-specific views and segmenting candidate papers into section-level representations for targeted retrieval. Use when the user wants to benchmark on SciFullBench, PatentFullBench, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  48. ▌
    Qa4ie Eval · qhjqhj00
    Evaluates document-level information extraction by framing it as a question answering task. It probes a model's ability to extract cross-sentence relation triples from large documents using entity-relation queries and knowledge base alignment. Use when the user wants to benchmark on QA4IE, or asks about evaluating this task. Reports Exact Match (EM), F1-score.
    3 repo stars
  49. ▌
    Qe4pe Eval · qhjqhj00
    Probes the practical usability and impact of word-level quality estimation highlights on professional translators' post-editing efficiency, accuracy, and workflow. It measures how different highlight modalities (oracle, supervised, unsupervised, none) affect editing effort, productivity, and final translation quality in real-world domain-specific settings. Use when the user wants to benchmark on QE4PE, or asks about evaluating this task. Reports ESA score.
    3 repo stars
  50. ▌
    Qilin Eval · qhjqhj00
    Evaluates multimodal information retrieval systems across search, recommendation, and deep query answering (DQA) tasks using real-world APP-level user sessions. It probes a model's ability to rank heterogeneous content (text, images, videos) and generate accurate answers augmented by retrieved documents. Use when the user wants to benchmark on Qilin, or asks about evaluating this task. Reports MRR@10.
    3 repo stars
  51. ▌
    Qmsum Eval · qhjqhj00
    Evaluates an agent's ability to distill and synthesize key information from high-noise, multi-turn meeting transcripts based on specific user queries. It probes the model's capacity to maintain global context while filtering irrelevant dialogue turns to produce a query-relevant summary. Use when the user wants to benchmark on QMSUM, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  52. ▌
    Qpain Eval · qhjqhj00
    Measures social bias in medical question-answering systems for pain management by evaluating treatment denial rates across intersectional race-gender profiles. It probes whether AI models exhibit discriminatory prescribing patterns when presented with clinical vignettes containing demographic attributes. Use when the user wants to benchmark on Q-Pain, or asks about evaluating this task. Reports probability_of_no.
    3 repo stars
  53. ▌
    Qwen2 Eval · qhjqhj00
    This protocol evaluates large language models across core competencies including general knowledge, reasoning, coding, and mathematics. It also assesses multilingual understanding, instruction following, and alignment with human preferences. Use when the user wants to benchmark on MMLU, MMLU-Pro, GPQA, HumanEval, GSM8K, MT-Bench, IFEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  54. ▌
    R Ice Eval · qhjqhj00
    Evaluates the accuracy of a regression framework (R-ICE) in estimating prompt-level inference carbon and energy emissions for LLMs using only token counts and publicly available performance data. It probes whether runtime can be reliably modeled as a piecewise linear function of input and output tokens without intrusive monitoring or architecture details. Use when the user wants to benchmark on HELM, or asks about evaluating this task. Reports average prediction error.
    3 repo stars
  55. ▌
    R2med Eval · qhjqhj00
    Evaluates retrieval models on reasoning-driven medical tasks where document relevance is determined by alignment with inferred clinical diagnoses or multi-step reasoning paths rather than lexical or semantic overlap. Covers three task types—Q&A reference, clinical evidence, and clinical case retrieval—spanning eight medical sub-domains. Use when the user wants to benchmark on R2MED, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  56. ▌
    Rand Score · qhjqhj00
    Compute the rand_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute rand_score, or asks how to score with rand_score.
    3 repo stars
  57. ▌
    Rar B Eval · qhjqhj00
    Evaluates whether dense retrievers and re-rankers can semantically encode and retrieve correct answers to reasoning problems across diverse tasks. It probes the retriever-LLM behavioral gap by testing performance with and without task instructions, and compares full-dataset retrieval against multiple-choice retrieval settings. Use when the user wants to benchmark on RAR-b, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  58. ▌
    Ravel Eval · qhjqhj00
    Evaluates interpretability methods' ability to disentangle polysemantic language model representations by isolating causal attributes through activation interventions on residual stream features. Use when the user wants to benchmark on RAVEL, or asks about evaluating this task. Reports Disentanglescore.
    3 repo stars
  59. ▌
    Re Mi Eval · qhjqhj00
    Evaluates multimodal LLMs' ability to reason across multiple images, including sequential and set-based consumption, interleaved text-image processing, and heterogeneous visual inputs like charts, equations, maps, and code. It probes cross-image contextual integration, precise visual reading, and step-by-step logical deduction. Use when the user wants to benchmark on ReMI, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  60. ▌
    Regen Eval · qhjqhj00
    Evaluates conversational recommender systems on next-item prediction and joint narrative generation, specifically testing how well models incorporate user interaction history and explicit natural language critiques to produce accurate recommendations and contextually grounded textual explanations. Use when the user wants to benchmark on REGEN, or asks about evaluating this task. Reports Recall@10.
    3 repo stars
  61. ▌
    Resyn Eval · qhjqhj00
    Evaluates the reasoning capabilities of language models on synthetically generated, code-verifiable tasks. It measures performance on both the custom ReSyn dataset and standard reasoning benchmarks using zero-shot generation with specific sampling parameters. Use when the user wants to benchmark on ReSyn, or asks about evaluating this task. Reports mean@4.
    3 repo stars
  62. ▌
    Rfuav Eval · qhjqhj00
    Evaluates deep learning models' ability to identify specific UAV models from radio-frequency signals by classifying time-frequency spectrograms. It probes robustness to varying signal-to-noise ratios (SNR) and sensitivity to preprocessing choices like color maps and frequency resolution. Use when the user wants to benchmark on RFUAV, or asks about evaluating this task. Reports Acc.
    3 repo stars
  63. ▌
    Rospr Eval · qhjqhj00
    Evaluates the zero-shot generalization capability of instruction-tuned language models by retrieving and applying task-specific soft prompt embeddings at inference time to adapt to unseen tasks. Use when the user wants to benchmark on BIG-bench, SuperGLUE/HellaSwag/StoryCloze/WiC suite, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  64. ▌
    Rougescore · qhjqhj00
    Compute the ROUGEScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ROUGEScore, or asks how to score with ROUGEScore.
    3 repo stars
  65. ▌
    Rover Eval · qhjqhj00
    Evaluates the accuracy and robustness of visual-inertial SLAM systems across diverse outdoor environments, seasons, and lighting conditions. It probes long-term trajectory consistency, scale estimation, and environmental adaptability under challenging visual degradation. Use when the user wants to benchmark on ROVER, or asks about evaluating this task. Reports mATE.
    3 repo stars
  66. ▌
    Rp 1k Eval · qhjqhj00
    Evaluates a model's ability to discover recurring visual patterns in a single image. It measures detection accuracy at both the individual pattern instance level and the whole pattern level against human annotations. Use when the user wants to benchmark on RP-1K, or asks about evaluating this task. Reports RP Instance Recall.
    3 repo stars
  67. ▌
    Rsmeb Eval · qhjqhj00
    Evaluates a vision-language model's ability to perform zero-shot classification, cross-modal retrieval, visual question answering, and fine-grained spatial grounding (including region-caption retrieval and geo-localization) on remote sensing imagery. It measures how well instruction-conditioned contrastive pretraining aligns multimodal features with geospatial metadata and textual prompts. Use when the user wants to benchmark on AID, Million-AID, RSI-CB, EuroSAT, UCM, PatternNet, RSITMD, RSICD, UCM-caption, LRBEN, HRBEN, or asks about evaluating this task. Reports Friedman score.
    3 repo stars
  68. ▌
    Rsrcc Eval · qhjqhj00
    This benchmark evaluates large language models' ability to perform fine-grained, region-specific semantic reasoning on remote sensing image pairs. It probes localized change comprehension by asking models to answer binary, multiple-choice, and open-ended questions about specific changes (e.g., new construction, vegetation loss) within satellite imagery. Use when the user wants to benchmark on RSRCC, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  69. ▌
    Ruler Eval · qhjqhj00
    This benchmark evaluates long-context language models' ability to retrieve, trace, aggregate, and answer questions across varying context lengths and task complexities. It probes whether models genuinely attend to injected information or rely on parametric knowledge and context copying as sequence length increases. Use when the user wants to benchmark on RULER, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  70. ▌
    Runningsum · qhjqhj00
    Compute the RunningSum metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RunningSum, or asks how to score with RunningSum.
    3 repo stars
  71. ▌
    Rxmrl Eval · qhjqhj00
    Evaluates a model's ability to maintain context and generate coherent responses in multi-turn dialogue, comparing stateful event-driven architectures against standard decoder-only LLMs. Use when the user wants to benchmark on MRL Curriculum Datasets (derived from TinyStories), or asks about evaluating this task. Reports MRL Reward Score.
    3 repo stars
  72. ▌
    Sciex Eval · qhjqhj00
    Evaluates large language models on solving university-level scientific exams in computer science. It probes capabilities in open-ended reasoning, mathematical proof writing, long-form explanations, and multimodal (image-text) understanding across English and German languages. Use when the user wants to benchmark on SciEx, or asks about evaluating this task. Reports Normalized score (0-100%).
    3 repo stars
  73. ▌
    Selqa Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant answer sentences from a document given a question (selection), and to determine whether a document section contains an answer at all (triggering). It probes open-domain QA robustness against paraphrasing, varying question types, and section lengths. Use when the user wants to benchmark on SelQA, or asks about evaluating this task. Reports MAP.
    3 repo stars
  74. ▌
    Sfiog Eval · qhjqhj00
    Evaluates large language models' ability to generate coherent, logically structured financial investment opinions based on company context and questions. It probes reasoning over mere knowledge retrieval by testing performance across varying degrees of familiarity and novelty in companies and questions. Use when the user wants to benchmark on sFIOG, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  75. ▌
    Shale Eval · qhjqhj00
    This benchmark evaluates fine-grained hallucination in Large Vision-Language Models (LVLMs) by testing their faithfulness to visual inputs and factuality against external knowledge. It measures model performance under clean conditions and across hierarchical input perturbations (image, instruction, and combination levels) to assess hallucination resistance. Use when the user wants to benchmark on SHALE, or asks about evaluating this task. Reports accuracy, non-hallucination rate.
    3 repo stars
  76. ▌
    Share Eval · qhjqhj00
    Evaluates a model's ability to predict the next item in an anonymous user session based on sequential click history. It probes the model's capacity to capture short-term user intent and higher-order item correlations within dynamic session contexts. Use when the user wants to benchmark on YooChoose, Diginetica, or asks about evaluating this task. Reports Hit@20.
    3 repo stars
  77. ▌
    Sim3d Eval · qhjqhj00
    Evaluates 3D anomaly detection and segmentation capabilities in industrial settings using multiview and multimodal (image + depth) inputs. It probes a model's ability to identify and localize defects across multiple object categories under both in-domain (real-to-real) and out-of-domain (synthetic-to-real) conditions. Use when the user wants to benchmark on SiM3D, or asks about evaluating this task. Reports I-AUROC.
    3 repo stars
  78. ▌
    Slump Eval · qhjqhj00
    Measures how much final code fidelity degrades when a system's design is progressively disclosed through multi-turn interaction rather than provided upfront. It evaluates semantic faithfulness to a committed design and structural integration of dependencies in long-horizon coding agents. Use when the user wants to benchmark on SLUMP benchmark, or asks about evaluating this task. Reports IF50.
    3 repo stars
  79. ▌
    Sparc Eval · qhjqhj00
    Probes cross-domain semantic parsing in context by requiring models to generate sequential SQL queries across multiple conversational turns. It evaluates the ability to maintain state, handle thematic evolution, and generalize to unseen databases while correctly resolving contextual dependencies. Use when the user wants to benchmark on SParC, or asks about evaluating this task. Reports question match.
    3 repo stars
  80. ▌
    Spert Eval · qhjqhj00
    Evaluates a model's ability to jointly identify named entity spans with their types and extract relational tuples between them from unstructured text. It probes span-based representation learning, localized context modeling, and joint classification without relying on sequential tagging schemes like BIO. Use when the user wants to benchmark on CoNLL04, SciERC, ADE, or asks about evaluating this task. Reports F1 score (micro/macro-averaged).
    3 repo stars
  81. ▌
    Spiqa Eval · qhjqhj00
    Evaluates multimodal long-context reasoning and figure/table comprehension on scientific papers. Tests direct question answering with images, full paper context, and chain-of-thought retrieval capabilities. Use when the user wants to benchmark on SPIQA, or asks about evaluating this task. Reports L3Score.
    3 repo stars
  82. ▌
    Squad Eval · qhjqhj00
    Measures a model's ability to extract precise answer spans from a given context paragraph in response to a natural language question, testing reading comprehension and span prediction. Use when the user wants to benchmark on SQuAD 1.1/2.0, or asks about evaluating this task. Reports F1.
    3 repo stars
  83. ▌
    Stark Eval · qhjqhj00
    Evaluates retrieval models on their ability to find relevant entities in semi-structured knowledge bases using complex queries that combine textual descriptions and relational constraints. It probes joint reasoning over mixed textual-relational semantics and user-intent modeling across product, academic, and medical domains. Use when the user wants to benchmark on STaRK, or asks about evaluating this task. Reports Hit@k.
    3 repo stars
  84. ▌
    Statscores · qhjqhj00
    Compute the StatScores metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute StatScores, or asks how to score with StatScores.
    3 repo stars
  85. ▌
    Svamp Eval · qhjqhj00
    Evaluates whether NLP models can genuinely solve simple math word problems through arithmetic reasoning versus relying on shallow heuristics like bag-of-words matching or positional cues. It probes model brittleness by testing performance on standard datasets alongside carefully perturbed variants that remove questions or alter operator types. Use when the user wants to benchmark on MAWPS, ASDiv-A, SVAMP, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  86. ▌
    Tcav Score · qhjqhj00
    Measures the influence of human-defined emotional concepts (physiognomy, utterance polarity, voice pitch) on a multimodal emotion recognition model's decisions using Concept Activation Vectors. It quantifies how much each concept drives the model's classification decisions across different network layers. Use when the user has predictions and gold and needs to compute TCAV score.
    3 repo stars
  87. ▌
    Tid 8 Eval · qhjqhj00
    Probes a model's ability to learn from inherently subjective or disagreed-upon annotations by treating each annotator's label as a separate example, rather than aggregating them into a single ground truth label. Use when the user wants to benchmark on TID-8, or asks about evaluating this task. Reports exact match accuracy.
    3 repo stars
  88. ▌
    Timer Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform temporal reasoning and extract accurate information from longitudinal electronic health records (EHRs). It probes whether models can correctly synthesize evidence across multiple time-stamped clinical visits, adhere to specified temporal boundaries, and maintain accuracy over long patient timelines. Use when the user wants to benchmark on TIMER-Bench, MedAlign, or asks about evaluating this task. Reports Correct.
    3 repo stars
  89. ▌
    Tnl2k Eval · qhjqhj00
    Evaluates natural language-based tracking on 2000 YouTube and surveillance videos, testing the model's ability to follow and adapt to language descriptions over time. Use when the user wants to benchmark on TNL2K, or asks about evaluating this task. Reports AUC.
    3 repo stars
  90. ▌
    Tnllt Eval · qhjqhj00
    Evaluates long-term vision-language tracking capability by measuring localization accuracy over extended video sequences while dynamically updating natural language descriptions to handle appearance changes and occlusions. Use when the user wants to benchmark on TNLLT, or asks about evaluating this task. Reports PR.
    3 repo stars
  91. ▌
    Trace Eval · qhjqhj00
    This protocol evaluates training-free partial audio deepfake detection by analyzing the temporal continuity of frozen speech foundation model embeddings. It probes a model's ability to detect splice boundaries and synthetic insertions in speech without requiring labeled training data or architectural modifications. Use when the user wants to benchmark on PartialSpoof, HalfTruth Audio Deepfake (HAD), ADD 2023 Track 2, LlamaPartialSpoof, or asks about evaluating this task. Reports EER.
    3 repo stars
  92. ▌
    Tsaia Eval · qhjqhj00
    Evaluates LLMs' ability to perform multi-step, constraint-aware reasoning and inference on real-world time series data. It probes compositional reasoning, numerical precision, and the capacity to assemble complex analytical or forecasting workflows via executable code generation. Use when the user wants to benchmark on TSAIA, or asks about evaluating this task. Reports Success Rate.
    3 repo stars
  93. ▌
    Tsaqa Eval · qhjqhj00
    Evaluates large language models' ability to perform time series analysis and reasoning across six tasks (anomaly detection, classification, characterization, comparison, data transformation, and temporal relationship) using three question formats (true-or-false, multiple-choice, and puzzling). Use when the user wants to benchmark on TSAQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  94. ▌
    Tsbow Eval · qhjqhj00
    Evaluates object detection models on traffic surveillance footage under diverse weather conditions and varying degrees of vehicle occlusion. It probes robustness to environmental degradation, scale variation, and dense urban traffic scenarios. Use when the user wants to benchmark on TSBOW, or asks about evaluating this task. Reports mAP50.
    3 repo stars
  95. ▌
    Tsrec Eval · qhjqhj00
    Evaluates a model's ability to perform sequential recommendation with a focus on capturing repeat-aware temporal patterns. It measures how well the model balances predicting new items versus recurring items based on user interaction history and time intervals. Use when the user wants to benchmark on RetailRocket, LastFM, Diginetica, or asks about evaluating this task. Reports HR@K.
    3 repo stars
  96. ▌
    Tsver Eval · qhjqhj00
    This benchmark evaluates an AI system's ability to perform fact verification using time-series evidence. It probes multi-timeframe temporal reasoning, cross-series numerical analysis, and the generation of factually consistent justifications aligned with human annotations. Use when the user wants to benchmark on TSVer, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  97. ▌
    Tt Df Eval · qhjqhj00
    Evaluates the ability of models to detect human body forgeries generated by diffusion models. It probes spatiotemporal motion inconsistencies and generalization across different generation configurations and unseen manipulation models. Use when the user wants to benchmark on TT-DF, or asks about evaluating this task. Reports AUC.
    3 repo stars
  98. ▌
    Ugr16 Eval · qhjqhj00
    Evaluates how data preprocessing choices—such as observation selection, flow directionality, and feature engineering—affect the performance of unsupervised anomaly detection models on network traffic. Use when the user wants to benchmark on UGR'16, or asks about evaluating this task. Reports AUC.
    3 repo stars
  99. ▌
    Unite Eval · qhjqhj00
    Evaluates text-to-SQL models on compositional generalization, out-of-domain robustness, and schema-question alignment across 18 diverse datasets and 12 domains. It probes the model's ability to handle long-form query decomposition and cross-domain SQL pattern diversity. Use when the user wants to benchmark on UNITE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Unpie Eval · qhjqhj00
    Assesses multimodal language models' ability to resolve lexical ambiguity in puns using visual context. It probes visual-textual alignment, multimodal literacy, and the capacity to disambiguate or reconstruct ambiguous text when provided with explanatory or disambiguating images. Use when the user wants to benchmark on UNPIE, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars