qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Mbgvc Eval · qhjqhj00Evaluates a model's ability to generate entity-aware, fine-grained text descriptions of basketball videos, specifically requiring accurate prediction of player names and precise action recognition. Use when the user wants to benchmark on MbgVC, or asks about evaluating this task. Reports GDS.
- ▌ Meanmetric · qhjqhj00Compute the MeanMetric metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MeanMetric, or asks how to score with MeanMetric.
- ▌ Mecat Eval · qhjqhj00This benchmark evaluates fine-grained audio understanding by testing models on generating detailed, multi-perspective captions and answering probing questions across diverse acoustic domains. It specifically probes a model's ability to distinguish between speech, music, and sound events, reason about acoustic scenes, and assess technical audio quality without relying on generic descriptions. Use when the user wants to benchmark on MECAT, or asks about evaluating this task. Reports DATE.
- ▌ Medec Eval · qhjqhj00Evaluates large language models' ability to detect and correct medical errors in clinical text. It probes the model's sensitivity to clinical inaccuracies, its precision in localizing erroneous sentences, and its capacity to generate semantically and lexically accurate corrections using different prompting strategies. Use when the user wants to benchmark on MEDEC, or asks about evaluating this task. Reports AggScore.
- ▌ Midog Eval · qhjqhj00Evaluates deep learning models for detecting mitotic figures in histopathology whole-slide images, specifically probing their ability to generalize across different scanner-induced domain shifts such as color distribution, contrast, and depth-of-field variations. Use when the user wants to benchmark on MIDOG, or asks about evaluating this task. Reports F_1 score.
- ▌ Mlgym Eval · qhjqhj00Evaluates LLM agents on open-ended AI research tasks across 13 diverse benchmarks. It measures the agent's ability to navigate codebases, run experiments, and improve model performance, assessing capabilities from reproducing existing research to achieving state-of-the-art results. Use when the user wants to benchmark on MLGym Benchmarks, or asks about evaluating this task. Reports AutoML-inspired optimization metric.
- ▌ Mlsum Eval · qhjqhj00Evaluates abstractive text summarization models across multiple languages (French, German, Spanish, Russian, Turkish) to measure generation quality and investigate cross-lingual performance gaps and model biases. Use when the user wants to benchmark on MLSUM, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Mm Iq Eval · qhjqhj00Probes human-like abstraction and visual reasoning capabilities in multimodal models across eight fine-grained paradigms, including logical operations, geometry, and spatial relationships. It measures how well models generalize to novel abstract patterns without relying on memorized visual features. Use when the user wants to benchmark on MM-IQ, or asks about evaluating this task. Reports accuracy.
- ▌ Mmmeb Eval · qhjqhj00Evaluates cross-lingual and cross-modal embedding alignment for image-text retrieval, classification, visual question answering, and visual grounding. Probes whether multilingual adaptation preserves semantic consistency across languages and modalities without degrading English performance. Use when the user wants to benchmark on MMMEB, or asks about evaluating this task. Reports P@1.
- ▌ Mmteb Eval · qhjqhj00Evaluates the quality of multilingual text embeddings across diverse tasks and languages. It probes capabilities like semantic similarity, classification, retrieval, and multilingual alignment. Use when the user wants to benchmark on MTEB(Multilingual), MTEB(Europe), MTEB(Indic), or asks about evaluating this task. Reports Borda count.
- ▌ Modad Eval · qhjqhj00Evaluates a model's ability to mitigate spurious correlations (bias) in image classification by measuring performance on both overall test sets and specifically on bias-conflicting samples where the spurious attribute contradicts the true label. Use when the user wants to benchmark on Corrupted CIFAR-10, BAR, BFFHQ, Waterbirds, or asks about evaluating this task. Reports Average Accuracy.
- ▌ Mokb6 Eval · qhjqhj00This benchmark evaluates multilingual knowledge graph embedding models on the task of completing missing facts across six languages. It specifically probes the model's ability to leverage cross-lingual information flow, benefit from translated training triples, and retain facts when queried in different scripts. Use when the user wants to benchmark on mOKB6, or asks about evaluating this task. Reports H@10.
- ▌ Mot16 Eval · qhjqhj00Evaluates multi-object tracking algorithms on video sequences by measuring detection accuracy, identity consistency, and localization precision. It assesses how well trackers maintain object identities over time while correctly handling occlusions, distractors, and varying crowd densities. Use when the user wants to benchmark on MOT16, or asks about evaluating this task. Reports MOTA.
- ▌ Motif Eval · qhjqhj00This benchmark evaluates the accuracy of malware family classification models and antivirus-based labeling tools on a large, expert-verified dataset. It probes a model's ability to correctly assign ground-truth family labels to malware samples, including handling open-set noise and alias resolution. Use when the user wants to benchmark on MOTIF, or asks about evaluating this task. Reports accuracy.
- ▌ Moverscore · qhjqhj00Automated evaluation of text generation quality by computing semantic distance between system outputs and human references using contextualized embeddings and Earth Mover's Distance (EMD). It probes a model's ability to capture meaning-based similarity rather than surface-level n-gram overlaps across machine translation, summarization, dialogue, and image captioning tasks. Use when the user has predictions and gold and needs to compute Pearson r.
- ▌ Mtvqa Eval · qhjqhj00This benchmark evaluates the multilingual visual-textual alignment and comprehension capabilities of multimodal large language models (MLLMs). It specifically probes whether models can accurately perceive, extract, and reason about text embedded within images across nine different languages without relying on translation. Use when the user wants to benchmark on MTVQA, or asks about evaluating this task. Reports Accuracy.
- ▌ Mucue Eval · qhjqhj00Evaluates a model's ability to understand music across a spectrum of tasks, ranging from low-level acoustic perception (e.g., pitch, chord, rhythm) to high-level cognitive reasoning (e.g., genre, mood, structure, lyrical comprehension). It probes whether foundation models can process long-context audio and lyrics jointly to answer standardized multiple-choice questions. Use when the user wants to benchmark on MuCUE, or asks about evaluating this task. Reports accuracy.
- ▌ Musdb Eval · qhjqhj00Evaluates a model's ability to isolate individual musical stems (vocals, drums, bass, other) from mixed audio recordings, testing long-range context modeling and cross-domain attention capabilities in source separation. Use when the user wants to benchmark on MUSDB, or asks about evaluating this task. Reports SDR.
- ▌ Mvtec Eval · qhjqhj00Unsupervised anomaly detection and pixel-level localization on industrial defect data. It probes the model's ability to distinguish normal from defective samples and precisely segment defect regions without using labeled anomalies during training. Use when the user wants to benchmark on MVTec, or asks about evaluating this task. Reports AUROC.
- ▌ Nacsp Eval · qhjqhj00Evaluates a model's ability to predict discrete neural audio codec parameters (quantizers, sampling rate, bits per second) from audio samples, enabling fine-grained source attribution of AI-generated speech. The protocol frames open-set attribution as a multi-task regression problem rather than binary classification, requiring the model to generalize across both seen and unseen codec configurations. Use when the user wants to benchmark on ST-Codecfake, CodecFake, or asks about evaluating this task. Reports MSE.
- ▌ Ndcg Score · qhjqhj00Compute the ndcg_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute ndcg_score, or asks how to score with ndcg_score.
- ▌ Nerel Eval · qhjqhj00Evaluates models on nested named entity recognition and relation extraction in Russian. It probes the ability to identify overlapping/contained entities and classify semantic relations between them, including cross-sentence and nested relations. Use when the user wants to benchmark on NEREL, or asks about evaluating this task. Reports F1.
- ▌ Nerrf Eval · qhjqhj00Evaluates the ability to reconstruct 3D geometry and synthesize novel views of transparent and specular objects from monocular RGB images and silhouettes. It probes physically accurate light path simulation, including refraction, reflection, and Fresnel effects, using a differentiable rendering framework. Use when the user wants to benchmark on Blender Synthetic Dataset, or asks about evaluating this task. Reports Chamfer Distance (CD).
- ▌ Nl2sh Eval · qhjqhj00This benchmark evaluates the ability of large language models to translate natural language instructions into executable Bash commands. It probes functional correctness by comparing model-generated commands against ground-truth commands using a functional equivalence heuristic that combines command execution with LLM-based output analysis. Use when the user wants to benchmark on NL2SH, InterCode-ALFA, or asks about evaluating this task. Reports accuracy.
- ▌ Nlb21 Eval · qhjqhj00Evaluates latent variable models for their ability to infer neural population dynamics and rates from spiking data. It probes model fidelity in predicting held-out spiking activity, decoding behavioral variables, and capturing autonomous forward dynamics without relying on external labels. Use when the user wants to benchmark on MC_Maze, MC_Maze-L, MC_Maze-M, MC_Maze-S, MC_RTT, Area2_Bump, DMFC_RSG, or asks about evaluating this task. Reports Co-smoothing bps.
- ▌ Norne Eval · qhjqhj00Evaluates Named Entity Recognition (NER) performance on Norwegian text, testing the model's ability to identify and classify entity boundaries and types (PER, ORG, LOC, GPE, PROD, EVT, DRV) across Bokmål and Nynorsk variants. Use when the user wants to benchmark on NorNE, or asks about evaluating this task. Reports F1 (strict).
- ▌ Nubia Eval · qhjqhj00Evaluates a learned neural metric's ability to correlate with human judgments of text generation quality. It probes semantic similarity, logical inference, and sentence likelihood capabilities across machine translation and image captioning domains. Use when the user wants to benchmark on WMT (Machine Translation), Flickr 8K, or asks about evaluating this task. Reports Pearson correlation.
- ▌ Nubot Eval · qhjqhj00Evaluates the ability of unbalanced optimal transport models to predict distributional shifts, mass creation (proliferation), and mass destruction (cell death) in heterogeneous populations under drug perturbation. Use when the user wants to benchmark on Synthetic Gaussian Mixture, Single-Cell Perturbation Response (Melanoma), or asks about evaluating this task. Reports weighted kernel MMD.
- ▌ Nusax Eval · qhjqhj00Evaluates sentiment classification and machine translation capabilities across 10 low-resource Indonesian local languages, Indonesian, and English. It probes cross-lingual transferability, multilingual training benefits, and data efficiency for underrepresented Austronesian languages. Use when the user wants to benchmark on NusaX, or asks about evaluating this task. Reports macro-F1.
- ▌ Nvtts Eval · qhjqhj00This evaluation protocol assesses the capability of zero-shot text-to-speech models to synthesize nonverbal vocalizations (NVs) like breathing, laughter, coughing, and sighs alongside emotional speech. It measures speech intelligibility, speaker and emotion fidelity, acoustic quality, and the precise alignment of generated NVs with reference audio. Use when the user wants to benchmark on NVTTS, or asks about evaluating this task. Reports WER.
- ▌ Occam Eval · qhjqhj00Evaluates whether object-centric representations derived from zero-shot segmentation masks enable robust zero-shot classification under spurious background correlations, and compares them against slot-based OCL methods on unsupervised object discovery. Use when the user wants to benchmark on Movi-C, Movi-E, UrbanCars, ImageNet-D, ImageNet-9, Waterbirds, CounterAnimals, or asks about evaluating this task. Reports accuracy.
- ▌ Oodis Eval · qhjqhj00Evaluates a model's ability to detect and segment anomalous objects (out-of-distribution instances) in real-world driving scenes. It measures instance-level segmentation and object detection performance on rare, unpredictable obstacles that are not part of the standard in-distribution classes. Use when the user wants to benchmark on OoDIS Benchmark, or asks about evaluating this task. Reports AP.
- ▌ Orbit Eval · qhjqhj00Evaluates recommendation models on candidate item ranking across multiple public sequential recommendation datasets and a large-scale synthetic hidden test (ClueWeb-Reco) to assess generalization to unseen item pools and real-world browsing scenarios. Use when the user wants to benchmark on ML-1M, Amazon Beauty, Amazon Toys, Amazon Sports, Amazon Books, ClueWeb-Reco, or asks about evaluating this task. Reports Recall@10, NDCG@10.
- ▌ Osbad Eval · qhjqhj00Evaluates the ability of statistical and machine learning models to detect anomalies in battery discharge capacity profiles across different chemistries. It probes cross-chemistry generalization and model robustness on imbalanced, rare-anomaly datasets typical of electrochemical systems. Use when the user wants to benchmark on MIT/Stanford (Severson), Tohoku, or asks about evaluating this task. Reports AUROC.
- ▌ Otb99 Eval · qhjqhj00Evaluates short-term vision-language tracking performance on a curated subset of OTB100 with added textual annotations, testing robustness to appearance changes and scale variations. Use when the user wants to benchmark on OTB99, or asks about evaluating this task. Reports PR.
- ▌ Ov Vg Eval · qhjqhj00Evaluates a model's ability to localize objects in images based on natural language descriptions without prior exposure to those specific categories. It probes visual-linguistic alignment, handling of novel vocabulary, and robustness to varying object scales and complex scenes. Use when the user wants to benchmark on OV-VG, or asks about evaluating this task. Reports Acc50.
- ▌ Ozone Eval · qhjqhj00Evaluates a unified platform for standardizing heterogeneous transportation trajectory data, automating cross-dataset conversion, and benchmarking safety and behavior models across multiple cities and datasets. Use when the user wants to benchmark on Ozone Standardized Trajectory Suite (NGSIM, highD, CitySim, UTE), or asks about evaluating this task. Reports cross-city F1 score.
- ▌ Pando Eval · qhjqhj00Evaluates whether mechanistic interpretability methods can recover decision-relevant signals from black-box models that lack faithful explanations. It probes the ability of gradient-based, representation-based, and black-box elicitation agents to predict held-out outcomes and identify correct decision-rule fields across varying explanation qualities and model complexities. Use when the user wants to benchmark on Pando, or asks about evaluating this task. Reports Held-out accuracy (%).
- ▌ Pdfqa Eval · qhjqhj00Evaluates end-to-end question answering over PDF documents, probing parsing, retrieval, and reasoning capabilities across diverse document types, modalities, and complexity dimensions. Use when the user wants to benchmark on pdfQA, or asks about evaluating this task. Reports G-Eval correctness.
- ▌ Pearl Eval · qhjqhj00Evaluates large vision-language models' ability to understand and generate culturally-aware Arabic content across multiple reasoning-centric question types. It probes hypothesis formation, comparative analysis, chronological reasoning, and explicit cultural grounding in both closed-form and open-ended multimodal tasks. Use when the user wants to benchmark on PeARL, or asks about evaluating this task. Reports relaxed-match accuracy (ACC).
- ▌ Perplexity · qhjqhj00This protocol evaluates language model memorisation and training data contamination by measuring how well the model predicts benchmark text compared to out-of-distribution baselines. Lower perplexity on benchmark passages relative to a clean baseline indicates the model has likely seen the text during training. Use when the user has predictions and gold and needs to compute perplexity.
- ▌ Phomt Eval · qhjqhj00This benchmark evaluates Vietnamese-English machine translation quality by comparing neural baselines and commercial engines. It probes translation accuracy across multiple domains and sentence lengths using both automatic metrics and human preference judgments. Use when the user wants to benchmark on PhoMT, or asks about evaluating this task. Reports BLEU.
- ▌ Phuma Eval · qhjqhj00Evaluates a humanoid robot's ability to imitate human motion and follow pelvis trajectories using physically-grounded retargeting. It probes full-body tracking accuracy and partial-state path-following control across diverse locomotion categories on Unitree G1 and H1-2 robots. Use when the user wants to benchmark on PHUMA, Unseen Video, or asks about evaluating this task. Reports success_rate.
- ▌ Phyre Eval · qhjqhj00Evaluates an agent's ability to reason about 2D Newtonian physics to solve goal-driven puzzles by placing dynamic objects. It probes sample-efficient learning and generalization across unseen task templates and action spaces. Use when the user wants to benchmark on PHYRE, or asks about evaluating this task. Reports AUCCESSION.
- ▌ Polqa Eval · qhjqhj00Evaluates open-domain question answering in Polish by measuring both passage retrieval accuracy and answer generation quality. It probes a model's ability to retrieve relevant evidence from a large corpus and accurately extract or generate answers from those passages. Use when the user wants to benchmark on PolQA, or asks about evaluating this task. Reports fuzzy_match.
- ▌ Powrl Eval · qhjqhj00Evaluates reinforcement learning agents for real-time power grid topology control under adversarial attacks and dynamic loads. It probes the agent's ability to maintain grid stability, minimize operational costs, and avoid blackouts across multiple challenging scenarios. Use when the user wants to benchmark on L2RPN NeurIPS 2020 (Robustness track) Offline, L2RPN NeurIPS 2020 (Robustness track) Online, L2RPN WCCI 2020 Offline, or asks about evaluating this task. Reports survival steps, scenario score.
- ▌ Prism Eval · qhjqhj00Evaluates fine-grained, multi-aspect-aware paper-to-paper retrieval by decomposing long-form query papers into aspect-specific views and segmenting candidate papers into section-level representations for targeted retrieval. Use when the user wants to benchmark on SciFullBench, PatentFullBench, or asks about evaluating this task. Reports Recall@K.
- ▌ Qa4ie Eval · qhjqhj00Evaluates document-level information extraction by framing it as a question answering task. It probes a model's ability to extract cross-sentence relation triples from large documents using entity-relation queries and knowledge base alignment. Use when the user wants to benchmark on QA4IE, or asks about evaluating this task. Reports Exact Match (EM), F1-score.
- ▌ Qe4pe Eval · qhjqhj00Probes the practical usability and impact of word-level quality estimation highlights on professional translators' post-editing efficiency, accuracy, and workflow. It measures how different highlight modalities (oracle, supervised, unsupervised, none) affect editing effort, productivity, and final translation quality in real-world domain-specific settings. Use when the user wants to benchmark on QE4PE, or asks about evaluating this task. Reports ESA score.
- ▌ Qilin Eval · qhjqhj00Evaluates multimodal information retrieval systems across search, recommendation, and deep query answering (DQA) tasks using real-world APP-level user sessions. It probes a model's ability to rank heterogeneous content (text, images, videos) and generate accurate answers augmented by retrieved documents. Use when the user wants to benchmark on Qilin, or asks about evaluating this task. Reports MRR@10.
- ▌ Qmsum Eval · qhjqhj00Evaluates an agent's ability to distill and synthesize key information from high-noise, multi-turn meeting transcripts based on specific user queries. It probes the model's capacity to maintain global context while filtering irrelevant dialogue turns to produce a query-relevant summary. Use when the user wants to benchmark on QMSUM, or asks about evaluating this task. Reports ROUGE-1.
- ▌ Qpain Eval · qhjqhj00Measures social bias in medical question-answering systems for pain management by evaluating treatment denial rates across intersectional race-gender profiles. It probes whether AI models exhibit discriminatory prescribing patterns when presented with clinical vignettes containing demographic attributes. Use when the user wants to benchmark on Q-Pain, or asks about evaluating this task. Reports probability_of_no.
- ▌ Qwen2 Eval · qhjqhj00This protocol evaluates large language models across core competencies including general knowledge, reasoning, coding, and mathematics. It also assesses multilingual understanding, instruction following, and alignment with human preferences. Use when the user wants to benchmark on MMLU, MMLU-Pro, GPQA, HumanEval, GSM8K, MT-Bench, IFEval, or asks about evaluating this task. Reports accuracy.
- ▌ R Ice Eval · qhjqhj00Evaluates the accuracy of a regression framework (R-ICE) in estimating prompt-level inference carbon and energy emissions for LLMs using only token counts and publicly available performance data. It probes whether runtime can be reliably modeled as a piecewise linear function of input and output tokens without intrusive monitoring or architecture details. Use when the user wants to benchmark on HELM, or asks about evaluating this task. Reports average prediction error.
- ▌ R2med Eval · qhjqhj00Evaluates retrieval models on reasoning-driven medical tasks where document relevance is determined by alignment with inferred clinical diagnoses or multi-step reasoning paths rather than lexical or semantic overlap. Covers three task types—Q&A reference, clinical evidence, and clinical case retrieval—spanning eight medical sub-domains. Use when the user wants to benchmark on R2MED, or asks about evaluating this task. Reports nDCG@10.
- ▌ Rand Score · qhjqhj00Compute the rand_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute rand_score, or asks how to score with rand_score.
- ▌ Rar B Eval · qhjqhj00Evaluates whether dense retrievers and re-rankers can semantically encode and retrieve correct answers to reasoning problems across diverse tasks. It probes the retriever-LLM behavioral gap by testing performance with and without task instructions, and compares full-dataset retrieval against multiple-choice retrieval settings. Use when the user wants to benchmark on RAR-b, or asks about evaluating this task. Reports nDCG@10.
- ▌ Ravel Eval · qhjqhj00Evaluates interpretability methods' ability to disentangle polysemantic language model representations by isolating causal attributes through activation interventions on residual stream features. Use when the user wants to benchmark on RAVEL, or asks about evaluating this task. Reports Disentanglescore.
- ▌ Re Mi Eval · qhjqhj00Evaluates multimodal LLMs' ability to reason across multiple images, including sequential and set-based consumption, interleaved text-image processing, and heterogeneous visual inputs like charts, equations, maps, and code. It probes cross-image contextual integration, precise visual reading, and step-by-step logical deduction. Use when the user wants to benchmark on ReMI, or asks about evaluating this task. Reports accuracy.
- ▌ Regen Eval · qhjqhj00Evaluates conversational recommender systems on next-item prediction and joint narrative generation, specifically testing how well models incorporate user interaction history and explicit natural language critiques to produce accurate recommendations and contextually grounded textual explanations. Use when the user wants to benchmark on REGEN, or asks about evaluating this task. Reports Recall@10.
- ▌ Resyn Eval · qhjqhj00Evaluates the reasoning capabilities of language models on synthetically generated, code-verifiable tasks. It measures performance on both the custom ReSyn dataset and standard reasoning benchmarks using zero-shot generation with specific sampling parameters. Use when the user wants to benchmark on ReSyn, or asks about evaluating this task. Reports mean@4.
- ▌ Rfuav Eval · qhjqhj00Evaluates deep learning models' ability to identify specific UAV models from radio-frequency signals by classifying time-frequency spectrograms. It probes robustness to varying signal-to-noise ratios (SNR) and sensitivity to preprocessing choices like color maps and frequency resolution. Use when the user wants to benchmark on RFUAV, or asks about evaluating this task. Reports Acc.
- ▌ Rospr Eval · qhjqhj00Evaluates the zero-shot generalization capability of instruction-tuned language models by retrieving and applying task-specific soft prompt embeddings at inference time to adapt to unseen tasks. Use when the user wants to benchmark on BIG-bench, SuperGLUE/HellaSwag/StoryCloze/WiC suite, or asks about evaluating this task. Reports accuracy.
- ▌ Rougescore · qhjqhj00Compute the ROUGEScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ROUGEScore, or asks how to score with ROUGEScore.
- ▌ Rover Eval · qhjqhj00Evaluates the accuracy and robustness of visual-inertial SLAM systems across diverse outdoor environments, seasons, and lighting conditions. It probes long-term trajectory consistency, scale estimation, and environmental adaptability under challenging visual degradation. Use when the user wants to benchmark on ROVER, or asks about evaluating this task. Reports mATE.
- ▌ Rp 1k Eval · qhjqhj00Evaluates a model's ability to discover recurring visual patterns in a single image. It measures detection accuracy at both the individual pattern instance level and the whole pattern level against human annotations. Use when the user wants to benchmark on RP-1K, or asks about evaluating this task. Reports RP Instance Recall.
- ▌ Rsmeb Eval · qhjqhj00Evaluates a vision-language model's ability to perform zero-shot classification, cross-modal retrieval, visual question answering, and fine-grained spatial grounding (including region-caption retrieval and geo-localization) on remote sensing imagery. It measures how well instruction-conditioned contrastive pretraining aligns multimodal features with geospatial metadata and textual prompts. Use when the user wants to benchmark on AID, Million-AID, RSI-CB, EuroSAT, UCM, PatternNet, RSITMD, RSICD, UCM-caption, LRBEN, HRBEN, or asks about evaluating this task. Reports Friedman score.
- ▌ Rsrcc Eval · qhjqhj00This benchmark evaluates large language models' ability to perform fine-grained, region-specific semantic reasoning on remote sensing image pairs. It probes localized change comprehension by asking models to answer binary, multiple-choice, and open-ended questions about specific changes (e.g., new construction, vegetation loss) within satellite imagery. Use when the user wants to benchmark on RSRCC, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Ruler Eval · qhjqhj00This benchmark evaluates long-context language models' ability to retrieve, trace, aggregate, and answer questions across varying context lengths and task complexities. It probes whether models genuinely attend to injected information or rely on parametric knowledge and context copying as sequence length increases. Use when the user wants to benchmark on RULER, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Runningsum · qhjqhj00Compute the RunningSum metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RunningSum, or asks how to score with RunningSum.
- ▌ Rxmrl Eval · qhjqhj00Evaluates a model's ability to maintain context and generate coherent responses in multi-turn dialogue, comparing stateful event-driven architectures against standard decoder-only LLMs. Use when the user wants to benchmark on MRL Curriculum Datasets (derived from TinyStories), or asks about evaluating this task. Reports MRL Reward Score.
- ▌ Sciex Eval · qhjqhj00Evaluates large language models on solving university-level scientific exams in computer science. It probes capabilities in open-ended reasoning, mathematical proof writing, long-form explanations, and multimodal (image-text) understanding across English and German languages. Use when the user wants to benchmark on SciEx, or asks about evaluating this task. Reports Normalized score (0-100%).
- ▌ Selqa Eval · qhjqhj00Evaluates a model's ability to retrieve relevant answer sentences from a document given a question (selection), and to determine whether a document section contains an answer at all (triggering). It probes open-domain QA robustness against paraphrasing, varying question types, and section lengths. Use when the user wants to benchmark on SelQA, or asks about evaluating this task. Reports MAP.
- ▌ Sfiog Eval · qhjqhj00Evaluates large language models' ability to generate coherent, logically structured financial investment opinions based on company context and questions. It probes reasoning over mere knowledge retrieval by testing performance across varying degrees of familiarity and novelty in companies and questions. Use when the user wants to benchmark on sFIOG, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Shale Eval · qhjqhj00This benchmark evaluates fine-grained hallucination in Large Vision-Language Models (LVLMs) by testing their faithfulness to visual inputs and factuality against external knowledge. It measures model performance under clean conditions and across hierarchical input perturbations (image, instruction, and combination levels) to assess hallucination resistance. Use when the user wants to benchmark on SHALE, or asks about evaluating this task. Reports accuracy, non-hallucination rate.
- ▌ Share Eval · qhjqhj00Evaluates a model's ability to predict the next item in an anonymous user session based on sequential click history. It probes the model's capacity to capture short-term user intent and higher-order item correlations within dynamic session contexts. Use when the user wants to benchmark on YooChoose, Diginetica, or asks about evaluating this task. Reports Hit@20.
- ▌ Sim3d Eval · qhjqhj00Evaluates 3D anomaly detection and segmentation capabilities in industrial settings using multiview and multimodal (image + depth) inputs. It probes a model's ability to identify and localize defects across multiple object categories under both in-domain (real-to-real) and out-of-domain (synthetic-to-real) conditions. Use when the user wants to benchmark on SiM3D, or asks about evaluating this task. Reports I-AUROC.
- ▌ Slump Eval · qhjqhj00Measures how much final code fidelity degrades when a system's design is progressively disclosed through multi-turn interaction rather than provided upfront. It evaluates semantic faithfulness to a committed design and structural integration of dependencies in long-horizon coding agents. Use when the user wants to benchmark on SLUMP benchmark, or asks about evaluating this task. Reports IF50.
- ▌ Sparc Eval · qhjqhj00Probes cross-domain semantic parsing in context by requiring models to generate sequential SQL queries across multiple conversational turns. It evaluates the ability to maintain state, handle thematic evolution, and generalize to unseen databases while correctly resolving contextual dependencies. Use when the user wants to benchmark on SParC, or asks about evaluating this task. Reports question match.
- ▌ Spert Eval · qhjqhj00Evaluates a model's ability to jointly identify named entity spans with their types and extract relational tuples between them from unstructured text. It probes span-based representation learning, localized context modeling, and joint classification without relying on sequential tagging schemes like BIO. Use when the user wants to benchmark on CoNLL04, SciERC, ADE, or asks about evaluating this task. Reports F1 score (micro/macro-averaged).
- ▌ Spiqa Eval · qhjqhj00Evaluates multimodal long-context reasoning and figure/table comprehension on scientific papers. Tests direct question answering with images, full paper context, and chain-of-thought retrieval capabilities. Use when the user wants to benchmark on SPIQA, or asks about evaluating this task. Reports L3Score.
- ▌ Squad Eval · qhjqhj00Measures a model's ability to extract precise answer spans from a given context paragraph in response to a natural language question, testing reading comprehension and span prediction. Use when the user wants to benchmark on SQuAD 1.1/2.0, or asks about evaluating this task. Reports F1.
- ▌ Stark Eval · qhjqhj00Evaluates retrieval models on their ability to find relevant entities in semi-structured knowledge bases using complex queries that combine textual descriptions and relational constraints. It probes joint reasoning over mixed textual-relational semantics and user-intent modeling across product, academic, and medical domains. Use when the user wants to benchmark on STaRK, or asks about evaluating this task. Reports Hit@k.
- ▌ Statscores · qhjqhj00Compute the StatScores metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute StatScores, or asks how to score with StatScores.
- ▌ Svamp Eval · qhjqhj00Evaluates whether NLP models can genuinely solve simple math word problems through arithmetic reasoning versus relying on shallow heuristics like bag-of-words matching or positional cues. It probes model brittleness by testing performance on standard datasets alongside carefully perturbed variants that remove questions or alter operator types. Use when the user wants to benchmark on MAWPS, ASDiv-A, SVAMP, or asks about evaluating this task. Reports accuracy.
- ▌ Tcav Score · qhjqhj00Measures the influence of human-defined emotional concepts (physiognomy, utterance polarity, voice pitch) on a multimodal emotion recognition model's decisions using Concept Activation Vectors. It quantifies how much each concept drives the model's classification decisions across different network layers. Use when the user has predictions and gold and needs to compute TCAV score.
- ▌ Tid 8 Eval · qhjqhj00Probes a model's ability to learn from inherently subjective or disagreed-upon annotations by treating each annotator's label as a separate example, rather than aggregating them into a single ground truth label. Use when the user wants to benchmark on TID-8, or asks about evaluating this task. Reports exact match accuracy.
- ▌ Timer Eval · qhjqhj00This benchmark evaluates a model's ability to perform temporal reasoning and extract accurate information from longitudinal electronic health records (EHRs). It probes whether models can correctly synthesize evidence across multiple time-stamped clinical visits, adhere to specified temporal boundaries, and maintain accuracy over long patient timelines. Use when the user wants to benchmark on TIMER-Bench, MedAlign, or asks about evaluating this task. Reports Correct.
- ▌ Tnl2k Eval · qhjqhj00Evaluates natural language-based tracking on 2000 YouTube and surveillance videos, testing the model's ability to follow and adapt to language descriptions over time. Use when the user wants to benchmark on TNL2K, or asks about evaluating this task. Reports AUC.
- ▌ Tnllt Eval · qhjqhj00Evaluates long-term vision-language tracking capability by measuring localization accuracy over extended video sequences while dynamically updating natural language descriptions to handle appearance changes and occlusions. Use when the user wants to benchmark on TNLLT, or asks about evaluating this task. Reports PR.
- ▌ Trace Eval · qhjqhj00This protocol evaluates training-free partial audio deepfake detection by analyzing the temporal continuity of frozen speech foundation model embeddings. It probes a model's ability to detect splice boundaries and synthetic insertions in speech without requiring labeled training data or architectural modifications. Use when the user wants to benchmark on PartialSpoof, HalfTruth Audio Deepfake (HAD), ADD 2023 Track 2, LlamaPartialSpoof, or asks about evaluating this task. Reports EER.
- ▌ Tsaia Eval · qhjqhj00Evaluates LLMs' ability to perform multi-step, constraint-aware reasoning and inference on real-world time series data. It probes compositional reasoning, numerical precision, and the capacity to assemble complex analytical or forecasting workflows via executable code generation. Use when the user wants to benchmark on TSAIA, or asks about evaluating this task. Reports Success Rate.
- ▌ Tsaqa Eval · qhjqhj00Evaluates large language models' ability to perform time series analysis and reasoning across six tasks (anomaly detection, classification, characterization, comparison, data transformation, and temporal relationship) using three question formats (true-or-false, multiple-choice, and puzzling). Use when the user wants to benchmark on TSAQA, or asks about evaluating this task. Reports accuracy.
- ▌ Tsbow Eval · qhjqhj00Evaluates object detection models on traffic surveillance footage under diverse weather conditions and varying degrees of vehicle occlusion. It probes robustness to environmental degradation, scale variation, and dense urban traffic scenarios. Use when the user wants to benchmark on TSBOW, or asks about evaluating this task. Reports mAP50.
- ▌ Tsrec Eval · qhjqhj00Evaluates a model's ability to perform sequential recommendation with a focus on capturing repeat-aware temporal patterns. It measures how well the model balances predicting new items versus recurring items based on user interaction history and time intervals. Use when the user wants to benchmark on RetailRocket, LastFM, Diginetica, or asks about evaluating this task. Reports HR@K.
- ▌ Tsver Eval · qhjqhj00This benchmark evaluates an AI system's ability to perform fact verification using time-series evidence. It probes multi-timeframe temporal reasoning, cross-series numerical analysis, and the generation of factually consistent justifications aligned with human annotations. Use when the user wants to benchmark on TSVer, or asks about evaluating this task. Reports Accuracy.
- ▌ Tt Df Eval · qhjqhj00Evaluates the ability of models to detect human body forgeries generated by diffusion models. It probes spatiotemporal motion inconsistencies and generalization across different generation configurations and unseen manipulation models. Use when the user wants to benchmark on TT-DF, or asks about evaluating this task. Reports AUC.
- ▌ Ugr16 Eval · qhjqhj00Evaluates how data preprocessing choices—such as observation selection, flow directionality, and feature engineering—affect the performance of unsupervised anomaly detection models on network traffic. Use when the user wants to benchmark on UGR'16, or asks about evaluating this task. Reports AUC.
- ▌ Unite Eval · qhjqhj00Evaluates text-to-SQL models on compositional generalization, out-of-domain robustness, and schema-question alignment across 18 diverse datasets and 12 domains. It probes the model's ability to handle long-form query decomposition and cross-domain SQL pattern diversity. Use when the user wants to benchmark on UNITE, or asks about evaluating this task. Reports accuracy.
- ▌ Unpie Eval · qhjqhj00Assesses multimodal language models' ability to resolve lexical ambiguity in puns using visual context. It probes visual-textual alignment, multimodal literacy, and the capacity to disambiguate or reconstruct ambiguous text when provided with explanatory or disambiguating images. Use when the user wants to benchmark on UNPIE, or asks about evaluating this task. Reports exact-match accuracy.