qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Pmmeval Eval · qhjqhj00Evaluates multilingual capabilities of LLMs across understanding, reasoning, and generation tasks in 10 languages. It probes prompt sensitivity and cross-lingual performance consistency to reveal benchmark origin bias and language-specific scaling trends. Use when the user wants to benchmark on MMMLU, MLogiQA, MGSM, MHellaSwag, XNLI, Flores-200, or asks about evaluating this task. Reports accuracy.
- ▌ Pokegym Eval · qhjqhj00PokeGym evaluates vision-language models' ability to perform long-horizon planning and spatial reasoning in a complex 3D open-world game using only raw RGB observations. It specifically probes visual grounding, autonomous goal decomposition, and physical deadlock recovery, revealing whether models can navigate cluttered environments, interact with objects, and recover from entrapment without explicit state feedback. Use when the user wants to benchmark on PokeGym, or asks about evaluating this task. Reports task_completion.
- ▌ Polaris Eval · qhjqhj00Evaluates the ability of machine learning models to distinguish between reference stars and circumstellar exoplanetary disks in high-contrast polarimetric imaging data. It probes representation learning quality through downstream supervised classification and unsupervised clustering tasks. Use when the user wants to benchmark on POLARIS, or asks about evaluating this task. Reports accuracy.
- ▌ Posenet Eval · qhjqhj00Evaluates a model's ability to estimate 6-DOF camera pose (translation and rotation) from a single monocular image across indoor and outdoor environments. It probes the network's robustness to challenging conditions like motion blur, low light, and dynamic objects, as well as its generalization to unseen scenes and varying training baselines. Use when the user wants to benchmark on 7 Scenes, Cambridge Landmarks, or asks about evaluating this task. Reports localization error.
- ▌ Probenc Eval · qhjqhj00Evaluates multimodal foundation models on open-ended, expert-level queries across 10 professional domains, probing visual perception, domain knowledge, and long-context reasoning in single-round, multi-lingual, and multi-turn settings. Use when the user wants to benchmark on ProBench, or asks about evaluating this task. Reports ELO rating.
- ▌ Progait Eval · qhjqhj00Evaluates vision models on prosthesis-specific video understanding, including instance segmentation of amputees and prosthetic limbs, 2D human pose estimation with focus on lower-body keypoints, and automated gait pattern classification from pose sequences. Use when the user wants to benchmark on ProGait, or asks about evaluating this task. Reports mIoU, AP@[0.5,0.95].
- ▌ Psbench Eval · qhjqhj00Evaluates the ability of Estimation of Model Accuracy (EMA) methods to predict the structural quality of protein complex models. It probes global and interface-level accuracy estimation using correlation, ranking, and classification metrics against reference structural scores. Use when the user wants to benchmark on CASP16_inhouse_TOP5_dataset, CASP16_community_dataset, or asks about evaluating this task. Reports Pearson’s correlation (CorrP).
- ▌ Psiloqa Eval · qhjqhj00Evaluates the ability of models to detect span-level hallucinations in multilingual question-answering contexts. It probes cross-lingual generalization and token-level inconsistency detection between generated answers and ground truth. Use when the user wants to benchmark on PsiloQA, or asks about evaluating this task. Reports IoU.
- ▌ Psyeval Eval · qhjqhj00Evaluates an AI model's ability to generate empathetic, principle-constrained psychological counseling dialogues in simulated multi-turn interactions. It probes competencies like accurate empathy, logical consistency, resistance handling, and ethical guidance beyond surface-level language features. Use when the user wants to benchmark on PsyEval, or asks about evaluating this task. Reports PsyEval.
- ▌ Pulselm Eval · qhjqhj00This benchmark evaluates multimodal physiological reasoning by testing whether large language models can accurately answer closed-ended questions conditioned on raw photoplethysmography (PPG) waveforms. It probes the model's ability to align continuous biosignal representations with natural language queries across diverse physiological domains and assesses cross-dataset generalization beyond the training distribution. Use when the user wants to benchmark on PulseLM, or asks about evaluating this task. Reports exact-match (EM) accuracy.
- ▌ Pytrial Eval · qhjqhj00Evaluates machine learning models across clinical trial tasks including patient and trial outcome prediction, trial search, and patient simulation, using standardized tabular and sequential data formats. Use when the user wants to benchmark on Tabular Clinical Trial Patient Datasets, TOP Benchmark, Trial Similarity Dataset, Sequential Trial Patient Data, or asks about evaluating this task. Reports AUROC.
- ▌ Quality Eval · qhjqhj00This benchmark evaluates a model's ability to comprehend and reason over long documents (2k–8k tokens) to answer multiple-choice questions. It specifically probes whether models can integrate global context rather than relying on local keyword matching or summaries, with a subset (HARD) filtering for questions that require full reading rather than skimming. Use when the user wants to benchmark on QuALITY, or asks about evaluating this task. Reports accuracy.
- ▌ R Bench Eval · qhjqhj00Evaluates complex reasoning capabilities of LLMs and MLLMs on graduate-level, multi-disciplinary academic questions in both English and Chinese. It probes the models' ability to handle rigorous, curriculum-based problems requiring extended chain-of-thought reasoning. Use when the user wants to benchmark on R-Bench-T, R-Bench-M, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ R Judge Eval · qhjqhj00Evaluates LLMs' ability to judge safety risks in multi-turn agent interactions by classifying whether a given interaction record poses a safety risk. It probes risk perception and binary safety classification under zero-shot and few-shot prompting conditions, with and without explicit risk descriptions. Use when the user wants to benchmark on R-Judge, or asks about evaluating this task. Reports F1.
- ▌ Racecar Eval · qhjqhj00Evaluates high-speed autonomous driving capabilities including precise localization, long-range object detection and tracking, and robust mapping/SLAM under extreme dynamic conditions (up to 170 mph). The protocol benchmarks how well models maintain accuracy and latency when processing multi-modal sensor data at racing speeds where motion blur, sensor dropout, and rapid ego-motion are prevalent. Use when the user wants to benchmark on RACECAR, or asks about evaluating this task. Reports Average Precision (AP).
- ▌ Radarqa Eval · qhjqhj00Evaluates multi-modal large language models on weather radar forecast quality analysis, specifically testing their ability to perform quantitative rating of radar frames/sequences and generate qualitative assessment reports. It probes domain-specific meteorological understanding, temporal pattern evolution tracking, and alignment with expert judgment. Use when the user wants to benchmark on RQA-70K, or asks about evaluating this task. Reports accuracy.
- ▌ RAG Har Eval · qhjqhj00Evaluates a training-free, retrieval-augmented framework for classifying human activities from wearable sensor time-series data. It probes the model's ability to perform open-world activity recognition by retrieving semantically similar sensor examples and using an LLM to predict activity labels without fine-tuning. Use when the user wants to benchmark on HHAR, PAMAP2, MHEALTH, GOTOV, SKODA, USC-HAD, or asks about evaluating this task. Reports accuracy.
- ▌ Rainnet Eval · qhjqhj00Evaluates deep learning models for spatial precipitation downscaling by measuring both static reconstruction accuracy and dynamic temporal evolution of rainfall patterns. It probes whether models can capture realistic meteorological properties like heavy rain coverage, cluster movement, and transition speeds. Use when the user wants to benchmark on RainNet, or asks about evaluating this task. Reports PEM.
- ▌ Realcqa Eval · qhjqhj00Evaluates scientific chart question answering capabilities, specifically testing a model's ability to extract, reason over, and answer questions about real-world scientific charts. It probes first-order logic reasoning and neuro-symbolic capabilities by requiring formal verification of logical inferences from complex visual data. Use when the user wants to benchmark on RealCQA, or asks about evaluating this task. Reports accuracy.
- ▌ Rec Auc Eval · qhjqhj00Evaluates the predictive performance of recommendation models on large-scale click-through rate datasets. It specifically probes how model scalability and embedding size affect ranking quality, revealing the phenomenon of embedding collapse when scaling up feature interactions. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.
- ▌ Recall Score · qhjqhj00Compute the recall_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute recall_score, or asks how to score with recall_score.
- ▌ Ref Adv Eval · qhjqhj00Probes multimodal large language models' ability to perform visual grounding and complex textual reasoning under challenging conditions. It specifically tests whether models rely on shortcut cues or genuinely comprehend referring expressions when faced with linguistically nontrivial descriptions and hard distractors. Use when the user wants to benchmark on Ref-Adv, or asks about evaluating this task. Reports Acc0.5.
- ▌ Ref Avs Eval · qhjqhj00Evaluates a model's ability to segment objects in audio-visual videos based on natural language referring expressions. It tests both seen categories and generalization to unseen categories, as well as handling null references where no object exists. Use when the user wants to benchmark on Ref-AVS Dataset, or asks about evaluating this task. Reports Jaccard Index ($\mathcal{J}$).
- ▌ Refedit Eval · qhjqhj00Evaluates instruction-based image editing models on referring expressions, measuring how well they align edits with text instructions while preserving background and maintaining perceptual quality. It probes multi-object scene editing, background preservation, and the ability to handle complex spatial grounding without relying on CLIP. Use when the user wants to benchmark on RefEdit-Bench, PIE-Bench, or asks about evaluating this task. Reports VIEScore.
- ▌ Refsgrs Eval · qhjqhj00Evaluates a model's ability to segment specific objects in remote sensing images guided by natural language expressions, with a focus on accurately localizing small, scattered targets that are characteristic of aerial and satellite imagery. Use when the user wants to benchmark on RefSegRS, or asks about evaluating this task. Reports mIoU.
- ▌ Repliqa Eval · qhjqhj00Evaluates LLMs' ability to read unseen reference documents and answer questions based solely on the provided context, as well as their ability to detect unanswerable questions and classify document topics. It specifically probes whether models rely on pre-training memory versus actual context-conditional reading and retrieval skills. Use when the user wants to benchmark on RepLiQA, TriviaQA, or asks about evaluating this task. Reports recall.
- ▌ Retrievalmap · qhjqhj00Compute the RetrievalMAP metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalMAP, or asks how to score with RetrievalMAP.
- ▌ Retrievalmrr · qhjqhj00Compute the RetrievalMRR metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalMRR, or asks how to score with RetrievalMRR.
- ▌ Rexrank Eval · qhjqhj00Evaluates AI models' ability to generate accurate and clinically relevant radiology reports from chest X-ray images. It assesses both linguistic quality and clinical entity extraction/alignment across diverse clinical datasets. Use when the user wants to benchmark on ReXGradient, MIMIC-CXR, IU X-ray, CheXpert Plus, or asks about evaluating this task. Reports 1/RadCliQ-v1.
- ▌ Rgb Har Eval · qhjqhj00Evaluates the ability of a skeleton-based BLSTM model to recognize human actions from RGB-only video streams under limited labeled data conditions, comparing against methods that use depth or inertial modalities. Use when the user wants to benchmark on UTD-MHAD, KTH, or asks about evaluating this task. Reports top-1 accuracy.
- ▌ Ris Lad Eval · qhjqhj00Evaluates referring image segmentation on low-altitude drone imagery, probing the model's ability to accurately localize and segment referred objects despite challenges like category drift (tiny objects) and object drift (dense same-category scenes). Use when the user wants to benchmark on RIS-LAD, or asks about evaluating this task. Reports oIoU, mIoU.
- ▌ Rjua QA Eval · qhjqhj00This benchmark evaluates large language models' ability to perform medical logical reasoning and urological disease diagnosis. It probes the model's capacity to handle complex, real-world clinical scenarios involving subjective patient queries and multi-disease comorbidity reasoning. Use when the user wants to benchmark on RJUA-QA, or asks about evaluating this task. Reports F1 score (diagnosis & advice).
- ▌ Rlirank Eval · qhjqhj00Evaluates a reinforcement learning framework for dynamic search ranking that adapts to evolving user intents over multiple search iterations using sequential feedback. It also benchmarks standard learning-to-rank performance on static datasets. Use when the user wants to benchmark on TREC 2016 Dynamic Domain, TREC 2017 Dynamic Domain, MQ2007, MQ2008, or asks about evaluating this task. Reports α-NDCG.
- ▌ Rmbench Eval · qhjqhj00This benchmark evaluates robotic manipulation policies on memory-dependent, non-Markovian tasks. It probes a model's ability to retain and utilize historical visual and state information over long horizons to complete multi-step dual-arm manipulation sequences. Use when the user wants to benchmark on RMBench, or asks about evaluating this task. Reports success rate.
- ▌ Robocse Eval · qhjqhj00Evaluates a robot's ability to generalize semantic knowledge by predicting object affordances, locations, and materials in unseen environments and inferring ranks of unseen semantic triples. It probes multi-relational embedding performance for common-sense reasoning in residential robotics. Use when the user wants to benchmark on AI2Thor, or asks about evaluating this task. Reports MRR.
- ▌ Robonar Eval · qhjqhj00Evaluates a multimodal robot narration framework's ability to select key events, generate natural language summaries, and perform failure analysis (risk estimation, localization, explanation, recovery) on real-world household robot tasks. Use when the user wants to benchmark on RoboNar, or asks about evaluating this task. Reports Accuracy on failure analysis tasks.
- ▌ Rsbench Eval · qhjqhj00Evaluates the ability of LLM-based evolutionary algorithms to optimize session-based recommendation prompts across multiple objectives (accuracy, diversity, and fairness) simultaneously. Use when the user wants to benchmark on RSBench, or asks about evaluating this task. Reports HV.
- ▌ Rwds Cz Eval · qhjqhj00Evaluates object detectors' robustness to real-world spatial domain shifts across different climate zones in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-CZ, or asks about evaluating this task. Reports mAP.
- ▌ Rwds Fr Eval · qhjqhj00Evaluates object detectors' robustness to real-world spatial domain shifts across flood-affected regions in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-FR, or asks about evaluating this task. Reports mAP.
- ▌ Rwds He Eval · qhjqhj00Evaluates object detectors' robustness to real-world spatial domain shifts across hurricane-affected regions in satellite imagery. It measures how well models generalize from in-domain training data to out-of-distribution target domains without fine-tuning. Use when the user wants to benchmark on RWDS-HE, or asks about evaluating this task. Reports mAP.
- ▌ Safavid Eval · qhjqhj00Evaluates the safety alignment and refusal capabilities of Video Large Multimodal Models (VLMMs) against everyday adversarial queries and covert, human-red-teamed prompts. It measures whether models can maintain safety guidelines across diverse harmful categories without compromising general utility or falling back to memorized refusals. Use when the user wants to benchmark on SafeVidBench, or asks about evaluating this task. Reports Safety Rate.
- ▌ Safeqil Eval · qhjqhj00This evaluation probes an agent's ability to learn safe navigation and manipulation policies from human demonstrations in environments with unknown safety constraints. It specifically tests the trade-off between maximizing task reward and minimizing safety violations (cost) under out-of-distribution conditions. Use when the user wants to benchmark on Safety-Gymnasium (SafetyPointGoal1-v0, SafetyPointCircle2-v0, SafetyCarButton1-v0, SafetyCarPush2-v0), or asks about evaluating this task. Reports episodic reward.
- ▌ Safety Score · qhjqhj00Quantifies implicit representational harms in pre-trained language models by measuring the disparity in language modeling probabilities between harmful and benign sentences targeting 13 marginalized demographics. It probes whether a model's internal likelihood estimates reflect toxic or stereotypical biases toward specific groups. Use when the user has predictions and gold and needs to compute safety score.
- ▌ Salt Kg Eval · qhjqhj00Probes whether tabular models can effectively leverage declarative business knowledge and metadata semantics for prediction tasks, rather than relying solely on statistical correlations in raw features. It evaluates the impact of schema-grounded semantic embeddings on model inductive biases and relative performance across different model families. Use when the user wants to benchmark on SALT-KG, or asks about evaluating this task. Reports ranking metrics.
- ▌ Scendi Score · qhjqhj00Evaluates the intrinsic diversity of text-to-image generative models by isolating model-driven variation from prompt-driven variation. It uses CLIP embeddings to construct a joint image-text kernel covariance matrix and applies Schur complement decomposition to remove text influence before computing spectral entropy. Use when the user has predictions and gold and needs to compute Scendi score.
- ▌ Scicode Eval · qhjqhj00Probes large language models' ability to perform scientific reasoning, domain-specific knowledge recall, and code synthesis on real-world research problems. It evaluates performance on both decomposed subproblems and full main problems under varying conditions of background knowledge and context carry-over. Use when the user wants to benchmark on SciCode, or asks about evaluating this task. Reports pass@1.
- ▌ Scieval Eval · qhjqhj00Evaluates a model's ability to automatically assess K-12 science instructional materials against pedagogical rubrics. It probes domain-aligned reasoning, long-context evidence grounding, and the capacity to generate rubric-consistent scores and justifications. Use when the user wants to benchmark on SciEval, or asks about evaluating this task. Reports Evidence Match Rate (EMR).
- ▌ Scitldr Eval · qhjqhj00This benchmark evaluates the ability of models to generate extreme, single-sentence summaries (TLDRs) of scientific papers, capturing key contributions while bypassing background details. It tests both automated overlap metrics and human-judged informativeness and correctness under multi-target and multi-input settings. Use when the user wants to benchmark on SCITLDR, or asks about evaluating this task. Reports Rouge-1.
- ▌ Scitrek Eval · qhjqhj00Evaluates long-context language models' ability to perform numerical aggregation, filtering, sorting, and logical operations across extended contexts (up to 1M tokens) using scientific article metadata and full-text articles. Use when the user wants to benchmark on SciTrek, or asks about evaluating this task. Reports exact match.
- ▌ Scizoom Eval · qhjqhj00Evaluates a model's ability to perform hierarchical scientific summarization by generating three distinct granularity levels (Abstract, Key Contributions, TL;DR) from a single full-text input. It probes multi-granularity text compression and the model's capacity to maintain coherence across varying compression ratios within a single inference pass. Use when the user wants to benchmark on SciZoom, or asks about evaluating this task. Reports unspecified summarization metric.
- ▌ Scrolls Eval · qhjqhj00Evaluates long-text understanding capabilities across summarization, question answering, and natural language inference tasks. It probes whether models can effectively process and extract information from documents exceeding standard context windows (up to 16K tokens) using chunked encoding and cross-chunk fusion. Use when the user wants to benchmark on SCROLLS, or asks about evaluating this task. Reports Avg SCROLLS score.
- ▌ Seaexam Eval · qhjqhj00Evaluates LLMs' ability to answer local, culturally grounded multiple-choice questions in Southeast Asian languages (Indonesian, Thai, Vietnamese). It probes regional knowledge, language comprehension, and alignment with actual local usage compared to translated benchmarks. Use when the user wants to benchmark on SeaExam, or asks about evaluating this task. Reports accuracy (%).
- ▌ Sec Gfd Eval · qhjqhj00Evaluates graph neural networks for fraud detection on real-world transaction and review graphs. It specifically probes a model's robustness to severe class imbalance and heterophily, where connected nodes often belong to different classes. Use when the user wants to benchmark on Amazon, YelpChi, T-Finance, T-Social, or asks about evaluating this task. Reports F1-macro, AUC.
- ▌ Seegull Eval · qhjqhj00Probes a model's propensity to generate or recognize stereotypical associations across diverse global and state-level identity groups. Evaluates the prevalence and cultural specificity of biases in English NLP models, highlighting regional disparities in stereotype content and offensiveness. Use when the user wants to benchmark on SeeGULL, or asks about evaluating this task. Reports stereotype_prevalence.
- ▌ Seephys Eval · qhjqhj00This benchmark evaluates multimodal LLMs' ability to perform physics reasoning using visual diagrams, text, or both. It probes visual interpretation, diagram-to-reasoning mapping, and the model's reliance on textual shortcuts versus actual visual perception across varying knowledge levels and diagram types. Use when the user wants to benchmark on SeePhys, or asks about evaluating this task. Reports accuracy.
- ▌ Segbook Eval · qhjqhj00Evaluates the transfer learning and fine-tuning capabilities of volumetric medical image segmentation models across diverse imaging modalities, anatomical targets, and dataset sizes. It probes how well pre-trained models generalize to downstream segmentation tasks in clinical imaging scenarios and reveals non-linear performance scaling with dataset scale. Use when the user wants to benchmark on SegBook, or asks about evaluating this task. Reports Dice Score (DSC).
- ▌ Sheriff Eval · qhjqhj00Evaluates EFCE solvers on a parametric sequential bargaining game modeling smuggling and inspection. It probes the solver's ability to handle multi-round negotiations, bribery, and deterrence to maximize social welfare. Use when the user wants to benchmark on Sheriff, or asks about evaluating this task. Reports Social Welfare (SW).
- ▌ Simmmdg Eval · qhjqhj00Evaluates a model's ability to generalize across unseen domains in multi-modal action recognition. It probes feature disentanglement, cross-modal translation for missing modalities, and robustness to domain shifts in video, audio, and optical flow inputs. Use when the user wants to benchmark on EPIC-Kitchens, HAC, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Siwarex Eval · qhjqhj00Evaluates LLM-based natural language question answering over heterogeneous data sources by testing the model's ability to generate SQL queries that correctly invoke both database tables and external APIs. It probes complex query planning, API sequencing, and routing across mixed data modalities. Use when the user wants to benchmark on Spider (modified with API-replaced tables), or asks about evaluating this task. Reports execution accuracy.
- ▌ Sparrta Eval · qhjqhj00This benchmark evaluates the spatial reasoning capabilities of Visual Foundation Models (VFMs) by testing their ability to recognize spatial relations between object triples in synthetic images. It specifically probes both egocentric (camera-perspective) and allocentric (world-perspective) spatial understanding across diverse semantic objects and environments. Use when the user wants to benchmark on SpaRRTa, or asks about evaluating this task. Reports accuracy.
- ▌ Spartqa Eval · qhjqhj00Probes a model's ability to perform multi-step spatial reasoning over natural language stories. It evaluates understanding of spatial relations (e.g., near, far, containment) and tests robustness against surface-level vocabulary changes and question phrasing variations. Use when the user wants to benchmark on SPARTQA-HUMAN, or asks about evaluating this task. Reports accuracy.
- ▌ Spce 10 Eval · qhjqhj00This benchmark evaluates multimodal large language models on compositional spatial intelligence by testing their ability to reason across 10 atomic spatial capabilities (e.g., counting, localization, spatial relations) combined into 8 complex tasks. It probes scene understanding using 2D and 3D inputs through multiple-choice questions, revealing how models handle hierarchical spatial reasoning and capability integration. Use when the user wants to benchmark on SpaCE-10, or asks about evaluating this task. Reports accuracy.
- ▌ Specweb Eval · qhjqhj00Evaluates the energy proportionality and power efficiency of enterprise server subsystems under varying PHP-based e-commerce web workloads. It measures how power consumption scales with session load and identifies non-proportional power draw in uncore components. Use when the user wants to benchmark on SPECweb2009, or asks about evaluating this task. Reports watts.
- ▌ Speechr Eval · qhjqhj00Probes speech reasoning capabilities in large audio-language models across factual, procedural, and normative dimensions. It tests whether models can perform multi-step inference, maintain logical coherence, and make normative judgments when processing spoken input under varying prosodic and emotional conditions. Use when the user wants to benchmark on SpeechR, or asks about evaluating this task. Reports Accuracy.
- ▌ Srbench Eval · qhjqhj00Evaluates sequential recommendation models across accuracy, fairness, stability, and efficiency dimensions. It tests whether models can correctly rank items based on user interaction history and assesses their robustness, bias, and computational cost. Use when the user wants to benchmark on Yelp, ML-100K, Beauty, or asks about evaluating this task. Reports Recall@5.
- ▌ Ssa Mte Eval · qhjqhj00Evaluates machine translation quality estimation metrics on under-resourced African languages by comparing their predicted scores against human-annotated Direct Assessment (DA) judgments. It probes a model's ability to correlate with human perception of translation adequacy across diverse language pairs, including both reference-based and reference-free settings. Use when the user wants to benchmark on SSA-MTE, or asks about evaluating this task. Reports Spearman correlation.
- ▌ Superni Eval · qhjqhj00Evaluates a language model's ability to follow diverse natural language instructions across various NLP tasks in a zero-shot setting. It measures how well the model generalizes to unseen tasks without in-context examples. Use when the user wants to benchmark on SUPER-NATURALINSTRUCTIONS, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Surgveo Eval · qhjqhj00Evaluates zero-shot surgical video generation models using a four-tiered Surgical Plausibility Pyramid. It probes the model's ability to maintain visual realism while correctly simulating domain-specific surgical causality, instrument handling, tissue feedback, and clinical intent over time. Use when the user wants to benchmark on SurgVeo benchmark, or asks about evaluating this task. Reports Visual Perceptual Plausibility.
- ▌ Svbench Eval · qhjqhj00Evaluates large vision-language models' ability to perform sustained temporal reasoning and context tracking across long-form streaming videos. It probes multi-turn dialogue continuity, temporal dependency handling, and complex reasoning skills like counterfactual analysis and spatio-temporal speculation. Use when the user wants to benchmark on SVBench, or asks about evaluating this task. Reports Overall Score (OS).
- ▌ Swim Ir Eval · qhjqhj00Evaluates multilingual dense retrieval models on cross-lingual and monolingual open retrieval tasks. It measures how effectively synthetic LLM-generated training data scales retrieval performance compared to human-labeled baselines across diverse languages and corpus sizes. Use when the user wants to benchmark on XOR-Retrieve, MIRACL, XTREME-UP, or asks about evaluating this task. Reports Recall@mkt.
- ▌ Symlink Eval · qhjqhj00Probes the ability to extract fine-grained mathematical symbols and their textual descriptions from LaTeX-formatted scientific documents. It evaluates both named entity recognition for identifying symbols and descriptions, and relation extraction for linking them according to specific semantic types. Use when the user wants to benchmark on Symlink, or asks about evaluating this task. Reports F-score.
- ▌ T3bench Eval · qhjqhj00Evaluates the visual quality and text-3D alignment of generated 3D scenes across varying prompt complexities (single object, object with surroundings, multiple objects). It specifically probes multi-view consistency (detecting the Janus problem) and the ability of 2D diffusion guidance to translate into coherent 3D structures. Use when the user wants to benchmark on T$^3$ Bench, or asks about evaluating this task. Reports Multi-view Quality (ImageReward), Alignment (GPT-4).
- ▌ Tabfact Eval · qhjqhj00Tests a model's ability to verify the truthfulness of a factual statement given a table. It probes structured reasoning capabilities by requiring the model to cross-reference table contents with a claim and output a binary label. Use when the user wants to benchmark on TabFact, or asks about evaluating this task. Reports binary classification accuracy.
- ▌ Tabshap Eval · qhjqhj00This protocol evaluates the faithfulness of feature attributions for LLM-based tabular classifiers. It measures how well an attribution method ranks features by sequentially masking them in importance order and tracking the resulting drop in the model's predicted class probability. Use when the user wants to benchmark on Adult Income, Heart Disease, or asks about evaluating this task. Reports faithfulness.
- ▌ Tat LLM Eval · qhjqhj00Evaluates discrete reasoning capabilities over hybrid tabular and textual data, specifically focusing on financial question answering tasks that require arithmetic operations, counting, and span extraction. Use when the user wants to benchmark on FinQA, TAT-QA, TAT-DQA, or asks about evaluating this task. Reports EM.
- ▌ Tdbench Eval · qhjqhj00Evaluates vision-language models on top-down (aerial) image understanding by testing their ability to answer questions about rotated views. It measures rotational consistency to filter out hallucinations and decomposes performance into true knowledge versus lucky guessing via a probabilistic reliability framework. Use when the user wants to benchmark on TDBench, or asks about evaluating this task. Reports RotationalEval (RE).
- ▌ Teleqna Eval · qhjqhj00Evaluates large language models' domain-specific knowledge in telecommunications, covering general terminology, research concepts, and complex technical standards. It also benchmarks model performance against human telecom professionals under strict no-search conditions. Use when the user wants to benchmark on TeleQnA, or asks about evaluating this task. Reports accuracy.
- ▌ Tgbsseq Eval · qhjqhj00Evaluates temporal graph neural networks on future link prediction tasks, specifically probing their ability to generalize to unseen edges and capture complex sequential dynamics rather than memorizing repeated interactions. Use when the user wants to benchmark on ML-20M, Taobao, Yelp, GoogleLocal, Wikipedia, Reddit, Flickr, YouTube, Patent, WikiLink, or asks about evaluating this task. Reports MRR.
- ▌ Time Ra Eval · qhjqhj00Evaluates the ability of LLMs and MLLMs to diagnose anomalies in univariate and multivariate time series data. It probes the models' capacity for structured reasoning (generating a 'Thought') and precise action classification ('ActionID') based on raw or visualized temporal data. Use when the user wants to benchmark on RATs40K, or asks about evaluating this task. Reports Label Matching F1.
- ▌ Timetom Eval · qhjqhj00Evaluates Large Language Models' Theory of Mind (ToM) reasoning capabilities across reading comprehension and interactive dialogue scenarios. It specifically probes the model's ability to track character beliefs over time, assess answerability, and determine information access, with a strong focus on first-order and higher-order (up to third-order) belief reasoning. Use when the user wants to benchmark on ToMI, BigToM, FanToM, or asks about evaluating this task. Reports accuracy.
- ▌ Tod Nlg Eval · qhjqhj00Evaluates the ability of task-oriented dialogue systems to generate natural language responses while maintaining entity consistency and completing user goals across multiple domains. It probes end-to-end dialogue generation, dialogue state tracking, and response quality under both automated simulation and human evaluation. Use when the user wants to benchmark on DSTC8 Track 1 End-to-End Multi-Domain Dialogue Challenge, MultiWOZ 2.0 benchmark, or asks about evaluating this task. Reports Success Rate.
- ▌ Toolemu Eval · qhjqhj00Evaluates the safety and helpfulness of language model agents interacting with tools in a simulated environment. It measures how well an automated emulator and evaluator align with human judgments, and quantifies agent failure rates under standard and adversarial conditions. Use when the user wants to benchmark on ToolEmu Agent Trajectories, or asks about evaluating this task. Reports Cohen's κ (Quadratic-weighted).
- ▌ Toximol Eval · qhjqhj00Evaluates whether Multimodal Large Language Models (MLLMs) can generate structurally valid, low-toxicity alternative molecules from toxic inputs while adhering to drug-likeness, synthetic feasibility, and structural similarity constraints. It probes the model's ability to perform structure-aware molecular editing and cross-modal scientific reasoning. Use when the user wants to benchmark on ToxiMol, or asks about evaluating this task. Reports Toxicity Repair Success Rate.
- ▌ Kairos Eval · qhjqhj00Evaluates a provenance-based intrusion detection system's ability to identify anomalous system behavior and reconstruct attack footprints from whole-system kernel-level logs. It probes the model's capacity to distinguish between benign and malicious activity in temporal windows without relying on attack signatures. Use when the user wants to benchmark on Manzoor et al., DARPA-E3-THEIA, DARPA-E3-CADETS, DARPA-E3-ClearScope, DARPA-E5-THEIA, DARPA-E5-CADETS, DARPA-E5-ClearScope, DARPA-OpTC, or asks about evaluating this task. Reports AUC.
- ▌ Kalahi Eval · qhjqhj00This benchmark probes an LLM's ability to understand and generate culturally appropriate responses for Filipino contexts. It evaluates whether models can align with the lived experiences, values, and preferred strategies of action of average native Filipino speakers across nuanced socio-cultural scenarios. Use when the user wants to benchmark on Kalahi, or asks about evaluating this task. Reports MC1.
- ▌ Kg Vip Eval · qhjqhj00Evaluates multi-modal LLMs' ability to perform visual question answering by grounding visual inputs with external knowledge graphs. It probes the model's capacity for multi-hop reasoning, visual perception, and knowledge retrieval-augmented generation. Use when the user wants to benchmark on FVQA 2.0+, MVQA, or asks about evaluating this task. Reports LLM-J.
- ▌ Kgquiz Eval · qhjqhj00Evaluates large language models' ability to store, retrieve, and reason over factual knowledge encoded in parametric memory across five progressively complex tasks. It probes basic fact verification, multiple-choice discrimination, open-ended entity generation, multi-hop factual editing, and comprehensive entity description generation. It measures how well models generalize encoded knowledge across commonsense, encyclopedic, and biomedical domains under increasing reasoning complexity. Use when the user wants to benchmark on KGQuiz, or asks about evaluating this task. Reports accuracy.
- ▌ Lae 1m Eval · qhjqhj00Evaluates open-vocabulary and closed-set object detection capabilities on remote sensing imagery. It probes a model's ability to detect novel Earth-based objects without prior training on them, as well as its efficiency when fine-tuned with limited labeled data. Use when the user wants to benchmark on LAE-1M, DIOR, DOTAv2.0, LAE-80C, or asks about evaluating this task. Reports mAP.
- ▌ Lasana Eval · qhjqhj00Evaluates deep learning models on video-based laparoscopic surgical training tasks. It probes the model's ability to recognize task-specific procedural errors and predict structured global skill ratings from synchronized stereo video streams. Use when the user wants to benchmark on LASANA, or asks about evaluating this task. Reports error_recognition.
- ▌ Lav Df Eval · qhjqhj00This benchmark evaluates the capability of models to detect and temporally localize content-driven audio-visual forgeries in long videos. It probes multimodal boundary matching and temporal manipulation detection by requiring models to identify fake segments and predict their precise start and end timestamps. Use when the user wants to benchmark on LAV-DF, or asks about evaluating this task. Reports AP@0.5.
- ▌ Layerd Eval · qhjqhj00Evaluates a model's ability to decompose raster graphic designs into a sequence of re-editable layers. It measures visual reconstruction quality and the number of edits required to match a ground-truth layer structure, accounting for the ill-posed nature of layer ordering. Use when the user wants to benchmark on Crello, or asks about evaluating this task. Reports RGB L1, Alpha IoU.
- ▌ Lexrel Eval · qhjqhj00Probes large language models' ability to extract structured legal relations (relation types and factual arguments) from Chinese civil court judgments. It evaluates both zero-shot prompting and fine-tuning capabilities, while also measuring performance on long-tail relation types and downstream legal reasoning tasks. Use when the user wants to benchmark on LexRel, or asks about evaluating this task. Reports micro-F1.
- ▌ Libero Eval · qhjqhj00Evaluates a robot policy's ability to sequentially learn multiple manipulation tasks while transferring knowledge and minimizing catastrophic forgetting. It measures forward transfer speed, backward transfer (forgetting), and overall performance across a curriculum of procedurally generated tasks. Use when the user wants to benchmark on LIBERO-LONG, LIBERO-SPATIAL, LIBERO-OBJECT, LIBERO-GOAL, or asks about evaluating this task. Reports FWT.
- ▌ Litqa2 Eval · qhjqhj00Evaluates an agentic system's ability to retrieve relevant scientific literature, re-rank passages, and answer multiple-choice questions based on the retrieved text. It probes retrieval coverage, passage localization, and question-answering accuracy under information loss constraints. Use when the user wants to benchmark on LitQA2, or asks about evaluating this task. Reports Accuracy.
- ▌ Llmbar Eval · qhjqhj00This benchmark probes an evaluator model's ability to accurately judge preference between two model responses while resisting bias towards superficial qualities like verbosity, fluency, and formality. It measures instruction-following accuracy by comparing judgments on natural preference data against adversarially crafted instances designed to confound less capable judges. Use when the user wants to benchmark on LLMBar, or asks about evaluating this task. Reports accuracy.
- ▌ Locomo Eval · qhjqhj00This benchmark evaluates the long-term conversational memory of LLM agents by testing their ability to answer questions, summarize events, and generate multi-modal dialogues over very long, multi-session conversations. It probes how well models retain and reason over temporal and causal information across hundreds of turns and thousands of tokens. Use when the user wants to benchmark on LoCoMo, or asks about evaluating this task. Reports F1-score.
- ▌ Loogle Eval · qhjqhj00Evaluates the ability of language models to comprehend and reason over long documents (up to 32k+ tokens) by testing short and long dependency tasks, including question answering, cloze completion, and summarization. Use when the user wants to benchmark on LooGLE, or asks about evaluating this task. Reports GPT4_score.
- ▌ Lumina Eval · qhjqhj00Evaluates deep learning models on multi-vendor full-field digital mammography for breast cancer diagnosis, BI-RADS classification, and breast density prediction. It specifically probes the model's robustness to domain shifts induced by different imaging vendors and X-ray energies. Use when the user wants to benchmark on LUMINA, or asks about evaluating this task. Reports AUC.
- ▌ M Absa Eval · qhjqhj00Evaluates multilingual aspect-based sentiment analysis (ABSA) models on triplet extraction (aspect term, category, sentiment) and pairwise extraction (aspect term, sentiment) across 21 languages and 7 domains. Probes cross-lingual transfer, cross-domain adaptation, and zero-shot LLM prompting capabilities. Use when the user wants to benchmark on M-ABSA, or asks about evaluating this task. Reports Micro-F1.
- ▌ M Beir Eval · qhjqhj00Evaluates multimodal information retrieval models across eight heterogeneous query-to-candidate modalities (text, image, image-text pairs) using instruction-tuned and fine-tuned vision-language models. Probes zero-shot generalization, cross-modality alignment, and the impact of instruction tuning on retrieval accuracy in large-scale candidate pools. Use when the user wants to benchmark on M-BEIR, or asks about evaluating this task. Reports Recall@5.