qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Moleculariq Eval · qhjqhj00Evaluates large language models' ability to perform symbolic reasoning on molecular graphs, including counting atomic features, indexing substructures, and generating constrained molecular structures. It probes whether models understand chemical topology and composition rather than relying on memorized token patterns or canonical SMILES conventions. Use when the user wants to benchmark on MOLECULARIQ, or asks about evaluating this task. Reports accuracy.
- ▌ Monkey Jump Eval · qhjqhj00Evaluates the performance and parameter/memory/throughput efficiency of a gradient-free MoE-style PEFT routing mechanism across text, image, and video benchmarks compared to standard and MoE-PEFT baselines. Use when the user wants to benchmark on 47 Benchmarks (Text/Image/Video), or asks about evaluating this task. Reports accuracy.
- ▌ Mos Rmbench Eval · qhjqhj00Evaluates the ability of speech quality reward models to correctly rank pairs of audio samples based on their Mean Opinion Score (MOS). It probes fine-grained perceptual discrimination and cross-dataset generalization in preference-based audio modeling. Use when the user wants to benchmark on BVCC, NISQA, SingMOS, SOMOS, TMHINT-QI, VMC’23, or asks about evaluating this task. Reports accuracy.
- ▌ Mosaic Cdsr Eval · qhjqhj00Evaluates a model's ability to perform next-item recommendation in cross-domain sequential settings by decomposing user intent into orthogonal preference components. It probes how well the model leverages shared and domain-specific signals across multiple item categories to predict future interactions. Use when the user wants to benchmark on Amazon Reviews (Movie–Book), Amazon Reviews (Movie–Music), Douban (Movie–Book), or asks about evaluating this task. Reports NDCG@10.
- ▌ Mtbi Speech Eval · qhjqhj00Evaluates the automatic speech recognition accuracy and zero-shot generalization capabilities of Speech Large Language Models. It probes robustness across mathematical reasoning, speaker role inference, and prompt adaptation tasks. Use when the user wants to benchmark on LibriSpeech, GSM8K, Generalization Test Set, or asks about evaluating this task. Reports WER.
- ▌ Mteb Eng V2 Eval · qhjqhj00Evaluates the semantic discrimination and generalization capabilities of text embedding models across diverse NLP tasks including retrieval, classification, clustering, and semantic similarity. It uses a zero-shot English-only benchmark to measure performance without task-specific fine-tuning. Use when the user wants to benchmark on MTEB(eng, v2), or asks about evaluating this task. Reports average score across all tasks.
- ▌ Mteb Subset Eval · qhjqhj00Evaluates text embedding models across diverse semantic tasks including retrieval, reranking, clustering, pair classification, classification, and semantic textual similarity to measure the quality of dense vector representations. Use when the user wants to benchmark on MTEB (15-task subset), or asks about evaluating this task. Reports MTEB average score.
- ▌ Muharaf Htr Eval · qhjqhj00Evaluates handwritten text recognition (HTR) systems on historical Arabic manuscripts. It probes the model's ability to accurately transcribe cursive text at both the page and line levels, handling contextual character variations and layout structures. Use when the user wants to benchmark on Muharaf, or asks about evaluating this task. Reports CER.
- ▌ Multi Bench Eval · qhjqhj00Evaluates the emotional intelligence (EI) capabilities of spoken dialogue models in multi-turn interactive settings. It probes basic emotion understanding, advanced emotion support, paralinguistic analysis, and style inference across both Chinese and English dialogues. Use when the user wants to benchmark on MULTI-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Multibanana Eval · qhjqhj00Evaluates text-to-image generation models on their ability to synthesize images from multiple reference images and text prompts. It probes adherence to complex instructions, consistency with reference attributes, and robustness to domain mismatches, scale discrepancies, rare concepts, and multilingual text. Use when the user wants to benchmark on MultiBanana, or asks about evaluating this task. Reports MultiBanana score.
- ▌ Multiclasslogauc · qhjqhj00Compute the MulticlassLogAUC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassLogAUC, or asks how to score with MulticlassLogAUC.
- ▌ Multiclassrecall · qhjqhj00Compute the MulticlassRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassRecall, or asks how to score with MulticlassRecall.
- ▌ Multiconer2 Eval · qhjqhj00Evaluates a model's ability to perform fine-grained named entity recognition and entity linking across multiple languages. It probes whether external knowledge retrieval improves classification of ambiguous or low-frequency entities compared to context-only baselines. Use when the user wants to benchmark on MultiCoNER2, or asks about evaluating this task. Reports macro-F1.
- ▌ Multidialog Eval · qhjqhj00This benchmark evaluates the semantic coherence, acoustic fidelity, and audio-visual synchronization of end-to-end spoken dialogue systems that generate face-to-face conversational audio and video. It probes a model's ability to maintain contextually appropriate dialogue while producing synchronized multimodal outputs without relying on intermediate text representations. Use when the user wants to benchmark on MultiDialog, or asks about evaluating this task. Reports PPL.
- ▌ Multifinben Eval · qhjqhj00Evaluates large language models on financial reasoning, comprehension, and generation across multiple modalities (text, vision, audio), languages (English, Chinese, Japanese, Spanish, Greek), and task types (IE, QA, summarization, etc.), using a difficulty-aware selection framework to ensure balanced and discriminative assessment. Use when the user wants to benchmark on IESC, FinRED, FINER-ORD, Headlines, TATSA, BRL-Math, FinQA, TATQA, CECTSUM, TGEDTSUM, RMCCF, BigData22, MDSFT, RRE, AIE, LNE, FinanceIQ, chabsa, MultiFin, EFPA, FNS-2023, GRFinNUM, GRMultiFin, GRFinQA, GRFNS-2023, DOLFIN, PolyFiQA-Easy, PolyFiQA-Expert, EnglishOCR, JapaneseOCR, SpanishOCR, GreekOCR, TableBench, MDRM-test, FinAudioSum, or asks about evaluating this task. Reports Accuracy.
- ▌ Multilabellogauc · qhjqhj00Compute the MultilabelLogAUC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelLogAUC, or asks how to score with MultilabelLogAUC.
- ▌ Multilabelrecall · qhjqhj00Compute the MultilabelRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelRecall, or asks how to score with MultilabelRecall.
- ▌ Multilexsum Eval · qhjqhj00Evaluates abstractive and extractive summarization models on real-world civil rights lawsuits, testing their ability to synthesize information from extremely long multi-document sources and generate summaries at three distinct length granularities (long, short, tiny). Use when the user wants to benchmark on Multi-LexSum, or asks about evaluating this task. Reports ROUGE-2 F1.
- ▌ Multimed St Eval · qhjqhj00Evaluates the capability of speech translation models to accurately convert medical speech across five languages (English, Vietnamese, German, French, Mandarin Chinese) into text. It probes both end-to-end and cascaded architectures, as well as the impact of multilingual vs. bilingual training and code-switching handling in a specialized medical domain. Use when the user wants to benchmark on MultiMed-ST, or asks about evaluating this task. Reports BLEU, BERTScore.
- ▌ Multisports Eval · qhjqhj00Evaluates multi-person spatio-temporal action detection in sports videos, probing the model's ability to localize fine-grained actions across multiple concurrent persons, handle occlusion, and model long-range temporal context. Use when the user wants to benchmark on MultiSports, or asks about evaluating this task. Reports frame-mAP@0.5.
- ▌ Multitaskwrapper · qhjqhj00Compute the MultitaskWrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultitaskWrapper, or asks how to score with MultitaskWrapper.
- ▌ Multiwoz2 1 Eval · qhjqhj00Evaluates a model's ability to track dialogue state (domain, slot, value triplets) across conversation turns, specifically probing its robustness to user mind-changes or 'turnback' utterances that modify previously stated intentions. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports joint goal accuracy.
- ▌ Musciclaims Eval · qhjqhj00Multimodal scientific claim verification, requiring models to read complex figures and captions to determine if a scientific claim is supported, neutral, or contradicted. It also probes evidence localization, basic visual understanding, cross-modal aggregation, and epistemic sensitivity. Use when the user wants to benchmark on MuSciClaims, or asks about evaluating this task. Reports F1 score.
- ▌ Musdb18 Sdr Eval · qhjqhj00Evaluates the quality of separated audio sources (vocals, drums, bass, other) from mixed music tracks using source-to-distortion ratio, probing the model's ability to perform supervised music source separation. Use when the user wants to benchmark on MUSDB18, or asks about evaluating this task. Reports SDR.
- ▌ Mvpaint T2t Eval · qhjqhj00This evaluation probes a model's ability to generate high-quality, multi-view consistent 3D textures on arbitrary meshes conditioned on text instructions. It measures visual fidelity, distributional similarity to ground truth, and cross-view consistency through both automated generative metrics and human preference studies. Use when the user wants to benchmark on Objaverse T2T benchmark, GSO T2T benchmark, or asks about evaluating this task. Reports FID.
- ▌ Narrativeqa Eval · qhjqhj00Evaluates English long-document retrieval on complex, narrative-style questions, probing deep comprehension and information extraction from lengthy texts. Use when the user wants to benchmark on NarrativeQA, or asks about evaluating this task. Reports nDCG@10.
- ▌ Needlebench Eval · qhjqhj00Evaluates large language models' ability to retrieve specific information and perform complex multi-point reasoning within long-context documents. It probes both information-sparse retrieval and information-dense reasoning (Ancestral Trace Challenge) across 32K and 128K token contexts. Use when the user wants to benchmark on NeedleBench, or asks about evaluating this task. Reports Overall.
- ▌ Neucir 2023 Eval · qhjqhj00Evaluates neural cross-language and multilingual information retrieval systems on news collections in Chinese, Persian, and Russian. It probes a model's ability to rank relevant documents when queries are in English and documents are in different languages, as well as its capacity to unify rankings across multiple languages. Use when the user wants to benchmark on NeuCLIR 2023, or asks about evaluating this task. Reports nDCG@20.
- ▌ Neural Stpp Eval · qhjqhj00Evaluates the ability of spatio-temporal point process models to accurately capture complex, history-dependent spatial and temporal distributions of discrete events. It probes how well models can compute exact likelihoods for sequences of events in continuous space and time across diverse domains like seismology, epidemiology, and neuroscience. Use when the user wants to benchmark on PINWHEEL, EARTHQUAKES, COVID-19 CASES, BOLD5000, or asks about evaluating this task. Reports log-likelihood per event.
- ▌ Nlu Service Eval · qhjqhj00Evaluates commercial and open-source NLU platforms on intent classification and named entity recognition across multiple dialogue domains, highlighting limitations in multi-intent support and contextual modeling. Use when the user wants to benchmark on NLU Evaluation Dataset, or asks about evaluating this task. Reports Intent classification accuracy, Entity recognition precision.
- ▌ Omniscience Eval · qhjqhj00Evaluates the semantic alignment and factual fidelity of automatically generated scientific image captions. It probes whether dense, context-aware captions can replace visual inputs for downstream reasoning tasks and how well they capture complex scientific figures compared to raw human-written captions. Use when the user wants to benchmark on OmniScience, AI2D, MMMU, MM-MT-Bench, MSEarth, or asks about evaluating this task. Reports cross-modal relevance score.
- ▌ Omnispatial Eval · qhjqhj00This benchmark evaluates vision-language models on advanced spatial reasoning capabilities, specifically probing dynamic reasoning, complex spatial logic, spatial interaction, and perspective-taking. It measures how well models understand and manipulate spatial relationships, temporal changes, and viewpoint shifts beyond basic object recognition. Use when the user wants to benchmark on OmniSpatial, or asks about evaluating this task. Reports accuracy.
- ▌ Omnispectra Eval · qhjqhj00Evaluates a foundation model's ability to process native-resolution astronomical spectra of variable lengths without resampling, and assesses its zero-shot, few-shot, and supervised performance on stellar property estimation and source classification tasks across diverse spectroscopic surveys. Use when the user wants to benchmark on OmniSpectra Multi-Survey Corpus, or asks about evaluating this task. Reports Mean-Squared Error.
- ▌ Onechart Se Eval · qhjqhj00Evaluates a model's ability to extract structured information from chart images, including textual OCR accuracy and precise numerical value parsing. It tests the model's capacity to convert visual chart elements into a standardized Python-dict representation, handling both annotated and unannotated charts across multiple languages and rendering styles. Use when the user wants to benchmark on ChartQA-SE, PlotQA-SE, ChartX-SE, ChartY-en, ChartY-zh, or asks about evaluating this task. Reports SCRM AP (strict/slight/high).
- ▌ Ood Mol Opt Eval · qhjqhj00Out-of-domain molecular property prediction and Bayesian optimization for molecular design. Tests transferability of learned representations to novel tasks. Use when the user wants to benchmark on Out-of-domain molecular design tasks, or asks about evaluating this task. Reports Top performing molecule property.
- ▌ Openfwi Fwi Eval · qhjqhj00Evaluates deep learning models for seismic full-waveform inversion (FWI) by predicting subsurface velocity models from seismic wavefield data. It probes the model's ability to generalize across varying geological complexities and out-of-distribution scenarios using parameter-efficient fine-tuning. Use when the user wants to benchmark on OpenFWI, or asks about evaluating this task. Reports SSIM.
- ▌ Openml Cc18 Eval · qhjqhj00Evaluates machine learning classifiers on a curated collection of standardized classification tasks. It probes the reproducibility and comparability of algorithm performance across diverse datasets under consistent, machine-readable evaluation protocols. Use when the user wants to benchmark on OpenML-CC18, or asks about evaluating this task. Reports accuracy_score.
- ▌ Openner 1 0 Eval · qhjqhj00Evaluates named entity recognition (NER) capabilities across 52 languages and 36 distinct corpora. It probes cross-lingual generalization, robustness to varying entity type ontologies, and the ability of both encoder-based models and LLMs to handle multilingual text with diverse annotation guidelines. Use when the user wants to benchmark on OpenNER 1.0, or asks about evaluating this task. Reports micro-averaged mention-level F1.
- ▌ Pas Dataset Eval · qhjqhj00Evaluates the transferability and pretraining quality of vision models trained on synthetic domain-specific datasets compared to manually curated and general-domain datasets. It probes the model's ability to generalize to fine-grained classification and object detection tasks within specific domains like birds and food. Use when the user wants to benchmark on CUB-200-2011, NABirds, iNatbirds, Food-101, FoodX-251, Food-2K, or asks about evaluating this task. Reports Top-1 k-NN accuracy.
- ▌ Pdbbind Lba Eval · qhjqhj00Predicts the binding affinity between a protein pocket and a ligand from 3D structural data. It probes the model's ability to quantify molecular interaction strength and generalize across protein sequence identities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports RMSE.
- ▌ Perrecbench Eval · qhjqhj00Evaluates whether LLMs can capture true personalized user preferences by ranking items or users in groups, while explicitly controlling for confounding factors like user rating bias and item quality. It probes the model's ability to perform comparative reasoning rather than simple rating prediction. Use when the user wants to benchmark on PerRecBench, or asks about evaluating this task. Reports Kendall’s tau.
- ▌ Persian Ner Eval · qhjqhj00Evaluates the quality and cross-lingual transferability of machine-translated Persian named entity recognition datasets by measuring model performance against original English benchmarks. It probes how well translation-based dataset generation preserves entity boundaries and labels across languages with different scripts and linguistic structures. Use when the user wants to benchmark on CoNLL 2003, OntoNotes 5.0, NCBI Disease, WNUT 2017, or asks about evaluating this task. Reports F1.
- ▌ Persian RAG Eval · qhjqhj00Evaluates retrieval-augmented generation (RAG) pipelines for Persian text across general, scientific, and formal domains. It probes the ability of sentence embeddings to retrieve relevant context and large language models to generate accurate, faithful, and relevant answers based on that context. Use when the user wants to benchmark on PQuad, Scientific-Specialized, Organizational Report, or asks about evaluating this task. Reports Context Recall.
- ▌ Phi3 Safety Eval · qhjqhj00Evaluates the safety and refusal capabilities of language models across multiple risk categories including harmful content generation, jailbreaking, stereotype bias, privacy leaks, and toxicity detection. It measures how well models balance harmlessness (refusing unsafe prompts) and helpfulness (complying with safe prompts) in both single- and multi-turn interactions. Use when the user wants to benchmark on XSTest, DecodingTrust, ToxiGen, XSafety, RTP-LX, Microsoft Internal Automated Measurement, or asks about evaluating this task. Reports IPRR.
- ▌ Phishnchips Eval · qhjqhj00Evaluates the security and robustness of autonomous LLM email agents against phishing attacks by measuring how different system prompt configurations affect detection sensitivity and operational false positive rates. It specifically probes the model's ability to maintain high recall while minimizing usability costs, and tests adversarial brittleness under infrastructure phishing conditions where attacker-controlled domains match sender addresses. Use when the user wants to benchmark on Synthetic Email Phishing Corpus, or asks about evaluating this task. Reports Net Effectiveness (Recall-FPR).
- ▌ Phreshphish Eval · qhjqhj00Evaluates phishing website detection models on temporally disjoint, real-world data with realistic base rates, while mitigating training-to-test leakage and varying difficulty levels. Use when the user wants to benchmark on PhreshPhish, or asks about evaluating this task. Reports Precision-Recall.
- ▌ Pii Masking Eval · qhjqhj00Evaluates the ability of PII masking models to correctly identify and classify sensitive information in text. It probes performance across diverse contexts, multilingual inputs, noisy formats, and evolving entity types. Use when the user wants to benchmark on PII Masking Dataset, or asks about evaluating this task. Reports non-identification.
- ▌ Pii Tagging Eval · qhjqhj00Evaluates a model's ability to identify and extract private or sensitive information spans from text across legal, healthcare, and finance domains. It probes the model's capacity for fine-grained named entity recognition under limited labeled data and domain-specific privacy definitions. Use when the user wants to benchmark on ECHR, MACCROBAT, PUPA (Finance Subset), or asks about evaluating this task. Reports F1.
- ▌ Plaba Track Eval · qhjqhj00Evaluates NLP systems and large language models on adapting biomedical abstracts to plain language for lay consumers. It probes capabilities in text simplification, term replacement, factual faithfulness, and conciseness while measuring alignment with human expert judgments. Use when the user wants to benchmark on TREC PLABA, or asks about evaluating this task. Reports SARI.
- ▌ Planetarium Eval · qhjqhj00Evaluates an LLM's ability to translate natural language planning task descriptions into valid, semantically equivalent Planning Domain Definition Language (PDDL) code. It specifically probes the model's capacity to accurately capture initial states, goal states, and object relationships while adhering to formal planning semantics. Use when the user wants to benchmark on Planetarium, or asks about evaluating this task. Reports equivalence.
- ▌ Pmindia Nmt Eval · qhjqhj00This evaluation probes the quality of automatic machine translation between English and 13 Indian languages using a parallel corpus. It measures how well NMT systems can handle diverse linguistic structures, including abugida scripts and agglutinative morphology, across low-resource language pairs. Use when the user wants to benchmark on PMIndia, or asks about evaluating this task. Reports BLEU.
- ▌ Polychartqa Eval · qhjqhj00Evaluates multimodal language models' ability to answer questions that require reasoning across multiple charts or images. It probes visual decomposition, sub-chart localization, and handling of complex multi-visual contexts versus single-chart inputs. Use when the user wants to benchmark on PolyChartQA, MultiChartQA-RQ1, or asks about evaluating this task. Reports L-Accuracy.
- ▌ Pope Nocaps Eval · qhjqhj00Tests object perception and hallucination on images without captions, evaluating whether LVLMs can ground object detection purely from visual input without textual priors. Use when the user wants to benchmark on POPE-NoCaps, or asks about evaluating this task. Reports Acc.
- ▌ Power Divergence · qhjqhj00Compute the power_divergence metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute power_divergence, or asks how to score with power_divergence.
- ▌ Privacylens Eval · qhjqhj00Assesses an LLM agent's ability to understand and follow privacy norms while performing real-world tasks. It measures both helpfulness and the rate at which sensitive information is incorrectly exposed. Use when the user wants to benchmark on PrivacyLens, or asks about evaluating this task. Reports privacy leakage rate.
- ▌ Progressftx Eval · qhjqhj00Evaluates a progressive feature transmission protocol for split inference at the wireless edge, measuring how efficiently features are transmitted to meet target inference accuracy or uncertainty thresholds under varying channel conditions. Use when the user wants to benchmark on GM dataset, MNIST, or asks about evaluating this task. Reports average communication latency.
- ▌ Propsegment Eval · qhjqhj00Evaluates a model's ability to decompose sentences into atomic semantic units (propositional segmentation) and determine entailment relationships between text spans. It probes fine-grained compositional semantic alignment and partial entailment recognition beyond sentence-level NLI. Use when the user wants to benchmark on PropSegmEnt, or asks about evaluating this task. Reports Precision/Recall/F1w (macro-averaged).
- ▌ Pulseimpute Eval · qhjqhj00Evaluates models on imputing missing values in pulsative physiological signals (ECG and PPG) under realistic, data-driven missingness patterns. It further assesses clinical utility by measuring downstream performance on heartbeat detection and cardiac classification tasks. Use when the user wants to benchmark on ECG, PPG, or asks about evaluating this task. Reports MSE, F1 Score.
- ▌ Pushupbench Eval · qhjqhj00Evaluates video-language models on long-form repetition counting and temporal reasoning. It probes whether models can accurately track state changes and count actions across extended video clips, revealing weaknesses in spatio-temporal tracking compared to supervised baselines. Use when the user wants to benchmark on PushupBench, or asks about evaluating this task. Reports Exact Match.
- ▌ Pyvision Rl Eval · qhjqhj00Evaluates open-weight multimodal agentic models on visual search, multimodal mathematical reasoning, multi-turn tool use, and video spatial reasoning. It probes the model's ability to dynamically construct context, invoke tools, and perform long-horizon reasoning with high visual token efficiency. Use when the user wants to benchmark on V*, HRBench-4K, HRBench-8K, MathVerse, MathVision, WeMath, DynaMath, TIR-Bench, VSI-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Qianfan Ocr Eval · qhjqhj00This evaluation probes a unified vision-language model's ability to perform end-to-end document intelligence, including specialized OCR, general text recognition, document understanding, and key information extraction across diverse document types and multilingual scenarios. Use when the user wants to benchmark on Omni-Doc-Bench v1.5, OLMOCRBench, OCRBench, DocVQA, ChartQA, Nanonets KIE, or asks about evaluating this task. Reports normalized accuracy (0-100).
- ▌ Qualispeech Eval · qhjqhj00Evaluates auditory large language models' ability to perceive and describe low-level speech quality aspects, including noise, distortion, speed, continuity, listening effort, naturalness, and overall quality. It probes both numerical score prediction and natural language reasoning/description generation. Use when the user wants to benchmark on QualiSpeech, or asks about evaluating this task. Reports PCC.
- ▌ Quito Bench Eval · qhjqhj00Evaluates time series forecasting models on a regime-balanced benchmark stratified by trend, seasonality, and forecastability. It probes a model's ability to handle varying context lengths, forecast horizons, and multivariate dependencies while mitigating domain bias and information leakage. Use when the user wants to benchmark on QuitoBench, or asks about evaluating this task. Reports MAE.
- ▌ Radioactive Eval · qhjqhj00Evaluates interactive 3D medical image segmentation by measuring how well models segment target structures under varying human-in-the-loop prompting strategies (points, boxes, scribbles) and iterative refinement protocols. It specifically probes the trade-off between interaction effort and segmentation accuracy across 2D and 3D architectures. Use when the user wants to benchmark on RadioActive, or asks about evaluating this task. Reports Dice.
- ▌ RAG Medical Eval · qhjqhj00Evaluates the effectiveness and efficiency of Retrieval-Augmented Generation (RAG) systems across medical and general knowledge domains. It probes how different RAG pipeline components (chunking, indexing, query classification, augmentation, and prompting) impact answer accuracy and response latency on question-answering and information extraction tasks. Use when the user wants to benchmark on MMLU, PubMedQA, PromptNER, Query Classification Dataset, or asks about evaluating this task. Reports accuracy (acc).
- ▌ Ragen Agent Eval · qhjqhj00Evaluates LLM agents' multi-turn decision-making and reasoning capabilities across symbolic planning, risk-sensitive reasoning, and realistic web interaction environments. It probes the agent's ability to complete interactive tasks under noisy or probabilistic feedback while maintaining exploration and training stability. Use when the user wants to benchmark on Bandit, Sokoban, Frozen Lake, WebShop, or asks about evaluating this task. Reports success rate.
- ▌ Randumb Ocl Eval · qhjqhj00Evaluates continual learning methods in online, exemplar-free, and low-exemplar regimes by measuring how well a model retains knowledge of previously seen classes after processing a single pass of sequential data. It specifically tests whether fixed random representations can match or exceed learned representations in these constrained settings. Use when the user wants to benchmark on MNIST, CIFAR10, CIFAR100, TinyImageNet200, miniImageNet100, or asks about evaluating this task. Reports average_accuracy.
- ▌ Recruitview Eval · qhjqhj00This benchmark evaluates multimodal models on predicting continuous personality traits and interview performance scores from video, audio, and text inputs. It probes the model's ability to perform fine-grained behavioral analysis and regression across psychometric targets. Use when the user wants to benchmark on RecruitView, or asks about evaluating this task. Reports Spearman's ρ.
- ▌ Repobench C Eval · qhjqhj00Evaluates autoregressive language models on predicting the next line of code using provided in-file and cross-file contexts. Use when the user wants to benchmark on RepoBench-C, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Repobench P Eval · qhjqhj00Evaluates an end-to-end pipeline that first retrieves cross-file snippets and then predicts the next line of code using both the in-file context and retrieved snippets. Use when the user wants to benchmark on RepoBench-P, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Repobench R Eval · qhjqhj00Evaluates a model's ability to retrieve relevant cross-file code snippets given an in-file context for predicting the next line of code. Use when the user wants to benchmark on RepoBench-R, or asks about evaluating this task. Reports acc@1.
- ▌ Repro Bench Eval · qhjqhj00Evaluates whether agentic AI systems can accurately assess the computational reproducibility of social science research by comparing original paper findings against results reproduced from provided raw data and code. It probes end-to-end agentic reasoning, including command execution, debugging, and result interpretation in a simulated research environment. Use when the user wants to benchmark on REPRO-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Researchgym Eval · qhjqhj00Evaluates the capability of LLM agents to conduct closed-loop scientific research by proposing hypotheses, executing experiments, and outperforming human baselines on repurposed real-world AI papers. It probes long-horizon planning, resource management, and autonomous experimentation under realistic tool constraints. Use when the user wants to benchmark on ResearchGym, or asks about evaluating this task. Reports improvement over baselines.
- ▌ Respondeoqa Eval · qhjqhj00This benchmark evaluates large language models on bilingual Latin-English question answering across knowledge-based, skill-based (grammar, scansion, literary devices), multihop reasoning, and translation tasks. It probes models' ability to handle classical language morphology, poetic meter analysis, and cross-lingual generation under constrained and unconstrained settings. Use when the user wants to benchmark on RespondeoQA, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Retrievalfallout · qhjqhj00Compute the RetrievalFallOut metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalFallOut, or asks how to score with RetrievalFallOut.
- ▌ Retrievalhitrate · qhjqhj00Compute the RetrievalHitRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalHitRate, or asks how to score with RetrievalHitRate.
- ▌ Retrotrucks Eval · qhjqhj00Evaluates video anomaly detection models on dashcam footage, specifically testing their ability to detect complex traffic anomalies like collisions and skidding in dynamic, real-world driving scenes. It also benchmarks performance against standard pedestrian anomaly detection datasets to highlight challenges posed by moving cameras and contextual anomalies. Use when the user wants to benchmark on RetroTrucks, UCSD Ped1, UCSD Ped2, ShanghaiTech, or asks about evaluating this task. Reports AUC-ROC.
- ▌ Reviewbench Eval · qhjqhj00Evaluates the quality and accuracy of automated peer reviews generated by LLMs. It probes both the semantic alignment of review text against paper-specific rubrics and the precision of predicted numerical ratings and acceptance decisions. Use when the user wants to benchmark on ReviewBench, or asks about evaluating this task. Reports Rubric Overall Score.
- ▌ Rextthewild Eval · qhjqhj00Evaluates multimodal models' ability to understand real-world medical photographs by answering clinician-verified multiple-choice questions across seven clinical domains. It probes capabilities in geometric perception, anatomical localization, clinical characterization, and causal reasoning. Use when the user wants to benchmark on ReXInTheWild, or asks about evaluating this task. Reports accuracy.
- ▌ Rgbt Ground Eval · qhjqhj00Evaluates multi-modal visual grounding capabilities by requiring models to localize objects in images using both RGB and thermal infrared (TIR) modalities guided by text queries. It specifically probes robustness under complex real-world conditions such as low-light environments, small object sizes, and diverse weather/illumination variations. Use when the user wants to benchmark on RGBT-Ground, or asks about evaluating this task. Reports Acc@0.5.
- ▌ Rgt Seismic Eval · qhjqhj00Evaluates a model's ability to perform continuous regression for Relative Geologic Time (RGT) estimation from 2D seismic images. It probes the model's capacity to learn stratigraphic continuity and structural consistency across diverse geological settings, testing generalization from synthetic labeled data to unlabeled real-world field data. Use when the user wants to benchmark on Field Seismic Dataset, Synthetic Seismic Dataset, or asks about evaluating this task. Reports regression.
- ▌ Safe Speech Eval · qhjqhj00Evaluates the capability of classifiers to detect sexist, abusive, offensive, and hate speech in conversational text across multiple granularity levels and established benchmarks. It probes fine-grained toxicity detection, cross-dataset generalization, and performance against strong supervised and LLM baselines. Use when the user wants to benchmark on EDOS (SemEval 2023), OffensEval 2019, AbusEval, HatEval, or asks about evaluating this task. Reports F1.
- ▌ Safeprotein Eval · qhjqhj00Evaluates the biosafety risks and jailbreak vulnerabilities of protein foundation models by measuring their ability to reconstruct harmful protein sequences and 3D structures from partially masked inputs. It probes whether models can bypass safety filters and generate biologically dangerous proteins when given sequence and structural prompts. Use when the user wants to benchmark on SafeProtein-Bench, or asks about evaluating this task. Reports jailbreak success rate.
- ▌ Safetunebed Eval · qhjqhj00Evaluates the safety alignment preservation and task utility of LLMs after parameter-efficient fine-tuning under data-poisoning attacks. It measures how well defenses maintain core capabilities while resisting harmful behavior injection. Use when the user wants to benchmark on MMLU, MT-Bench, AdvBench, PolicyEval, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Sagalee Asr Eval · qhjqhj00Evaluates automatic speech recognition (ASR) performance on the Oromo language using real-world, crowd-sourced audio data. It measures how well different model architectures (Conformer trained from scratch, Whisper fine-tuned) transcribe spoken Oromo into text under varying acoustic conditions. Use when the user wants to benchmark on Sagalee, or asks about evaluating this task. Reports WER.
- ▌ Sailcompass Eval · qhjqhj00Evaluates large language models on language proficiency, reading comprehension, reasoning, and cultural understanding across three Southeast Asian languages (Indonesian, Vietnamese, Thai). It covers eight diverse tasks including question answering, machine translation, text summarization, multiple-choice exams, commonsense reasoning, machine reading comprehension, natural language inference, and sentiment analysis. Use when the user wants to benchmark on XQuAD, TyDiQA, Flores-200, ThaiSum, IndoSum, XLSUM, M3Exam, XCOPA, BELEBELE, XNLI, IndoNLI, Wisesight, Indolem, VSMEC, or asks about evaluating this task. Reports Exact Match.
- ▌ Salad Bench Eval · qhjqhj00Evaluates the safety, robustness, and helpfulness of Large Language Models across a hierarchical taxonomy of 6 domains, 16 tasks, and 66 categories. It measures performance on benign, adversarial (attack-enhanced), and defense-enhanced prompts, as well as multiple-choice safety questions, while also benchmarking the effectiveness of various attack and defense strategies. Use when the user wants to benchmark on SALAD-Bench, ToxicChat, Beavertails, SafeRLHF, Harmbench, Lifetox, AdvBench-50, or asks about evaluating this task. Reports Safety Rate, Attack Success Rate (ASR).
- ▌ Sarena Icon Eval · qhjqhj00Evaluates a model's ability to generate scalable vector graphics (SVG) from text prompts and reference images, measuring visual fidelity, semantic alignment, structural success, and code efficiency. Use when the user wants to benchmark on SArena-Icon, or asks about evaluating this task. Reports SR.
- ▌ Scatspotter Eval · qhjqhj00Evaluates object detection and instance segmentation capabilities on real-world images of dog feces. It specifically probes model robustness to camouflage, occlusion, varying lighting conditions, and small object detection in outdoor urban environments. Use when the user wants to benchmark on ScatSpotter, or asks about evaluating this task. Reports mAP.
- ▌ Scene Bench Eval · qhjqhj00Evaluates the factual consistency and scene graph adherence of text-to-image generation models. It probes whether generated images accurately preserve specified objects and their spatial/relational configurations as defined by input scene graphs, rather than just measuring aesthetic quality or text-image alignment. Use when the user wants to benchmark on Visual Genome (VG) test set, MegaSG, or asks about evaluating this task. Reports SGScore.
- ▌ Scene Smith Eval · qhjqhj00Evaluates text-to-3D indoor scene generation systems on their ability to produce dense, physically plausible, and prompt-faithful environments. It probes both visual realism and simulation-readiness, measuring collision-free layouts and stable physics properties required for robotics policy testing. Use when the user wants to benchmark on SceneSmith Prompt Corpus, or asks about evaluating this task. Reports Realism Win%.
- ▌ Scenicrules Eval · qhjqhj00Evaluates autonomous driving agents on their ability to navigate stochastic traffic scenarios while satisfying a hierarchical set of multi-objective specifications. It probes how well agents balance conflicting goals like collision avoidance, road compliance, passenger comfort, and progress under varying priority constraints. Use when the user wants to benchmark on ScenicRules Benchmark, or asks about evaluating this task. Reports Violation Score (VS).
- ▌ Scholawrite Eval · qhjqhj00Evaluates an LLM's ability to predict human scholarly writing intentions from a LaTeX draft and to iteratively edit the draft according to those intentions. It measures lexical diversity, topic consistency, and intention coverage across a 100-iteration self-writing process. Use when the user wants to benchmark on SCHOLAWRITE, or asks about evaluating this task. Reports intention coverage.
- ▌ Scigenbench Eval · qhjqhj00Evaluates the logical correctness, structural fidelity, and information utility of AI-generated scientific images. It probes whether generated visuals accurately encode domain-specific facts and geometric relationships, and whether they are indispensable for solving visually grounded scientific quizzes. Use when the user wants to benchmark on SciGenBench, or asks about evaluating this task. Reports inverse_validation_rate.
- ▌ Screen Spot Eval · qhjqhj00This evaluation probes a vision-language model's ability to localize specific UI elements within graphical user interfaces based on natural language instructions. It tests precise coordinate prediction and cross-resolution generalization across mobile, desktop, and web platforms. Use when the user wants to benchmark on ScreenSpot, ScreenSpot-v2, ScreenSpot-Pro, or asks about evaluating this task. Reports accuracy.
- ▌ Se Toxicity Eval · qhjqhj00Evaluates the ability of contemporary toxicity detection models to correctly identify toxic language in software engineering contexts, such as code reviews and developer chat logs. It probes whether general-purpose classifiers can handle domain-specific terminology and contextual nuances without significant performance degradation. Use when the user wants to benchmark on Jigsaw Sample, Code Review, Gitter Ethereum, or asks about evaluating this task. Reports F-Score.
- ▌ Seas Safety Eval · qhjqhj00This evaluation probes the safety alignment and refusal capabilities of LLMs when exposed to harmful or adversarial prompts. It measures the frequency of unsafe model outputs to quantify vulnerability, while simultaneously tracking general instruction-following scores to ensure that safety hardening does not degrade overall utility. Use when the user wants to benchmark on SEAS-Test, BeaverTrail, HH-RLHF, XSTest, or asks about evaluating this task. Reports Attack Success Rate (ASR).
- ▌ Secretbench Eval · qhjqhj00Evaluates the capability of automated secret detection tools to accurately identify hardcoded secrets (e.g., API keys, passwords, private keys) in source code repositories. It probes the tools' ability to balance high recall for true secrets against low false positive rates to mitigate alert fatigue. Use when the user wants to benchmark on SecretBench, or asks about evaluating this task. Reports Precision.
- ▌ Sega Layout Eval · qhjqhj00Evaluates a model's ability to generate content-aware graphic layouts from background images and instructions. It probes spatial reasoning, adherence to design principles (alignment, overlap, occlusion), and aesthetic quality. Use when the user wants to benchmark on PKU, CGL, Crello, or asks about evaluating this task. Reports Ali.
- ▌ Semantic Kg Eval · qhjqhj00Evaluates the ability of semantic similarity methods to correctly classify pairs of natural language statements as semantically similar (label 1) or dissimilar (label 0). It specifically probes how well models handle controlled semantic variations (node and edge perturbations) across general and domain-specific knowledge domains. Use when the user wants to benchmark on Semantic-KG Benchmark, or asks about evaluating this task. Reports F1-score.