qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Ultralink Eval · qhjqhj00Evaluates multilingual LLMs on chat, math reasoning, and code generation across five languages (English, Chinese, Spanish, Russian, French) to measure the effectiveness of knowledge-enhanced supervised fine-tuning. Use when the user wants to benchmark on OMGEval, MGSM, Multilingual HumanEval, or asks about evaluating this task. Reports OMGEval score.
- ▌ Univbench Eval · qhjqhj00Evaluates video foundation models across six core tasks (understanding, generation, editing, reconstruction) by scoring their outputs on eight cinematic and semantic dimensions, including subject consistency, action dynamics, camera movement, and lighting. Use when the user wants to benchmark on UniVBench, or asks about evaluating this task. Reports UniV-Eval.
- ▌ Univearth Eval · qhjqhj00This benchmark probes an LLM agent's ability to perform spatial-temporal reasoning and generate executable code for Earth Observation tasks. It evaluates whether models can correctly answer yes/no questions derived from scientific articles by leveraging remote sensing data via Google Earth Engine. Use when the user wants to benchmark on UnivEARTH, or asks about evaluating this task. Reports accuracy.
- ▌ Unrealzoo Eval · qhjqhj00Evaluates embodied AI agents' capabilities in complex, photo-realistic 3D open-world environments. Specifically probes visual navigation on unstructured terrain, active visual tracking across diverse scenes, and social tracking under dynamic distractions, varying morphologies, and different control frequencies. Use when the user wants to benchmark on UnrealZoo, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Unsw Nb15 Eval · qhjqhj00Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on UNSW-NB15, or asks about evaluating this task. Reports accuracy.
- ▌ Urdu Mner Eval · qhjqhj00Evaluates the ability of models to recognize named entities (Person, Location, Organization, Miscellaneous) in Urdu social media posts by jointly processing textual and visual inputs. It probes cross-modal alignment, handling of low-resource language morphological complexity, and ambiguity resolution using visual context. Use when the user wants to benchmark on Twitter2015-Urdu, or asks about evaluating this task. Reports F1 score.
- ▌ V2x Radar Eval · qhjqhj00Evaluates 3D object detection capabilities for autonomous driving using multi-modal sensors (LiDAR, camera, 4D radar) in single-agent (roadside and vehicle-mounted) and cooperative perception setups. It probes robustness to adverse weather conditions and communication delays in cooperative scenarios. Use when the user wants to benchmark on V2X-Radar, or asks about evaluating this task. Reports AP@IoU.
- ▌ Vaexbench Eval · qhjqhj00Evaluates multimodal large language models' ability to perform extractive and abstractive spatiotemporal reasoning on egocentric videos. It probes long-horizon memory, object tracking, spatial orientation, and metric distance estimation under both multiple-choice and free-form generation settings. Use when the user wants to benchmark on VAEX-Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Valerie22 Eval · qhjqhj00This protocol evaluates the perceptual fidelity and cross-domain generalization capability of the VALERIE22 synthetic urban dataset by training a semantic segmentation model on it and testing on real-world automotive datasets. It specifically probes how dataset diversity (unique 3D assets) and training scale affect downstream perception performance. Use when the user wants to benchmark on VALERIE22, Cityscapes, A2D2, BDD100K, India Driving Dataset, Mapillary Vistas, or asks about evaluating this task. Reports mIoU.
- ▌ Vc Ifeval Eval · qhjqhj00Evaluates multimodal large language models' ability to follow instructions containing vision-dependent constraints, such as spatial, stylistic, and structural requirements. It isolates the contribution of visual input to instruction adherence and assesses generalization on standard visual reasoning tasks. Use when the user wants to benchmark on VC-IFEval, MM-IFEval, IFEval, or asks about evaluating this task. Reports instruction-following accuracy.
- ▌ Vcb Bench Eval · qhjqhj00VCB Bench evaluates audio-grounded large language models on instruction following with speech-level controls, knowledge reasoning, and robustness under real-world acoustic perturbations. It probes how well models understand and generate spoken responses in Chinese and English using authentic human speech rather than synthetic data. Use when the user wants to benchmark on VCB Bench, or asks about evaluating this task. Reports 1-5 scale score.
- ▌ Vibe Eval Eval · qhjqhj00Probes multimodal reasoning and visual understanding on real-world images. Specifically designed with a 'hard' subset of prompts that are unsolvable by current frontier models to measure genuine performance gaps and contamination-free generalization. Use when the user wants to benchmark on Vibe-Eval, or asks about evaluating this task. Reports Vibe-Eval Score.
- ▌ Video Mme Eval · qhjqhj00Evaluates long-horizon video understanding and reasoning capabilities of multimodal models on extended video sequences. Use when the user wants to benchmark on Video-MME, or asks about evaluating this task. Reports accuracy.
- ▌ Videocube Eval · qhjqhj00Evaluates a model's ability to track arbitrary visual instances across complex, unstructured real-world videos without assuming motion continuity or fixed categories. It measures both local search accuracy and global robustness against challenges like occlusion, fast motion, and scene transitions. Use when the user wants to benchmark on VideoCube, or asks about evaluating this task. Reports PRE.
- ▌ Vidore V3 Eval · qhjqhj00This benchmark evaluates end-to-end Retrieval Augmented Generation (RAG) systems on visually rich, real-world documents across multiple professional domains. It probes a model's ability to retrieve relevant pages, generate accurate answers to complex open-ended and multi-hop queries, and precisely ground those answers with bounding boxes in multimodal content. Use when the user wants to benchmark on ViDoRe V3, or asks about evaluating this task. Reports F1 score (Dice coefficient).
- ▌ Vimed Pet Eval · qhjqhj00Evaluates vision-language models on generating Vietnamese clinical reports from paired PET/CT images and answering medical questions about them. Probes the model's ability to align 3D medical imaging features with low-resource language text and produce clinically accurate descriptions. Use when the user wants to benchmark on ViMed-PET, or asks about evaluating this task. Reports BLEU-4.
- ▌ Vision R1 Eval · qhjqhj00Evaluates Large Vision-Language Models on their ability to detect, localize, and ground objects in images across diverse and challenging scenarios, including in-domain dense detection, out-of-domain real-world settings, and generalization to unseen categories or scenes. Use when the user wants to benchmark on MSCOCO Val2017, ODINW-13, or asks about evaluating this task. Reports mAP.
- ▌ Visonlyqa Eval · qhjqhj00This benchmark probes a model's ability to accurately perceive basic geometric information—such as shape, angle, length, area, and intersections—in scientific figures and diagrams. It isolates visual perception from higher-level reasoning or domain knowledge by using direct, low-reasoning questions on synthetic and real-world images. Use when the user wants to benchmark on VisOnlyQA, or asks about evaluating this task. Reports accuracy.
- ▌ Visulogic Eval · qhjqhj00VisuLogic probes vision-centric reasoning in multimodal large language models by presenting problems that require retaining critical visual cues during image description. It eliminates text-based reasoning shortcuts, forcing models to perform genuine visual inference across categories like spatial relations, quantitative shifts, and stylistic details. Use when the user wants to benchmark on VisuLogic, or asks about evaluating this task. Reports accuracy.
- ▌ Voxpopuli Eval · qhjqhj00Benchmarks multilingual speech representation learning and semi-supervised ASR/ST performance across multiple languages and domains, measuring phoneme discriminability, recognition accuracy, and translation quality. Use when the user wants to benchmark on VoxPopuli, Common Voice, ZeroSpeech 2017, EuroParl-ST, CoVoST 2, or asks about evaluating this task. Reports WER.
- ▌ Voxtream2 Eval · qhjqhj00This evaluation probes a full-stream text-to-speech model's ability to generate intelligible, natural-sounding speech while dynamically controlling the speaking rate in real-time. It measures objective intelligibility, speaker similarity, audio quality, generation latency, and the accuracy of speaking-rate control against target rates. Use when the user wants to benchmark on Emilia speaking-rate dataset, or asks about evaluating this task. Reports WER (%).
- ▌ Vprochart Eval · qhjqhj00Evaluates a model's ability to understand chart visuals and perform multi-step numerical and logical reasoning to answer natural language questions. It specifically probes visual perception alignment and programmatic solution reasoning over structured chart data. Use when the user wants to benchmark on ChartQA, PlotQA, DVQA, or asks about evaluating this task. Reports accuracy.
- ▌ Vsi Bench Eval · qhjqhj00Evaluates multimodal large language models' ability to reason about 3D spatial relationships from visual inputs. It probes capabilities across configurational tasks (counting, relative distance/direction, route planning), measurement estimation (object/room size, absolute distance), and spatiotemporal ordering. Use when the user wants to benchmark on VSI-Bench, or asks about evaluating this task. Reports score.
- ▌ Vtc Bench Eval · qhjqhj00Evaluates multimodal large language models' ability to perform agentic visual reasoning by composing multiple OpenCV-based tool calls. It probes long-horizon planning, precise tool selection, and the capacity to chain coarse- to fine-grained visual operations to solve complex multi-step problems. Use when the user wants to benchmark on VTC-Bench, or asks about evaluating this task. Reports Average Pass Rate (APR).
- ▌ Wa Hls4ml Eval · qhjqhj00Probes the ability of surrogate models to accurately predict FPGA resource usage (e.g., LUTs, BRAMs) and inference latency (clock cycles) for neural networks synthesized via hls4ml. It evaluates how well learned models approximate time-consuming hardware synthesis processes without running the full compilation pipeline. Use when the user wants to benchmark on wa-hls4ml benchmark, or asks about evaluating this task. Reports R².
- ▌ Weather2k Eval · qhjqhj00Evaluates the ability of deep learning models to forecast multivariate and univariate meteorological factors (temperature, visibility, humidity) using historical time-series and spatio-temporal data from ground weather stations. Use when the user wants to benchmark on Weather2K, or asks about evaluating this task. Reports MAE.
- ▌ Weatherqa Eval · qhjqhj00Evaluates multimodal reasoning capabilities in the meteorological domain, specifically testing a model's ability to interpret weather maps and answer domain-specific multiple-choice questions. It also measures cross-task generalization and the logical consistency of the model's reasoning chains versus final answers. Use when the user wants to benchmark on WeatherQA, ScienceQA, or asks about evaluating this task. Reports Multiple-choice accuracy.
- ▌ Webcode2m Eval · qhjqhj00This benchmark evaluates multimodal models' ability to translate webpage design screenshots into functional HTML/CSS code. It probes visual fidelity, structural hierarchy recall, and the capacity to generate long, complex, real-world front-end code from visual inputs. Use when the user wants to benchmark on WebCode2M, or asks about evaluating this task. Reports TreeBLEU.
- ▌ Webuav 3m Eval · qhjqhj00Evaluates the robustness and accuracy of deep visual tracking algorithms on large-scale, real-world UAV video sequences. It probes how well trackers handle diverse environmental conditions, motion dynamics, and target appearance changes without parameter tuning. Use when the user wants to benchmark on WebUAV-3M, or asks about evaluating this task. Reports AUC.
- ▌ Webuot 1m Eval · qhjqhj00Evaluates the robustness and accuracy of deep object trackers in challenging underwater environments. It probes cross-domain adaptation from open-air to underwater domains, as well as within-domain fine-tuning capabilities, while also assessing performance under varying frame rates and complex visual conditions like occlusion and low visibility. Use when the user wants to benchmark on WebUOT-1M, or asks about evaluating this task. Reports AUC.
- ▌ Webwalker Eval · qhjqhj00Evaluates LLM-based agents' ability to systematically navigate multi-layered, real-world websites to extract buried information. It probes long-range reasoning, memory management, and step-by-step click-based navigation under strict action limits. Use when the user wants to benchmark on WebWalkerQA, or asks about evaluating this task. Reports accuracy.
- ▌ Weedsense Eval · qhjqhj00Evaluates a model's ability to jointly perform fine-grained weed species segmentation, continuous plant height regression, and discrete temporal growth stage classification from single RGB images. Use when the user wants to benchmark on WeedSense, or asks about evaluating this task. Reports mIoU, MAE, Accuracy.
- ▌ Wiki Eval Eval · qhjqhj00Measures how well automated RAG scoring metrics align with human preferences in pairwise comparison tasks. It probes the ability of reference-free faithfulness, answer relevance, and context relevance estimators to replicate human judgment on answer and context quality. Use when the user wants to benchmark on WikiEval, or asks about evaluating this task. Reports accuracy.
- ▌ Wilddet3d Eval · qhjqhj00Evaluates open-vocabulary monocular 3D object detection across diverse real-world and synthetic scenes. It probes the model's ability to localize and regress 3D bounding boxes using text or geometric prompts, measuring generalization to unseen categories and datasets with and without depth cues. Use when the user wants to benchmark on WildDet3D-Bench, Omni3D, Argoverse 2, ScanNet, Stereo4D, or asks about evaluating this task. Reports AP_3D.
- ▌ Wildscore Eval · qhjqhj00This benchmark evaluates multimodal large language models' ability to perform multi-step, context-sensitive reasoning over symbolic musical notation. It probes capabilities in harmonic analysis, rhythmic interpretation, structural form recognition, and expressive markings through multiple-choice questions derived from real-world compositions and forum queries. Use when the user wants to benchmark on WildScore, or asks about evaluating this task. Reports accuracy.
- ▌ Wili 2018 Eval · qhjqhj00Evaluates the ability of models to correctly identify the language of monolingual text paragraphs. It probes language identification capabilities across a wide range of languages (235) with balanced representation. Use when the user wants to benchmark on WiLI-2018, or asks about evaluating this task. Reports F1.
- ▌ Winoqueer Eval · qhjqhj00Evaluates anti-LGBTQ+ bias in language models by measuring their tendency to prefer stereotypical completions over counterfactual ones when prompted with identity-specific contexts. Use when the user wants to benchmark on WinoQueer, or asks about evaluating this task. Reports bias score.
- ▌ Wmt19 Slt Eval · qhjqhj00Evaluates machine translation quality between similar languages (Czech to Polish) using a multi-encoder transformer trained on out-of-domain data filtered by cross-entropy differences. Probes the model's ability to adapt to low-resource similar language pairs via domain adaptation and data selection. Use when the user wants to benchmark on WMT19 SLT Shared Task dataset, or asks about evaluating this task. Reports BLEU.
- ▌ Wmt21 Nmt Eval · qhjqhj00Evaluates neural machine translation performance across news and biomedical domains for English-German and English-Russian language pairs. It probes the model's ability to handle domain-specific vocabulary, cross-lingual alignment, and translation quality under constrained data conditions typical of shared task tracks. Use when the user wants to benchmark on WMT21 News & Biomedical Shared Tasks, or asks about evaluating this task. Reports BLEU.
- ▌ Wmt22 Slt Eval · qhjqhj00Evaluates sign language translation from video to spoken text. It probes the model's ability to handle long videos, large vocabularies, and high singleton rates by leveraging full-body and lip-reading visual features. Use when the user wants to benchmark on WMT 2022 Shared Task, PHOENIX 2014T, or asks about evaluating this task. Reports BLEU.
- ▌ Workarena Eval · qhjqhj00Evaluates web agents' ability to perform complex, knowledge-worker tasks on enterprise UIs (ServiceNow) and standard web benchmarks. It probes multimodal browser observation processing, large DOM navigation, and action execution in interactive environments. Use when the user wants to benchmark on WorkArena, MiniWoB, WebGum Subset, or asks about evaluating this task. Reports success rate.
- ▌ Xcodeeval Eval · qhjqhj00Evaluates large language models on multilingual code understanding, generation, translation, and retrieval across 11 programming languages. It probes the model's ability to produce executable, correct code by validating outputs against unit tests rather than relying on lexical overlap. Use when the user wants to benchmark on xCodeEval, or asks about evaluating this task. Reports pass@5.
- ▌ Xiyan SQL Eval · qhjqhj00Evaluates a Text-to-SQL framework's ability to generate correct and efficient SQL queries from natural language questions across complex, cross-domain databases. It probes schema filtering, multi-generator candidate creation, and selection robustness. Use when the user wants to benchmark on BIRD, Spider, or asks about evaluating this task. Reports Execution Accuracy (EX).
- ▌ Xmodbench Eval · qhjqhj00This benchmark probes the cross-modal consistency and reasoning capabilities of omni-language models by evaluating semantic equivalence across all six possible modality combinations (text, vision, audio) for both context and candidate inputs. It measures how well models maintain performance when modalities are swapped or combined, highlighting modality-specific biases and directional asymmetries. Use when the user wants to benchmark on XModBench, or asks about evaluating this task. Reports accuracy.
- ▌ Xrl Bench Eval · qhjqhj00Evaluates the fidelity and stability of state-explaining methods in Reinforcement Learning across tabular and image-based environments. It measures how accurately explanations identify critical states and how consistently they perform under perturbations. Use when the user wants to benchmark on XRL-Bench Environments (DunkCityDynasty-v1, LunarLander-v2, CartPole-v0, FlappyBird-v0, Breakout-v0, Pong-v0), or asks about evaluating this task. Reports AIM.
- ▌ Xtc Bench Eval · qhjqhj00Evaluates cross-task semantic consistency in unified multimodal models by measuring how well generation and understanding tasks align on shared scene-graph facts. It probes whether architectural unification leads to representation-level coherence or merely independent task accuracy, specifically highlighting failures like consistent hallucination. Use when the user wants to benchmark on XTC-Bench, or asks about evaluating this task. Reports CCTA, AW-CCTA.
- ▌ Zeroquant Eval · qhjqhj00Evaluates the accuracy and inference latency of post-training quantized Transformer models (BERT and GPT-3-style) on standard NLP benchmarks and language modeling tasks. Use when the user wants to benchmark on GLUE benchmark, 20 zero-shot evaluation tasks, PTB / Wikitext-2 / Wikitext-103, or asks about evaluating this task. Reports average accuracy.
- ▌ Zerosense Eval · qhjqhj00Evaluates a model's ability to perform visual-text compression (OCR) by measuring raw text retention on rendered documents where the textual content has been deliberately stripped of semantic meaning. It isolates pure visual decoding capability from downstream linguistic priors or contextual inference. Use when the user wants to benchmark on ZeroSense, or asks about evaluating this task. Reports text preservation capability.
- ▌ Zoombench Eval · qhjqhj00Evaluates fine-grained multimodal perception, visual grounding, and reasoning capabilities of vision-language models. It measures performance across general perception, specific perception (color, counting), and out-of-distribution generalization tasks using a suite of established and custom benchmarks. Use when the user wants to benchmark on ZoomBench, HR-Bench, VStar, CV-Bench, MME-RealWorld, ColorBench, CountQA, MMStar, BabyVision, or asks about evaluating this task. Reports accuracy.
- ▌ Finch Data Analysis · qhjqhj00 bundleHosted biological-data-analysis agent (Finch) on the FutureHouse Platform. Hands a dataset + question to Finch, which builds a Jupyter notebook that explores, analyzes, and interprets the data. Use when the user has a biological dataset (omics, imaging, clinical) and a research question, and wants a multi-step analysis with code + results, not just a literature answer.
- ▌ 3doc Bench Eval · qhjqhj00This benchmark evaluates a model's ability to generate text-to-image outputs that strictly adhere to 3D layout constraints, handle complex inter-object occlusions, and maintain correct object orientations and visibility orders. It probes depth-consistent scene composition, attribute binding to specific objects, and overall image fidelity under varying camera viewpoints. Use when the user wants to benchmark on 3DOc-Bench, or asks about evaluating this task. Reports depth ordering.
- ▌ Abstain QA Eval · qhjqhj00Evaluates large language models' ability to abstain from answering when questions are unanswerable or when uncertain, while maintaining accuracy on answerable questions. It measures how well models balance abstention with correct answer selection under different prompting strategies and uncertainty calibration methods. Use when the user wants to benchmark on Abstain-QA, or asks about evaluating this task. Reports Abstention Rate (AR).
- ▌ Adamerging Eval · qhjqhj00Evaluates multi-task model merging methods on image classification tasks by measuring average accuracy across multiple datasets, generalization to unseen tasks, and robustness to image corruptions. Use when the user wants to benchmark on Image Classification Bench (SUN397, Cars, RESISC45, EuroSAT, SVHN, GTSRB, MNIST, DTD), or asks about evaluating this task. Reports Avg Acc.
- ▌ Adrd Bench Eval · qhjqhj00Evaluates LLMs on domain-specific knowledge and clinical reasoning for Alzheimer's Disease and Related Dementias (ADRD), as well as practical daily caregiving scenarios. It probes both factual recall and error detection capabilities in a medical context. Use when the user wants to benchmark on ADRD-Bench, or asks about evaluating this task. Reports exact match accuracy.
- ▌ Agent Spec Eval · qhjqhj00Evaluates the cross-framework portability and reusability of declarative agent specifications by executing identical agentic workflows across four different runtime frameworks (AutoGen, CrewAI, LangGraph, WayFlow) on three distinct task benchmarks. Use when the user wants to benchmark on SimpleQA Verified, BIRD-SQL, $\tau^{2}$-Bench, or asks about evaluating this task. Reports F1 score, EX%, Passˆk.
- ▌ Agentbench Eval · qhjqhj00Evaluates LLMs as autonomous agents across eight diverse, real-world environments requiring multi-turn interaction, long-term reasoning, decision-making, and strict instruction following. The benchmark measures success rates across code, game, and web-based tasks to identify performance gaps between commercial and open-source models. Use when the user wants to benchmark on AgentBench, or asks about evaluating this task. Reports overall_score.
- ▌ Agentboard Eval · qhjqhj00Evaluates an agent's ability to complete multi-round interactive tasks and achieve target goals across diverse task types. It simulates real-world environments where the model must navigate sequential decision-making to reach a defined endpoint. Use when the user wants to benchmark on AgentBoard, or asks about evaluating this task. Reports target achievement rate.
- ▌ Agentquest Eval · qhjqhj00This evaluation protocol measures LLM agent performance on multi-step reasoning tasks by tracking step-wise progress toward goal completion and the frequency of repetitive actions or states. It enables fine-grained debugging and architectural refinement beyond simple pass/fail success rates. Use when the user wants to benchmark on ALFWorld, Sudoku, or asks about evaluating this task. Reports progress rate.
- ▌ Agentseval Eval · qhjqhj00Probes the clinical faithfulness, factual accuracy, and diagnostic logic of medical imaging report generation systems. It evaluates robustness to paraphrasing and semantic perturbations by decomposing assessment into interpretable reasoning stages that mimic radiologist workflows. Use when the user wants to benchmark on Five medical imaging datasets (names not provided in excerpt), or asks about evaluating this task. Reports AgentsEval Score.
- ▌ Agentsynth Eval · qhjqhj00Evaluates the ability of multimodal language models to execute long-horizon, multi-step computer-use tasks on a desktop environment. It probes visual grounding, precise GUI interaction, state tracking, and error recovery across varying task complexities and software domains. Use when the user wants to benchmark on AgentSynth, or asks about evaluating this task. Reports success rate.
- ▌ Agentvista Eval · qhjqhj00Evaluates the ability of multimodal agents to perform long-horizon, multi-step tool use in complex, realistic visual environments. It probes cross-image reasoning, constraint tracking, and robust grounding when interacting with dynamic tools like web search and code execution. Use when the user wants to benchmark on AgentVista, or asks about evaluating this task. Reports accuracy.
- ▌ Aigvdbench Eval · qhjqhj00Evaluates the ability of AI-generated video detectors to distinguish between real and synthetically generated videos across diverse generation models, tasks (T2V, I2V, V2V), and temporal/spatial artifacts. Use when the user wants to benchmark on AIGVDBench, or asks about evaluating this task. Reports accuracy.
- ▌ Airs Bench Eval · qhjqhj00Evaluates AI research agents across the full scientific lifecycle, including idea generation, experiment design, and iterative refinement. Agents must generate and execute code to train models on specified datasets without baseline code, testing reasoning, generalization, and solution exploration capabilities. Use when the user wants to benchmark on AIRS-Bench, or asks about evaluating this task. Reports average normalized score.
- ▌ Aist Dance Eval · qhjqhj00Evaluates the quality, diversity, and music-motion synchronization of generated 3D dance sequences. It probes a model's ability to synthesize physically plausible, choreographically diverse, and rhythm-aligned human motion from audio input. Use when the user wants to benchmark on AIST++, or asks about evaluating this task. Reports FID_k.
- ▌ Alden Vrdu Eval · qhjqhj00Evaluates vision-language models' ability to actively navigate long, visually rich documents to gather evidence and answer complex queries. It probes multi-turn reasoning, retrieval accuracy, and the effectiveness of direct page-index access versus semantic search. Use when the user wants to benchmark on MMLongBench, LongDocURL, PaperTab, PaperText, FetaTab, DUDE-sub, or asks about evaluating this task. Reports GPT-4o–judged answer accuracy (Acc).
- ▌ Alora Peft Eval · qhjqhj00Evaluates the effectiveness of dynamic low-rank adaptation (LoRA) for fine-tuning large language models across classification, question answering, and instruction generation tasks. It measures how well rank allocation strategies preserve performance while maintaining or reducing tunable parameter counts. Use when the user wants to benchmark on SQuAD, BoolQ, COPA, ReCoRD, SST-2, RTE, QNLI, Alpaca, MT-Bench, E2E, or asks about evaluating this task. Reports accuracy, GPT-4 score.
- ▌ Alpacaeval Eval · qhjqhj00Evaluates LLM response quality via pairwise win rates against a baseline, while specifically probing the metric's susceptibility to length bias, gameability via verbosity prompting, and robustness to adversarial truncation. Use when the user wants to benchmark on AlpacaEval, or asks about evaluating this task. Reports Win rate.
- ▌ Am Lora Cl Eval · qhjqhj00Evaluates a model's ability to continuously learn multiple text classification tasks without catastrophic forgetting, measuring how well it retains knowledge of previous tasks while adapting to new ones. Use when the user wants to benchmark on Standard CL benchmarks, Large number of tasks benchmark, or asks about evaluating this task. Reports average results.
- ▌ Anomalygen Eval · qhjqhj00This benchmark evaluates log-based anomaly detection models by measuring their ability to classify log sequences as normal or anomalous. It specifically probes how well different model paradigms (classical ML, supervised/unsupervised deep learning, and LLM-based) generalize when trained on code-guided synthetic data augmentation across varying augmentation ratios. Use when the user wants to benchmark on HDFS, Zookeeper, or asks about evaluating this task. Reports F1-score.
- ▌ Answersumm Eval · qhjqhj00Evaluates multi-perspective answer summarization for community question answering, probing content selection, perspective clustering, abstractive summarization, and factual consistency/coverage. Use when the user wants to benchmark on AnswerSumm, or asks about evaluating this task. Reports F1, ROUGE-1/2/L.
- ▌ Ape Prompt Eval · qhjqhj00Evaluates the effectiveness of automatically generated prompts (instructions) from the APE framework compared to human-designed or baseline prompts across various natural language processing tasks. Use when the user wants to benchmark on Instruction Induction, BIG-Bench Instruction Induction (BBII), MultiArith, GSM8K, or asks about evaluating this task. Reports zero-shot execution accuracy.
- ▌ Apibench Q Eval · qhjqhj00Evaluates the retrieval accuracy and ranking quality of query-based API recommendation systems for Java APIs at both class and method levels. It also measures how query reformulation techniques impact recommendation performance. Use when the user wants to benchmark on APIBench-Q, or asks about evaluating this task. Reports Success Rate@k.
- ▌ Arabicmmlu Eval · qhjqhj00ArabicMMLU probes massive multitask language understanding in Modern Standard Arabic across 40 educational subjects spanning STEM, social sciences, humanities, Arabic language, and culturally specific domains. It evaluates models on cross-lingual transfer, cultural localization, and robustness to linguistic phenomena like negation across primary, middle, high school, and university levels. The benchmark specifically tests how well models handle Arabic-specific knowledge and exam-style multiple-choice questions. Use when the user wants to benchmark on ArabicMMLU, or asks about evaluating this task. Reports accuracy.
- ▌ Arabicmteb Eval · qhjqhj00Evaluates Arabic-centric and cross-lingual text embedding models across multiple linguistic, cultural, and domain-specific capabilities. It probes how well models capture dialectal variations, regional cultural knowledge, and specialized domain terminology in Arabic. Use when the user wants to benchmark on ArabicMTEB, or asks about evaluating this task. Reports Avg..
- ▌ Argscichat Eval · qhjqhj00Evaluates the ability of dialogue agents to select supportive facts from scientific papers and generate contextually appropriate responses in argumentative scientific dialogues. Probes document-grounded response generation and fact selection under expert-level, opinion-driven interactions. Use when the user wants to benchmark on ArgSciChat, or asks about evaluating this task. Reports Fact-F1.
- ▌ Asd And Se Eval · qhjqhj00Evaluates a unified audio-visual model's ability to detect which speaker is actively speaking in multi-person video scenes and to enhance speech signals by removing background noise and interference. Use when the user wants to benchmark on AVA-ActiveSpeaker, LRS2, TalkSet, Columbia, MUSAN, or asks about evaluating this task. Reports mAP.
- ▌ Astro Mcad Eval · qhjqhj00Evaluates a classifier-based anomaly detection pipeline on simulated astronomical transient light curves. It probes the model's ability to identify rare, out-of-distribution events in real-time without prior exposure to the anomalous classes during training. Use when the user wants to benchmark on Simulated LSST-like transient light curves, or asks about evaluating this task. Reports anomaly score.
- ▌ Astrochart Eval · qhjqhj00Evaluates multimodal large language models' ability to comprehend scientific charts and perform knowledge-intensive reasoning in astronomy. It probes visual understanding, data extraction, numerical calculation, and domain-specific inference. Use when the user wants to benchmark on AstroChart, or asks about evaluating this task. Reports Accuracy (%).
- ▌ Atlas Chat Eval · qhjqhj00Evaluates large language models on Moroccan Arabic (Darija) across multiple-choice reasoning, instruction following, translation, summarization, and sentiment analysis. It probes dialect-specific linguistic features, script variability (Arabic vs. Arabizi), and real-world instruction-following capabilities in a low-resource setting. Use when the user wants to benchmark on DarijaMMLU, DarijaHellaSwag, Belebele_Ary, DarijaBench, DarijaAlpacaEval, or asks about evaluating this task. Reports Accuracy.
- ▌ Audiomnist Eval · qhjqhj00Evaluates audio classification performance on spoken digits and speaker sex using raw waveforms and spectrograms, serving as a benchmark for explainable AI (XAI) methods in the audio domain. Use when the user wants to benchmark on AudioMNIST, or asks about evaluating this task. Reports accuracy.
- ▌ Auto Split Eval · qhjqhj00Evaluates the latency, accuracy, and model size of a collaborative edge-cloud DNN splitting framework (Auto-Split) compared to baselines like QDMP and Neurosurgeon across image classification and object detection tasks. Use when the user wants to benchmark on ImageNet, COCO 2017, or asks about evaluating this task. Reports End-to-end latency (normalized).
- ▌ Averimavec Eval · qhjqhj00This benchmark probes a model's ability to perform multimodal fact-checking by verifying real-world image-text claims. It requires the system to retrieve cross-modal evidence, analyze inconsistencies, and produce a justified verdict that aligns with ground truth labels. Use when the user wants to benchmark on AVerImaTeC, or asks about evaluating this task. Reports verdict_correctness.
- ▌ Babyvision Eval · qhjqhj00Evaluates fundamental visual reasoning capabilities in multimodal large language models independent of linguistic priors. It probes early-vision abilities such as visual tracking, spatial perception, fine-grained discrimination, and visual pattern recognition through image-based tasks. Use when the user wants to benchmark on BabyVision, or asks about evaluating this task. Reports Avg@3.
- ▌ Basqueglue Eval · qhjqhj00This benchmark evaluates Basque language models across classical NLP tasks including topic classification, stance detection, coreference detection, and natural language inference. It measures both task-specific performance and overall linguistic competence in a low-resource agglutinative language setting. Use when the user wants to benchmark on BasqueGLUE, or asks about evaluating this task. Reports Avg.
- ▌ Batonvoice Eval · qhjqhj00Evaluates a controllable text-to-speech model's ability to generate intelligible speech and accurately convey specific emotional tones based on text instructions. It probes zero-shot cross-lingual generalization and instruction-following capabilities in speech synthesis. Use when the user wants to benchmark on Seed-TTS, Emotion dataset, or asks about evaluating this task. Reports Emotion Classification Accuracy.
- ▌ Battleship Eval · qhjqhj00Evaluates EFCE solvers on a parametric sequential conflict-resolution game where players place ships and fire shots. It probes the solver's ability to construct incentive-compatible correlation plans that maximize social welfare through deterrence and punishment mechanisms. Use when the user wants to benchmark on Battleship, or asks about evaluating this task. Reports Social Welfare (SW).
- ▌ Beans Zero Eval · qhjqhj00Evaluates zero-shot generalization of audio-language models on bioacoustic tasks, including species classification, multilabel detection, call-type prediction, lifestage classification, captioning, and individual counting across diverse taxa. Use when the user wants to benchmark on BEANS-Zero, or asks about evaluating this task. Reports accuracy.
- ▌ Behavior1k Eval · qhjqhj00Evaluates embodied AI agents on long-horizon, human-centered manipulation tasks in a realistic physics-based simulation. It probes the agent's ability to plan and execute complex sequences of action primitives (pick, place, navigate, etc.) while handling rigid, articulated, and deformable objects. Use when the user wants to benchmark on BEHAVIOR-1K, or asks about evaluating this task. Reports task success rate.
- ▌ Bench Push Eval · qhjqhj00Evaluates the transferability and performance of reinforcement learning policies for mobile robot navigation and pushing-based manipulation tasks. It probes how well policies trained in simulation handle real-world sim-to-real gaps, clutter, and sparse rewards across varying obstacle densities. Use when the user wants to benchmark on Bench-Push (Maze & Box-Delivery), or asks about evaluating this task. Reports $S_{\text{manip}}$.
- ▌ Benchie Fl Eval · qhjqhj00Evaluates Open Information Extraction (OIE) systems on their ability to extract fact-based triples from text. It uses a conservative exact-matching function with synset-based clustering to penalize non-informative copies and reward precise fact extraction, while also measuring correlation with downstream QA and knowledge base tasks. Use when the user wants to benchmark on BenchIE^FL, or asks about evaluating this task. Reports exact-match.
- ▌ Ber Performance · qhjqhj00Evaluates the bit error rate (BER) performance of a Reconfigurable Intelligent Surface (RIS) aided spatial media-based modulation system compared to baseline schemes (SM, MBM, QSM) under uncorrelated Rayleigh fading channels. Use when the user has predictions and gold and needs to compute Bit Error Rate (BER).
- ▌ Binaryhingeloss · qhjqhj00Compute the BinaryHingeLoss metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryHingeLoss, or asks how to score with BinaryHingeLoss.
- ▌ Binaryprecision · qhjqhj00Compute the BinaryPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryPrecision, or asks how to score with BinaryPrecision.
- ▌ Biobert QA Eval · qhjqhj00Evaluates factoid question answering performance on small biomedical datasets, testing transfer learning effectiveness from domain-specific pre-training. It measures how well the model retrieves exact or lenient answers to biomedical queries. Use when the user wants to benchmark on BioASQ 4b, BioASQ 5b, BioASQ 6b, or asks about evaluating this task. Reports Mean Reciprocal Rank (MRR).
- ▌ Biobert Re Eval · qhjqhj00Tests a model's capability to extract biomedical relations (gene-disease, gene-chemical) from text using minimal task-specific modifications. It probes whether domain-specific pre-training improves relation classification on small-scale biomedical corpora. Use when the user wants to benchmark on GAD, EU-ADR, CHEMPROT, or asks about evaluating this task. Reports entity-level F1.
- ▌ Biomed Vqa Eval · qhjqhj00Evaluates the domain-adaptive post-training of multimodal large language models on biomedical visual question answering tasks, measuring how well models generalize to specialized medical domains using both open and closed evaluation splits. Use when the user wants to benchmark on SLAKE, PathVQA, VQA-RAD, PMC-VQA, or asks about evaluating this task. Reports accuracy.
- ▌ Bionli 300 Eval · qhjqhj00This evaluation probes a model's ability to verify scientific claims against provided or retrieved evidence in a binary classification setting. It measures performance on Supported vs. Refuted labels, testing factual grounding, uncertainty calibration, and the impact of atomic decomposition and web corroboration. Use when the user wants to benchmark on BIONLI-300, or asks about evaluating this task. Reports Balanced Accuracy.
- ▌ Bioscan 5m Eval · qhjqhj00Evaluates models on insect biodiversity monitoring by testing closed-world species identification, open-world genus-level grouping for novel species, and zero-shot clustering of multimodal embeddings against taxonomic ground truth. Use when the user wants to benchmark on BIOSCAN-5M, or asks about evaluating this task. Reports Fine-tuned accuracy.
- ▌ Bird Bench Eval · qhjqhj00Evaluates the capability of large language models to generate correct SQL queries from natural language questions. It specifically probes how annotation noise and errors in benchmark datasets affect model performance and reliability. Use when the user wants to benchmark on BIRD-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Bouquet Mt Eval · qhjqhj00Evaluates machine translation systems on a contamination-free, multilingual dataset covering diverse domains and registers. It measures translation quality at both sentence and paragraph levels to assess how well models handle linguistic diversity and cultural authenticity across 8 major languages. Use when the user wants to benchmark on BOUQuET, or asks about evaluating this task. Reports CometKiwi.