all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 52 of 76

  1. ▌
    Ultralink Eval · qhjqhj00
    Evaluates multilingual LLMs on chat, math reasoning, and code generation across five languages (English, Chinese, Spanish, Russian, French) to measure the effectiveness of knowledge-enhanced supervised fine-tuning. Use when the user wants to benchmark on OMGEval, MGSM, Multilingual HumanEval, or asks about evaluating this task. Reports OMGEval score.
    3 repo stars
  2. ▌
    Univbench Eval · qhjqhj00
    Evaluates video foundation models across six core tasks (understanding, generation, editing, reconstruction) by scoring their outputs on eight cinematic and semantic dimensions, including subject consistency, action dynamics, camera movement, and lighting. Use when the user wants to benchmark on UniVBench, or asks about evaluating this task. Reports UniV-Eval.
    3 repo stars
  3. ▌
    Univearth Eval · qhjqhj00
    This benchmark probes an LLM agent's ability to perform spatial-temporal reasoning and generate executable code for Earth Observation tasks. It evaluates whether models can correctly answer yes/no questions derived from scientific articles by leveraging remote sensing data via Google Earth Engine. Use when the user wants to benchmark on UnivEARTH, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  4. ▌
    Unrealzoo Eval · qhjqhj00
    Evaluates embodied AI agents' capabilities in complex, photo-realistic 3D open-world environments. Specifically probes visual navigation on unstructured terrain, active visual tracking across diverse scenes, and social tracking under dynamic distractions, varying morphologies, and different control frequencies. Use when the user wants to benchmark on UnrealZoo, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  5. ▌
    Unsw Nb15 Eval · qhjqhj00
    Evaluates network intrusion detection capability by classifying network traffic flows as benign or malicious (or specific attack types) using graph-structured representations of network connections. It probes the model's ability to learn from adaptive graph construction and contrastive learning under resource-constrained conditions. Use when the user wants to benchmark on UNSW-NB15, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  6. ▌
    Urdu Mner Eval · qhjqhj00
    Evaluates the ability of models to recognize named entities (Person, Location, Organization, Miscellaneous) in Urdu social media posts by jointly processing textual and visual inputs. It probes cross-modal alignment, handling of low-resource language morphological complexity, and ambiguity resolution using visual context. Use when the user wants to benchmark on Twitter2015-Urdu, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  7. ▌
    V2x Radar Eval · qhjqhj00
    Evaluates 3D object detection capabilities for autonomous driving using multi-modal sensors (LiDAR, camera, 4D radar) in single-agent (roadside and vehicle-mounted) and cooperative perception setups. It probes robustness to adverse weather conditions and communication delays in cooperative scenarios. Use when the user wants to benchmark on V2X-Radar, or asks about evaluating this task. Reports AP@IoU.
    3 repo stars
  8. ▌
    Vaexbench Eval · qhjqhj00
    Evaluates multimodal large language models' ability to perform extractive and abstractive spatiotemporal reasoning on egocentric videos. It probes long-horizon memory, object tracking, spatial orientation, and metric distance estimation under both multiple-choice and free-form generation settings. Use when the user wants to benchmark on VAEX-Bench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  9. ▌
    Valerie22 Eval · qhjqhj00
    This protocol evaluates the perceptual fidelity and cross-domain generalization capability of the VALERIE22 synthetic urban dataset by training a semantic segmentation model on it and testing on real-world automotive datasets. It specifically probes how dataset diversity (unique 3D assets) and training scale affect downstream perception performance. Use when the user wants to benchmark on VALERIE22, Cityscapes, A2D2, BDD100K, India Driving Dataset, Mapillary Vistas, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  10. ▌
    Vc Ifeval Eval · qhjqhj00
    Evaluates multimodal large language models' ability to follow instructions containing vision-dependent constraints, such as spatial, stylistic, and structural requirements. It isolates the contribution of visual input to instruction adherence and assesses generalization on standard visual reasoning tasks. Use when the user wants to benchmark on VC-IFEval, MM-IFEval, IFEval, or asks about evaluating this task. Reports instruction-following accuracy.
    3 repo stars
  11. ▌
    Vcb Bench Eval · qhjqhj00
    VCB Bench evaluates audio-grounded large language models on instruction following with speech-level controls, knowledge reasoning, and robustness under real-world acoustic perturbations. It probes how well models understand and generate spoken responses in Chinese and English using authentic human speech rather than synthetic data. Use when the user wants to benchmark on VCB Bench, or asks about evaluating this task. Reports 1-5 scale score.
    3 repo stars
  12. ▌
    Vibe Eval Eval · qhjqhj00
    Probes multimodal reasoning and visual understanding on real-world images. Specifically designed with a 'hard' subset of prompts that are unsolvable by current frontier models to measure genuine performance gaps and contamination-free generalization. Use when the user wants to benchmark on Vibe-Eval, or asks about evaluating this task. Reports Vibe-Eval Score.
    3 repo stars
  13. ▌
    Video Mme Eval · qhjqhj00
    Evaluates long-horizon video understanding and reasoning capabilities of multimodal models on extended video sequences. Use when the user wants to benchmark on Video-MME, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  14. ▌
    Videocube Eval · qhjqhj00
    Evaluates a model's ability to track arbitrary visual instances across complex, unstructured real-world videos without assuming motion continuity or fixed categories. It measures both local search accuracy and global robustness against challenges like occlusion, fast motion, and scene transitions. Use when the user wants to benchmark on VideoCube, or asks about evaluating this task. Reports PRE.
    3 repo stars
  15. ▌
    Vidore V3 Eval · qhjqhj00
    This benchmark evaluates end-to-end Retrieval Augmented Generation (RAG) systems on visually rich, real-world documents across multiple professional domains. It probes a model's ability to retrieve relevant pages, generate accurate answers to complex open-ended and multi-hop queries, and precisely ground those answers with bounding boxes in multimodal content. Use when the user wants to benchmark on ViDoRe V3, or asks about evaluating this task. Reports F1 score (Dice coefficient).
    3 repo stars
  16. ▌
    Vimed Pet Eval · qhjqhj00
    Evaluates vision-language models on generating Vietnamese clinical reports from paired PET/CT images and answering medical questions about them. Probes the model's ability to align 3D medical imaging features with low-resource language text and produce clinically accurate descriptions. Use when the user wants to benchmark on ViMed-PET, or asks about evaluating this task. Reports BLEU-4.
    3 repo stars
  17. ▌
    Vision R1 Eval · qhjqhj00
    Evaluates Large Vision-Language Models on their ability to detect, localize, and ground objects in images across diverse and challenging scenarios, including in-domain dense detection, out-of-domain real-world settings, and generalization to unseen categories or scenes. Use when the user wants to benchmark on MSCOCO Val2017, ODINW-13, or asks about evaluating this task. Reports mAP.
    3 repo stars
  18. ▌
    Visonlyqa Eval · qhjqhj00
    This benchmark probes a model's ability to accurately perceive basic geometric information—such as shape, angle, length, area, and intersections—in scientific figures and diagrams. It isolates visual perception from higher-level reasoning or domain knowledge by using direct, low-reasoning questions on synthetic and real-world images. Use when the user wants to benchmark on VisOnlyQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Visulogic Eval · qhjqhj00
    VisuLogic probes vision-centric reasoning in multimodal large language models by presenting problems that require retaining critical visual cues during image description. It eliminates text-based reasoning shortcuts, forcing models to perform genuine visual inference across categories like spatial relations, quantitative shifts, and stylistic details. Use when the user wants to benchmark on VisuLogic, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  20. ▌
    Voxpopuli Eval · qhjqhj00
    Benchmarks multilingual speech representation learning and semi-supervised ASR/ST performance across multiple languages and domains, measuring phoneme discriminability, recognition accuracy, and translation quality. Use when the user wants to benchmark on VoxPopuli, Common Voice, ZeroSpeech 2017, EuroParl-ST, CoVoST 2, or asks about evaluating this task. Reports WER.
    3 repo stars
  21. ▌
    Voxtream2 Eval · qhjqhj00
    This evaluation probes a full-stream text-to-speech model's ability to generate intelligible, natural-sounding speech while dynamically controlling the speaking rate in real-time. It measures objective intelligibility, speaker similarity, audio quality, generation latency, and the accuracy of speaking-rate control against target rates. Use when the user wants to benchmark on Emilia speaking-rate dataset, or asks about evaluating this task. Reports WER (%).
    3 repo stars
  22. ▌
    Vprochart Eval · qhjqhj00
    Evaluates a model's ability to understand chart visuals and perform multi-step numerical and logical reasoning to answer natural language questions. It specifically probes visual perception alignment and programmatic solution reasoning over structured chart data. Use when the user wants to benchmark on ChartQA, PlotQA, DVQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  23. ▌
    Vsi Bench Eval · qhjqhj00
    Evaluates multimodal large language models' ability to reason about 3D spatial relationships from visual inputs. It probes capabilities across configurational tasks (counting, relative distance/direction, route planning), measurement estimation (object/room size, absolute distance), and spatiotemporal ordering. Use when the user wants to benchmark on VSI-Bench, or asks about evaluating this task. Reports score.
    3 repo stars
  24. ▌
    Vtc Bench Eval · qhjqhj00
    Evaluates multimodal large language models' ability to perform agentic visual reasoning by composing multiple OpenCV-based tool calls. It probes long-horizon planning, precise tool selection, and the capacity to chain coarse- to fine-grained visual operations to solve complex multi-step problems. Use when the user wants to benchmark on VTC-Bench, or asks about evaluating this task. Reports Average Pass Rate (APR).
    3 repo stars
  25. ▌
    Wa Hls4ml Eval · qhjqhj00
    Probes the ability of surrogate models to accurately predict FPGA resource usage (e.g., LUTs, BRAMs) and inference latency (clock cycles) for neural networks synthesized via hls4ml. It evaluates how well learned models approximate time-consuming hardware synthesis processes without running the full compilation pipeline. Use when the user wants to benchmark on wa-hls4ml benchmark, or asks about evaluating this task. Reports R².
    3 repo stars
  26. ▌
    Weather2k Eval · qhjqhj00
    Evaluates the ability of deep learning models to forecast multivariate and univariate meteorological factors (temperature, visibility, humidity) using historical time-series and spatio-temporal data from ground weather stations. Use when the user wants to benchmark on Weather2K, or asks about evaluating this task. Reports MAE.
    3 repo stars
  27. ▌
    Weatherqa Eval · qhjqhj00
    Evaluates multimodal reasoning capabilities in the meteorological domain, specifically testing a model's ability to interpret weather maps and answer domain-specific multiple-choice questions. It also measures cross-task generalization and the logical consistency of the model's reasoning chains versus final answers. Use when the user wants to benchmark on WeatherQA, ScienceQA, or asks about evaluating this task. Reports Multiple-choice accuracy.
    3 repo stars
  28. ▌
    Webcode2m Eval · qhjqhj00
    This benchmark evaluates multimodal models' ability to translate webpage design screenshots into functional HTML/CSS code. It probes visual fidelity, structural hierarchy recall, and the capacity to generate long, complex, real-world front-end code from visual inputs. Use when the user wants to benchmark on WebCode2M, or asks about evaluating this task. Reports TreeBLEU.
    3 repo stars
  29. ▌
    Webuav 3m Eval · qhjqhj00
    Evaluates the robustness and accuracy of deep visual tracking algorithms on large-scale, real-world UAV video sequences. It probes how well trackers handle diverse environmental conditions, motion dynamics, and target appearance changes without parameter tuning. Use when the user wants to benchmark on WebUAV-3M, or asks about evaluating this task. Reports AUC.
    3 repo stars
  30. ▌
    Webuot 1m Eval · qhjqhj00
    Evaluates the robustness and accuracy of deep object trackers in challenging underwater environments. It probes cross-domain adaptation from open-air to underwater domains, as well as within-domain fine-tuning capabilities, while also assessing performance under varying frame rates and complex visual conditions like occlusion and low visibility. Use when the user wants to benchmark on WebUOT-1M, or asks about evaluating this task. Reports AUC.
    3 repo stars
  31. ▌
    Webwalker Eval · qhjqhj00
    Evaluates LLM-based agents' ability to systematically navigate multi-layered, real-world websites to extract buried information. It probes long-range reasoning, memory management, and step-by-step click-based navigation under strict action limits. Use when the user wants to benchmark on WebWalkerQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  32. ▌
    Weedsense Eval · qhjqhj00
    Evaluates a model's ability to jointly perform fine-grained weed species segmentation, continuous plant height regression, and discrete temporal growth stage classification from single RGB images. Use when the user wants to benchmark on WeedSense, or asks about evaluating this task. Reports mIoU, MAE, Accuracy.
    3 repo stars
  33. ▌
    Wiki Eval Eval · qhjqhj00
    Measures how well automated RAG scoring metrics align with human preferences in pairwise comparison tasks. It probes the ability of reference-free faithfulness, answer relevance, and context relevance estimators to replicate human judgment on answer and context quality. Use when the user wants to benchmark on WikiEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  34. ▌
    Wilddet3d Eval · qhjqhj00
    Evaluates open-vocabulary monocular 3D object detection across diverse real-world and synthetic scenes. It probes the model's ability to localize and regress 3D bounding boxes using text or geometric prompts, measuring generalization to unseen categories and datasets with and without depth cues. Use when the user wants to benchmark on WildDet3D-Bench, Omni3D, Argoverse 2, ScanNet, Stereo4D, or asks about evaluating this task. Reports AP_3D.
    3 repo stars
  35. ▌
    Wildscore Eval · qhjqhj00
    This benchmark evaluates multimodal large language models' ability to perform multi-step, context-sensitive reasoning over symbolic musical notation. It probes capabilities in harmonic analysis, rhythmic interpretation, structural form recognition, and expressive markings through multiple-choice questions derived from real-world compositions and forum queries. Use when the user wants to benchmark on WildScore, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  36. ▌
    Wili 2018 Eval · qhjqhj00
    Evaluates the ability of models to correctly identify the language of monolingual text paragraphs. It probes language identification capabilities across a wide range of languages (235) with balanced representation. Use when the user wants to benchmark on WiLI-2018, or asks about evaluating this task. Reports F1.
    3 repo stars
  37. ▌
    Winoqueer Eval · qhjqhj00
    Evaluates anti-LGBTQ+ bias in language models by measuring their tendency to prefer stereotypical completions over counterfactual ones when prompted with identity-specific contexts. Use when the user wants to benchmark on WinoQueer, or asks about evaluating this task. Reports bias score.
    3 repo stars
  38. ▌
    Wmt19 Slt Eval · qhjqhj00
    Evaluates machine translation quality between similar languages (Czech to Polish) using a multi-encoder transformer trained on out-of-domain data filtered by cross-entropy differences. Probes the model's ability to adapt to low-resource similar language pairs via domain adaptation and data selection. Use when the user wants to benchmark on WMT19 SLT Shared Task dataset, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  39. ▌
    Wmt21 Nmt Eval · qhjqhj00
    Evaluates neural machine translation performance across news and biomedical domains for English-German and English-Russian language pairs. It probes the model's ability to handle domain-specific vocabulary, cross-lingual alignment, and translation quality under constrained data conditions typical of shared task tracks. Use when the user wants to benchmark on WMT21 News & Biomedical Shared Tasks, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  40. ▌
    Wmt22 Slt Eval · qhjqhj00
    Evaluates sign language translation from video to spoken text. It probes the model's ability to handle long videos, large vocabularies, and high singleton rates by leveraging full-body and lip-reading visual features. Use when the user wants to benchmark on WMT 2022 Shared Task, PHOENIX 2014T, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  41. ▌
    Workarena Eval · qhjqhj00
    Evaluates web agents' ability to perform complex, knowledge-worker tasks on enterprise UIs (ServiceNow) and standard web benchmarks. It probes multimodal browser observation processing, large DOM navigation, and action execution in interactive environments. Use when the user wants to benchmark on WorkArena, MiniWoB, WebGum Subset, or asks about evaluating this task. Reports success rate.
    3 repo stars
  42. ▌
    Xcodeeval Eval · qhjqhj00
    Evaluates large language models on multilingual code understanding, generation, translation, and retrieval across 11 programming languages. It probes the model's ability to produce executable, correct code by validating outputs against unit tests rather than relying on lexical overlap. Use when the user wants to benchmark on xCodeEval, or asks about evaluating this task. Reports pass@5.
    3 repo stars
  43. ▌
    Xiyan SQL Eval · qhjqhj00
    Evaluates a Text-to-SQL framework's ability to generate correct and efficient SQL queries from natural language questions across complex, cross-domain databases. It probes schema filtering, multi-generator candidate creation, and selection robustness. Use when the user wants to benchmark on BIRD, Spider, or asks about evaluating this task. Reports Execution Accuracy (EX).
    3 repo stars
  44. ▌
    Xmodbench Eval · qhjqhj00
    This benchmark probes the cross-modal consistency and reasoning capabilities of omni-language models by evaluating semantic equivalence across all six possible modality combinations (text, vision, audio) for both context and candidate inputs. It measures how well models maintain performance when modalities are swapped or combined, highlighting modality-specific biases and directional asymmetries. Use when the user wants to benchmark on XModBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  45. ▌
    Xrl Bench Eval · qhjqhj00
    Evaluates the fidelity and stability of state-explaining methods in Reinforcement Learning across tabular and image-based environments. It measures how accurately explanations identify critical states and how consistently they perform under perturbations. Use when the user wants to benchmark on XRL-Bench Environments (DunkCityDynasty-v1, LunarLander-v2, CartPole-v0, FlappyBird-v0, Breakout-v0, Pong-v0), or asks about evaluating this task. Reports AIM.
    3 repo stars
  46. ▌
    Xtc Bench Eval · qhjqhj00
    Evaluates cross-task semantic consistency in unified multimodal models by measuring how well generation and understanding tasks align on shared scene-graph facts. It probes whether architectural unification leads to representation-level coherence or merely independent task accuracy, specifically highlighting failures like consistent hallucination. Use when the user wants to benchmark on XTC-Bench, or asks about evaluating this task. Reports CCTA, AW-CCTA.
    3 repo stars
  47. ▌
    Zeroquant Eval · qhjqhj00
    Evaluates the accuracy and inference latency of post-training quantized Transformer models (BERT and GPT-3-style) on standard NLP benchmarks and language modeling tasks. Use when the user wants to benchmark on GLUE benchmark, 20 zero-shot evaluation tasks, PTB / Wikitext-2 / Wikitext-103, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  48. ▌
    Zerosense Eval · qhjqhj00
    Evaluates a model's ability to perform visual-text compression (OCR) by measuring raw text retention on rendered documents where the textual content has been deliberately stripped of semantic meaning. It isolates pure visual decoding capability from downstream linguistic priors or contextual inference. Use when the user wants to benchmark on ZeroSense, or asks about evaluating this task. Reports text preservation capability.
    3 repo stars
  49. ▌
    Zoombench Eval · qhjqhj00
    Evaluates fine-grained multimodal perception, visual grounding, and reasoning capabilities of vision-language models. It measures performance across general perception, specific perception (color, counting), and out-of-distribution generalization tasks using a suite of established and custom benchmarks. Use when the user wants to benchmark on ZoomBench, HR-Bench, VStar, CV-Bench, MME-RealWorld, ColorBench, CountQA, MMStar, BabyVision, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  50. ▌
    Finch Data Analysis · qhjqhj00 bundle
    Hosted biological-data-analysis agent (Finch) on the FutureHouse Platform. Hands a dataset + question to Finch, which builds a Jupyter notebook that explores, analyzes, and interprets the data. Use when the user has a biological dataset (omics, imaging, clinical) and a research question, and wants a multi-step analysis with code + results, not just a literature answer.
    3 repo stars
  51. ▌
    3doc Bench Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate text-to-image outputs that strictly adhere to 3D layout constraints, handle complex inter-object occlusions, and maintain correct object orientations and visibility orders. It probes depth-consistent scene composition, attribute binding to specific objects, and overall image fidelity under varying camera viewpoints. Use when the user wants to benchmark on 3DOc-Bench, or asks about evaluating this task. Reports depth ordering.
    3 repo stars
  52. ▌
    Abstain QA Eval · qhjqhj00
    Evaluates large language models' ability to abstain from answering when questions are unanswerable or when uncertain, while maintaining accuracy on answerable questions. It measures how well models balance abstention with correct answer selection under different prompting strategies and uncertainty calibration methods. Use when the user wants to benchmark on Abstain-QA, or asks about evaluating this task. Reports Abstention Rate (AR).
    3 repo stars
  53. ▌
    Adamerging Eval · qhjqhj00
    Evaluates multi-task model merging methods on image classification tasks by measuring average accuracy across multiple datasets, generalization to unseen tasks, and robustness to image corruptions. Use when the user wants to benchmark on Image Classification Bench (SUN397, Cars, RESISC45, EuroSAT, SVHN, GTSRB, MNIST, DTD), or asks about evaluating this task. Reports Avg Acc.
    3 repo stars
  54. ▌
    Adrd Bench Eval · qhjqhj00
    Evaluates LLMs on domain-specific knowledge and clinical reasoning for Alzheimer's Disease and Related Dementias (ADRD), as well as practical daily caregiving scenarios. It probes both factual recall and error detection capabilities in a medical context. Use when the user wants to benchmark on ADRD-Bench, or asks about evaluating this task. Reports exact match accuracy.
    3 repo stars
  55. ▌
    Agent Spec Eval · qhjqhj00
    Evaluates the cross-framework portability and reusability of declarative agent specifications by executing identical agentic workflows across four different runtime frameworks (AutoGen, CrewAI, LangGraph, WayFlow) on three distinct task benchmarks. Use when the user wants to benchmark on SimpleQA Verified, BIRD-SQL, $\tau^{2}$-Bench, or asks about evaluating this task. Reports F1 score, EX%, Passˆk.
    3 repo stars
  56. ▌
    Agentbench Eval · qhjqhj00
    Evaluates LLMs as autonomous agents across eight diverse, real-world environments requiring multi-turn interaction, long-term reasoning, decision-making, and strict instruction following. The benchmark measures success rates across code, game, and web-based tasks to identify performance gaps between commercial and open-source models. Use when the user wants to benchmark on AgentBench, or asks about evaluating this task. Reports overall_score.
    3 repo stars
  57. ▌
    Agentboard Eval · qhjqhj00
    Evaluates an agent's ability to complete multi-round interactive tasks and achieve target goals across diverse task types. It simulates real-world environments where the model must navigate sequential decision-making to reach a defined endpoint. Use when the user wants to benchmark on AgentBoard, or asks about evaluating this task. Reports target achievement rate.
    3 repo stars
  58. ▌
    Agentquest Eval · qhjqhj00
    This evaluation protocol measures LLM agent performance on multi-step reasoning tasks by tracking step-wise progress toward goal completion and the frequency of repetitive actions or states. It enables fine-grained debugging and architectural refinement beyond simple pass/fail success rates. Use when the user wants to benchmark on ALFWorld, Sudoku, or asks about evaluating this task. Reports progress rate.
    3 repo stars
  59. ▌
    Agentseval Eval · qhjqhj00
    Probes the clinical faithfulness, factual accuracy, and diagnostic logic of medical imaging report generation systems. It evaluates robustness to paraphrasing and semantic perturbations by decomposing assessment into interpretable reasoning stages that mimic radiologist workflows. Use when the user wants to benchmark on Five medical imaging datasets (names not provided in excerpt), or asks about evaluating this task. Reports AgentsEval Score.
    3 repo stars
  60. ▌
    Agentsynth Eval · qhjqhj00
    Evaluates the ability of multimodal language models to execute long-horizon, multi-step computer-use tasks on a desktop environment. It probes visual grounding, precise GUI interaction, state tracking, and error recovery across varying task complexities and software domains. Use when the user wants to benchmark on AgentSynth, or asks about evaluating this task. Reports success rate.
    3 repo stars
  61. ▌
    Agentvista Eval · qhjqhj00
    Evaluates the ability of multimodal agents to perform long-horizon, multi-step tool use in complex, realistic visual environments. It probes cross-image reasoning, constraint tracking, and robust grounding when interacting with dynamic tools like web search and code execution. Use when the user wants to benchmark on AgentVista, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  62. ▌
    Aigvdbench Eval · qhjqhj00
    Evaluates the ability of AI-generated video detectors to distinguish between real and synthetically generated videos across diverse generation models, tasks (T2V, I2V, V2V), and temporal/spatial artifacts. Use when the user wants to benchmark on AIGVDBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  63. ▌
    Airs Bench Eval · qhjqhj00
    Evaluates AI research agents across the full scientific lifecycle, including idea generation, experiment design, and iterative refinement. Agents must generate and execute code to train models on specified datasets without baseline code, testing reasoning, generalization, and solution exploration capabilities. Use when the user wants to benchmark on AIRS-Bench, or asks about evaluating this task. Reports average normalized score.
    3 repo stars
  64. ▌
    Aist Dance Eval · qhjqhj00
    Evaluates the quality, diversity, and music-motion synchronization of generated 3D dance sequences. It probes a model's ability to synthesize physically plausible, choreographically diverse, and rhythm-aligned human motion from audio input. Use when the user wants to benchmark on AIST++, or asks about evaluating this task. Reports FID_k.
    3 repo stars
  65. ▌
    Alden Vrdu Eval · qhjqhj00
    Evaluates vision-language models' ability to actively navigate long, visually rich documents to gather evidence and answer complex queries. It probes multi-turn reasoning, retrieval accuracy, and the effectiveness of direct page-index access versus semantic search. Use when the user wants to benchmark on MMLongBench, LongDocURL, PaperTab, PaperText, FetaTab, DUDE-sub, or asks about evaluating this task. Reports GPT-4o–judged answer accuracy (Acc).
    3 repo stars
  66. ▌
    Alora Peft Eval · qhjqhj00
    Evaluates the effectiveness of dynamic low-rank adaptation (LoRA) for fine-tuning large language models across classification, question answering, and instruction generation tasks. It measures how well rank allocation strategies preserve performance while maintaining or reducing tunable parameter counts. Use when the user wants to benchmark on SQuAD, BoolQ, COPA, ReCoRD, SST-2, RTE, QNLI, Alpaca, MT-Bench, E2E, or asks about evaluating this task. Reports accuracy, GPT-4 score.
    3 repo stars
  67. ▌
    Alpacaeval Eval · qhjqhj00
    Evaluates LLM response quality via pairwise win rates against a baseline, while specifically probing the metric's susceptibility to length bias, gameability via verbosity prompting, and robustness to adversarial truncation. Use when the user wants to benchmark on AlpacaEval, or asks about evaluating this task. Reports Win rate.
    3 repo stars
  68. ▌
    Am Lora Cl Eval · qhjqhj00
    Evaluates a model's ability to continuously learn multiple text classification tasks without catastrophic forgetting, measuring how well it retains knowledge of previous tasks while adapting to new ones. Use when the user wants to benchmark on Standard CL benchmarks, Large number of tasks benchmark, or asks about evaluating this task. Reports average results.
    3 repo stars
  69. ▌
    Anomalygen Eval · qhjqhj00
    This benchmark evaluates log-based anomaly detection models by measuring their ability to classify log sequences as normal or anomalous. It specifically probes how well different model paradigms (classical ML, supervised/unsupervised deep learning, and LLM-based) generalize when trained on code-guided synthetic data augmentation across varying augmentation ratios. Use when the user wants to benchmark on HDFS, Zookeeper, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  70. ▌
    Answersumm Eval · qhjqhj00
    Evaluates multi-perspective answer summarization for community question answering, probing content selection, perspective clustering, abstractive summarization, and factual consistency/coverage. Use when the user wants to benchmark on AnswerSumm, or asks about evaluating this task. Reports F1, ROUGE-1/2/L.
    3 repo stars
  71. ▌
    Ape Prompt Eval · qhjqhj00
    Evaluates the effectiveness of automatically generated prompts (instructions) from the APE framework compared to human-designed or baseline prompts across various natural language processing tasks. Use when the user wants to benchmark on Instruction Induction, BIG-Bench Instruction Induction (BBII), MultiArith, GSM8K, or asks about evaluating this task. Reports zero-shot execution accuracy.
    3 repo stars
  72. ▌
    Apibench Q Eval · qhjqhj00
    Evaluates the retrieval accuracy and ranking quality of query-based API recommendation systems for Java APIs at both class and method levels. It also measures how query reformulation techniques impact recommendation performance. Use when the user wants to benchmark on APIBench-Q, or asks about evaluating this task. Reports Success Rate@k.
    3 repo stars
  73. ▌
    Arabicmmlu Eval · qhjqhj00
    ArabicMMLU probes massive multitask language understanding in Modern Standard Arabic across 40 educational subjects spanning STEM, social sciences, humanities, Arabic language, and culturally specific domains. It evaluates models on cross-lingual transfer, cultural localization, and robustness to linguistic phenomena like negation across primary, middle, high school, and university levels. The benchmark specifically tests how well models handle Arabic-specific knowledge and exam-style multiple-choice questions. Use when the user wants to benchmark on ArabicMMLU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  74. ▌
    Arabicmteb Eval · qhjqhj00
    Evaluates Arabic-centric and cross-lingual text embedding models across multiple linguistic, cultural, and domain-specific capabilities. It probes how well models capture dialectal variations, regional cultural knowledge, and specialized domain terminology in Arabic. Use when the user wants to benchmark on ArabicMTEB, or asks about evaluating this task. Reports Avg..
    3 repo stars
  75. ▌
    Argscichat Eval · qhjqhj00
    Evaluates the ability of dialogue agents to select supportive facts from scientific papers and generate contextually appropriate responses in argumentative scientific dialogues. Probes document-grounded response generation and fact selection under expert-level, opinion-driven interactions. Use when the user wants to benchmark on ArgSciChat, or asks about evaluating this task. Reports Fact-F1.
    3 repo stars
  76. ▌
    Asd And Se Eval · qhjqhj00
    Evaluates a unified audio-visual model's ability to detect which speaker is actively speaking in multi-person video scenes and to enhance speech signals by removing background noise and interference. Use when the user wants to benchmark on AVA-ActiveSpeaker, LRS2, TalkSet, Columbia, MUSAN, or asks about evaluating this task. Reports mAP.
    3 repo stars
  77. ▌
    Astro Mcad Eval · qhjqhj00
    Evaluates a classifier-based anomaly detection pipeline on simulated astronomical transient light curves. It probes the model's ability to identify rare, out-of-distribution events in real-time without prior exposure to the anomalous classes during training. Use when the user wants to benchmark on Simulated LSST-like transient light curves, or asks about evaluating this task. Reports anomaly score.
    3 repo stars
  78. ▌
    Astrochart Eval · qhjqhj00
    Evaluates multimodal large language models' ability to comprehend scientific charts and perform knowledge-intensive reasoning in astronomy. It probes visual understanding, data extraction, numerical calculation, and domain-specific inference. Use when the user wants to benchmark on AstroChart, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  79. ▌
    Atlas Chat Eval · qhjqhj00
    Evaluates large language models on Moroccan Arabic (Darija) across multiple-choice reasoning, instruction following, translation, summarization, and sentiment analysis. It probes dialect-specific linguistic features, script variability (Arabic vs. Arabizi), and real-world instruction-following capabilities in a low-resource setting. Use when the user wants to benchmark on DarijaMMLU, DarijaHellaSwag, Belebele_Ary, DarijaBench, DarijaAlpacaEval, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  80. ▌
    Audiomnist Eval · qhjqhj00
    Evaluates audio classification performance on spoken digits and speaker sex using raw waveforms and spectrograms, serving as a benchmark for explainable AI (XAI) methods in the audio domain. Use when the user wants to benchmark on AudioMNIST, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Auto Split Eval · qhjqhj00
    Evaluates the latency, accuracy, and model size of a collaborative edge-cloud DNN splitting framework (Auto-Split) compared to baselines like QDMP and Neurosurgeon across image classification and object detection tasks. Use when the user wants to benchmark on ImageNet, COCO 2017, or asks about evaluating this task. Reports End-to-end latency (normalized).
    3 repo stars
  82. ▌
    Averimavec Eval · qhjqhj00
    This benchmark probes a model's ability to perform multimodal fact-checking by verifying real-world image-text claims. It requires the system to retrieve cross-modal evidence, analyze inconsistencies, and produce a justified verdict that aligns with ground truth labels. Use when the user wants to benchmark on AVerImaTeC, or asks about evaluating this task. Reports verdict_correctness.
    3 repo stars
  83. ▌
    Babyvision Eval · qhjqhj00
    Evaluates fundamental visual reasoning capabilities in multimodal large language models independent of linguistic priors. It probes early-vision abilities such as visual tracking, spatial perception, fine-grained discrimination, and visual pattern recognition through image-based tasks. Use when the user wants to benchmark on BabyVision, or asks about evaluating this task. Reports Avg@3.
    3 repo stars
  84. ▌
    Basqueglue Eval · qhjqhj00
    This benchmark evaluates Basque language models across classical NLP tasks including topic classification, stance detection, coreference detection, and natural language inference. It measures both task-specific performance and overall linguistic competence in a low-resource agglutinative language setting. Use when the user wants to benchmark on BasqueGLUE, or asks about evaluating this task. Reports Avg.
    3 repo stars
  85. ▌
    Batonvoice Eval · qhjqhj00
    Evaluates a controllable text-to-speech model's ability to generate intelligible speech and accurately convey specific emotional tones based on text instructions. It probes zero-shot cross-lingual generalization and instruction-following capabilities in speech synthesis. Use when the user wants to benchmark on Seed-TTS, Emotion dataset, or asks about evaluating this task. Reports Emotion Classification Accuracy.
    3 repo stars
  86. ▌
    Battleship Eval · qhjqhj00
    Evaluates EFCE solvers on a parametric sequential conflict-resolution game where players place ships and fire shots. It probes the solver's ability to construct incentive-compatible correlation plans that maximize social welfare through deterrence and punishment mechanisms. Use when the user wants to benchmark on Battleship, or asks about evaluating this task. Reports Social Welfare (SW).
    3 repo stars
  87. ▌
    Beans Zero Eval · qhjqhj00
    Evaluates zero-shot generalization of audio-language models on bioacoustic tasks, including species classification, multilabel detection, call-type prediction, lifestage classification, captioning, and individual counting across diverse taxa. Use when the user wants to benchmark on BEANS-Zero, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  88. ▌
    Behavior1k Eval · qhjqhj00
    Evaluates embodied AI agents on long-horizon, human-centered manipulation tasks in a realistic physics-based simulation. It probes the agent's ability to plan and execute complex sequences of action primitives (pick, place, navigate, etc.) while handling rigid, articulated, and deformable objects. Use when the user wants to benchmark on BEHAVIOR-1K, or asks about evaluating this task. Reports task success rate.
    3 repo stars
  89. ▌
    Bench Push Eval · qhjqhj00
    Evaluates the transferability and performance of reinforcement learning policies for mobile robot navigation and pushing-based manipulation tasks. It probes how well policies trained in simulation handle real-world sim-to-real gaps, clutter, and sparse rewards across varying obstacle densities. Use when the user wants to benchmark on Bench-Push (Maze & Box-Delivery), or asks about evaluating this task. Reports $S_{\text{manip}}$.
    3 repo stars
  90. ▌
    Benchie Fl Eval · qhjqhj00
    Evaluates Open Information Extraction (OIE) systems on their ability to extract fact-based triples from text. It uses a conservative exact-matching function with synset-based clustering to penalize non-informative copies and reward precise fact extraction, while also measuring correlation with downstream QA and knowledge base tasks. Use when the user wants to benchmark on BenchIE^FL, or asks about evaluating this task. Reports exact-match.
    3 repo stars
  91. ▌
    Ber Performance · qhjqhj00
    Evaluates the bit error rate (BER) performance of a Reconfigurable Intelligent Surface (RIS) aided spatial media-based modulation system compared to baseline schemes (SM, MBM, QSM) under uncorrelated Rayleigh fading channels. Use when the user has predictions and gold and needs to compute Bit Error Rate (BER).
    3 repo stars
  92. ▌
    Binaryhingeloss · qhjqhj00
    Compute the BinaryHingeLoss metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryHingeLoss, or asks how to score with BinaryHingeLoss.
    3 repo stars
  93. ▌
    Binaryprecision · qhjqhj00
    Compute the BinaryPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryPrecision, or asks how to score with BinaryPrecision.
    3 repo stars
  94. ▌
    Biobert QA Eval · qhjqhj00
    Evaluates factoid question answering performance on small biomedical datasets, testing transfer learning effectiveness from domain-specific pre-training. It measures how well the model retrieves exact or lenient answers to biomedical queries. Use when the user wants to benchmark on BioASQ 4b, BioASQ 5b, BioASQ 6b, or asks about evaluating this task. Reports Mean Reciprocal Rank (MRR).
    3 repo stars
  95. ▌
    Biobert Re Eval · qhjqhj00
    Tests a model's capability to extract biomedical relations (gene-disease, gene-chemical) from text using minimal task-specific modifications. It probes whether domain-specific pre-training improves relation classification on small-scale biomedical corpora. Use when the user wants to benchmark on GAD, EU-ADR, CHEMPROT, or asks about evaluating this task. Reports entity-level F1.
    3 repo stars
  96. ▌
    Biomed Vqa Eval · qhjqhj00
    Evaluates the domain-adaptive post-training of multimodal large language models on biomedical visual question answering tasks, measuring how well models generalize to specialized medical domains using both open and closed evaluation splits. Use when the user wants to benchmark on SLAKE, PathVQA, VQA-RAD, PMC-VQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  97. ▌
    Bionli 300 Eval · qhjqhj00
    This evaluation probes a model's ability to verify scientific claims against provided or retrieved evidence in a binary classification setting. It measures performance on Supported vs. Refuted labels, testing factual grounding, uncertainty calibration, and the impact of atomic decomposition and web corroboration. Use when the user wants to benchmark on BIONLI-300, or asks about evaluating this task. Reports Balanced Accuracy.
    3 repo stars
  98. ▌
    Bioscan 5m Eval · qhjqhj00
    Evaluates models on insect biodiversity monitoring by testing closed-world species identification, open-world genus-level grouping for novel species, and zero-shot clustering of multimodal embeddings against taxonomic ground truth. Use when the user wants to benchmark on BIOSCAN-5M, or asks about evaluating this task. Reports Fine-tuned accuracy.
    3 repo stars
  99. ▌
    Bird Bench Eval · qhjqhj00
    Evaluates the capability of large language models to generate correct SQL queries from natural language questions. It specifically probes how annotation noise and errors in benchmark datasets affect model performance and reliability. Use when the user wants to benchmark on BIRD-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  100. ▌
    Bouquet Mt Eval · qhjqhj00
    Evaluates machine translation systems on a contamination-free, multilingual dataset covering diverse domains and registers. It measures translation quality at both sentence and paragraph levels to assess how well models handle linguistic diversity and cultural authenticity across 8 major languages. Use when the user wants to benchmark on BOUQuET, or asks about evaluating this task. Reports CometKiwi.
    3 repo stars