qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Session Rec Rnn Eval · qhjqhj00Evaluates a model's ability to predict the next item in a user session based on sequential click or watch history. It probes the capability to capture short-term sequential dependencies and maintain context within session-based recommendation scenarios. Use when the user wants to benchmark on RSC15, VIDEO, or asks about evaluating this task. Reports recall@20.
- ▌ Shawshank Bench Eval · qhjqhj00Evaluates the vulnerability of embodied AI agents (Vision-Language Models) to indirect environmental jailbreaks, where malicious instructions are physically embedded in the environment (e.g., on walls or tables) rather than provided as direct text prompts. It measures both the agent's susceptibility to harmful behavior and the collateral impact on benign task execution. Use when the user wants to benchmark on Shawshank-Bench, or asks about evaluating this task. Reports ASR.
- ▌ Showdown Clicks Eval · qhjqhj00Evaluates a vision-language model's ability to accurately predict the correct UI element to click based on a screenshot and a task instruction. It isolates low-level visual grounding and interaction skills in ambiguous or icon-heavy interfaces. Use when the user wants to benchmark on Showdown-Clicks, or asks about evaluating this task. Reports Top-1 Accuracy.
- ▌ Skilllearnbench Eval · qhjqhj00This benchmark evaluates continual learning methods for generating reusable procedural skills in LLM agents. It probes the quality of generated skills, their alignment with execution trajectories, and the ultimate task-solving accuracy and efficiency of a fixed solving agent. Use when the user wants to benchmark on SkillLearnBench, or asks about evaluating this task. Reports Acc..
- ▌ Smoldocling Doc Eval · qhjqhj00This evaluation probes a vision-language model's ability to perform end-to-end document conversion, including text recognition, layout analysis, table and chart structure extraction, and code/formula parsing. It measures how accurately the model reconstructs document content and spatial structure from page images into standardized markup formats. Use when the user wants to benchmark on DocLayNet, SynthCodeNet, Im2Latex-230k, FinTabNet, PubTables-1M, or asks about evaluating this task. Reports mAP@0.5:0.95, TEDS.
- ▌ Smtm Mobile Cnn Eval · qhjqhj00Evaluates the latency reduction, accuracy loss, memory overhead, energy saving, and early exit performance of a semantic memory caching mechanism (SMTM) for accelerating CNN inference on mobile devices. Use when the user wants to benchmark on UCF101, CIFAR-100 (long-tail), or asks about evaluating this task. Reports latency reduction.
- ▌ Snn Dfe Optical Eval · qhjqhj00Evaluates the communication performance and hardware efficiency of Spiking Neural Network-based Decision-Feedback Equalizers (DFEs) for optical channels compared to traditional Artificial Neural Network baselines. It probes the trade-off between bit error rate, computational complexity, and energy efficiency under varying quantization levels and FPGA resource constraints. Use when the user wants to benchmark on Custom Optical Communication DFE Benchmark, or asks about evaluating this task. Reports BER.
- ▌ Social Chem 101 Eval · qhjqhj00Evaluates a model's ability to reason about and generate rules-of-thumb (RoTs) that capture social and moral norms across 12 distinct dimensions of judgment, such as cultural pressure, legality, and moral foundations. It probes whether neural models can produce attribute-aware, context-sensitive normative judgments for unseen social scenarios. Use when the user wants to benchmark on Social-Chem-101, or asks about evaluating this task. Reports micro-F1.
- ▌ Spearman Correlation · qhjqhj00Measures the rank-order agreement between an NLG evaluator's predicted scores and human reference judgments. It probes the evaluator's ability to capture task-specific quality dimensions such as fluency, coherence, consistency, and groundedness across summarization and dialogue generation. Use when the user has predictions and gold and needs to compute Spearman correlation.
- ▌ Spec Compliance Eval · qhjqhj00Evaluates whether frontier LLM responses adhere to their published model specifications when faced with generated value tradeoff scenarios. It also measures the consistency of model-based judges in detecting specification violations and identifies specification flaws like contradictions and ambiguities. Use when the user wants to benchmark on Generated Value Tradeoff Scenarios, or asks about evaluating this task. Reports compliance.
- ▌ Speech Commands Eval · qhjqhj00This benchmark evaluates limited-vocabulary keyword spotting models for on-device speech recognition. It probes a model's ability to correctly identify isolated spoken words from a fixed set of 10 commands, while also handling background silence and unrecognized speech in both aligned and continuous streaming audio contexts. Use when the user wants to benchmark on Speech Commands, or asks about evaluating this task. Reports Top-One Error.
- ▌ Speech Df Arena Eval · qhjqhj00This benchmark evaluates the robustness and cross-domain generalization of speech deepfake detection models across diverse synthetic speech generation techniques, including TTS, voice conversion, neural codecs, and real-world social media leaks. It measures how well models maintain performance when faced with unseen attack types, languages, and distribution shifts. Use when the user wants to benchmark on ASVspoof 2019, ASVspoof 2021, ASVspoof 2024, ADD 2022, ADD 2023, CodecFake, LibriSeVoc, SONAR, Fake or Real (FoR), DFADD, In-the-wild, or asks about evaluating this task. Reports EER.
- ▌ Spice Reasoning Eval · qhjqhj00This evaluation probes a model's ability to solve challenging mathematical and general reasoning tasks, both from standard benchmarks and document-grounded self-play generated questions. It measures how well the model can extract information, perform multi-step logical deduction, and produce verifiable answers across diverse academic and competition-level datasets. Use when the user wants to benchmark on MATH-500, OlympiadBench, Minerva Math, GSM8K, AMC, AIME'24, AIME'25, SuperGPQA, GPQA-Diamond, MMLU-Pro, BBEH, or asks about evaluating this task. Reports pass rate.
- ▌ Split Computing Eval · qhjqhj00Evaluates a multi-task supervised compression model for split computing across image classification, object detection, and semantic segmentation. It measures predictive accuracy alongside system-level metrics like end-to-end latency and energy consumption on resource-constrained edge devices with simulated wireless links. Use when the user wants to benchmark on ILSVRC 2012, COCO 2017, PASCAL VOC 2012, or asks about evaluating this task. Reports model accuracy.
- ▌ Starcoder2 Code Eval · qhjqhj00Evaluates code generation, completion, and bug-fixing capabilities across multiple programming languages and libraries. It probes a model's ability to write correct functions from prompts, translate code across languages, and fix existing buggy code using standard and enhanced benchmarks. Use when the user wants to benchmark on HumanEval, MBPP, EvalPlus, MultiPL-E, DS-1000, HumanEvalFix, or asks about evaluating this task. Reports pass@1.
- ▌ Stereo Ycb V Ds Eval · qhjqhj00Evaluates 6D object pose estimation methods under stereo vision conditions, specifically probing robustness to occlusion and scale ambiguity by leveraging dense 2D-3D correspondences and stereo disparity. Use when the user wants to benchmark on Stereo PBR YCB-V DS, or asks about evaluating this task. Reports ADD0.1.
- ▌ Supervised Sadp Eval · qhjqhj00Evaluates the ability of a gradient-free, locally-updated spiking neural network to classify images using population-level spike agreement metrics. It probes whether replacing backpropagation with supervised Spike Agreement-Dependent Plasticity (SADP) and Cohen’s κ can achieve competitive vision and biomedical classification performance while maintaining biological plausibility and hardware compatibility. Use when the user wants to benchmark on MNIST, Fashion-MNIST, CIFAR-10, LC25000, Brain MRI Tumor, or asks about evaluating this task. Reports accuracy.
- ▌ Surgery Mapping Eval · qhjqhj00Evaluates the correctness and efficiency of gradient-based versus boolean logic-based methods for identifying feature-parameter interactions and transferring trained weights when new features are added to a reinforcement learning model. It measures how well each mapping technique preserves model performance and computational speed during architectural surgery. Use when the user has predictions and gold and needs to compute Interactions Found.
- ▌ Suvach Hindi QA Eval · qhjqhj00Evaluates Hindi extractive question answering capabilities using multiple-choice questions generated from Wikipedia contexts. It probes a model's ability to comprehend Hindi text, locate relevant information, and select the correct answer from four options under varying context availability settings. Use when the user wants to benchmark on Suvach, or asks about evaluating this task. Reports accuracy.
- ▌ Sygus Comp 2016 Eval · qhjqhj00Evaluates syntax-guided program synthesis solvers on their ability to generate correct programs from logical constraints and grammars. It probes capabilities in conditional linear integer arithmetic, invariant generation, and programming-by-example with bit-vectors and strings. Use when the user wants to benchmark on SyGuS-Comp 2016, or asks about evaluating this task. Reports number_of_benchmarks_solved.
- ▌ Sygus Comp 2017 Eval · qhjqhj00Evaluates the ability of synthesis solvers to generate correct programs or expressions that satisfy given grammatical and semantic constraints across multiple domain-specific tracks. Use when the user wants to benchmark on SyGuS-Comp 2017, or asks about evaluating this task. Reports correctness.
- ▌ Sygus Comp 2018 Eval · qhjqhj00Evaluates syntax-guided synthesis solvers on their ability to generate correct programs or specifications across multiple domains, including general synthesis, conditional linear integer arithmetic, invariant generation, and programming by examples. Use when the user wants to benchmark on SyGuS-Comp 2018, or asks about evaluating this task. Reports correctness.
- ▌ Synthref Refvos Eval · qhjqhj00Evaluates the effectiveness of a synthetic referring expression dataset for training language-guided video object segmentation models. It measures segmentation accuracy when models are trained on synthetic versus human annotations and evaluated on standard referring video segmentation benchmarks. Use when the user wants to benchmark on DAVIS-2017, Refer-YouTube-VOS, or asks about evaluating this task. Reports J&F.
- ▌ Text Clustering Eval · qhjqhj00Evaluates the ability of centroid-based clustering algorithms to group unlabeled text documents into semantically coherent clusters. It measures clustering accuracy, label alignment with ground truth, and how closely learned centroids match true cluster centers. Use when the user wants to benchmark on Bank77, CLINC, GoEmo, MASSIVE, StackExchange, or asks about evaluating this task. Reports ACC, NMI.
- ▌ Thiomi Baseline Eval · qhjqhj00Evaluates the quality and utility of a multimodal corpus for low-resource African languages. It does so by training and testing baseline models for automatic speech recognition, machine translation, and text-to-speech across multiple languages. Use when the user wants to benchmark on Thiomi Dataset, or asks about evaluating this task. Reports WER.
- ▌ Timeseries Exam Eval · qhjqhj00This benchmark evaluates vision-language models' ability to reason over time series data across medical, financial, and meteorological domains. It probes pattern recognition, anomaly detection, and causal reasoning using synthetically generated multiple-choice questions derived from real-world datasets. Use when the user wants to benchmark on PTB-XL, MIT-BIH, MIMIC-IV Waveform, Yahoo Finance, WeatherBench 2, or asks about evaluating this task. Reports accuracy.
- ▌ Top K Accuracy Score · qhjqhj00Compute the top_k_accuracy_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute top_k_accuracy_score, or asks how to score with top_k_accuracy_score.
- ▌ Tox21 Challenge Eval · qhjqhj00Evaluates molecular toxicity prediction capabilities across diverse AI architectures (descriptor-based models, neural networks, tabular transformers, and zero-shot LLMs) on a standardized chemical safety benchmark. Use when the user wants to benchmark on Tox21 Challenge dataset, or asks about evaluating this task. Reports performance.
- ▌ Track Any State Eval · qhjqhj00Evaluates a model's ability to track objects through appearance-changing state transformations and explicitly model those transformations as a state graph. It probes spatiotemporal continuity, zero-shot object recovery, and semantic reasoning about object interactions. Use when the user wants to benchmark on VOST, VSCOS, M3-VOS, DAVIS 2017, VOST-TAS, or asks about evaluating this task. Reports Jaccard (J).
- ▌ Trade The Event Eval · qhjqhj00This evaluation probes a model's ability to detect objective corporate events in financial news and translate those detections into actionable, timely trading signals. It measures how effectively the detected events predict short-term stock price movements and generate excess returns compared to a market benchmark. Use when the user wants to benchmark on EDT, or asks about evaluating this task. Reports Winning Rate.
- ▌ Tts Nl Guidance Eval · qhjqhj00Evaluates a text-to-speech model's ability to generate audio that matches natural language descriptions of speaker attributes (gender, accent, pitch, speaking rate, recording quality) and overall audio fidelity. It measures both objective acoustic metrics and subjective human ratings of relevance and naturalness. Use when the user wants to benchmark on MLS, LibriTTS-R, or asks about evaluating this task. Reports MOS.
- ▌ Tvm Dl Compiler Eval · qhjqhj00Evaluates the end-to-end performance and optimization capability of a deep learning compiler across diverse hardware back-ends (GPU, CPU, embedded GPU, FPGA) on standard inference workloads. It measures how effectively the compiler automatically generates high-performance kernels compared to hand-tuned vendor libraries and existing frameworks. Use when the user wants to benchmark on DL Inference Workloads (ResNet-18, MobileNet, LSTM, DQN, DCGAN), or asks about evaluating this task. Reports speedup.
- ▌ Tweediedeviancescore · qhjqhj00Compute the TweedieDevianceScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TweedieDevianceScore, or asks how to score with TweedieDevianceScore.
- ▌ Ultraeval Audio Eval · qhjqhj00Evaluates audio foundation models across understanding, generation, and codec capabilities. It probes semantic accuracy, timbre fidelity, acoustic quality, and multilingual speech comprehension using a unified taxonomy and standardized benchmarks. Use when the user wants to benchmark on SpeechCMMLU, SpeechHSK, LibriSpeech, AISHELL-1, or asks about evaluating this task. Reports WER.
- ▌ Unclearinstruct Eval · qhjqhj00Evaluates a robot's ability to infer human goals and execute household tasks from noisy, accented, or mispronounced spoken instructions. It probes robust speech perception, joint planning, and Theory of Mind in embodied human-robot collaboration under mixed-observability conditions. Use when the user wants to benchmark on UnclearInstruct, or asks about evaluating this task. Reports Accuracy.
- ▌ Vdr Compression Eval · qhjqhj00Evaluates multi-vector visual document retrieval (VDR) models under varying compression ratios. It measures how well pruning and merging strategies maintain retrieval accuracy while reducing storage and computational overhead. Use when the user wants to benchmark on ViDoRe-V1, ViDoRe-V2, JinaVDR-Bench, REAL-MM-RAG, ViDoSeek, MMLongBench-Doc, or asks about evaluating this task. Reports nDCG@5.
- ▌ Video Analytics Eval · qhjqhj00Evaluates long-video understanding and agentic retrieval capabilities by testing models on temporal grounding, summarization, reasoning, and event causality across ultra-long video streams. Use when the user wants to benchmark on LVBench, VideoMME-Long, Ava-100, or asks about evaluating this task. Reports accuracy.
- ▌ Videoconviction Eval · qhjqhj00Evaluates whether LLMs and MLLMs can accurately extract stock tickers, identify explicit investment actions, and quantify human conviction levels from financial influencer videos and transcripts. It probes multimodal reasoning, financial domain understanding, and the ability to filter out noisy or promotional content. Use when the user wants to benchmark on VideoConviction, or asks about evaluating this task. Reports F1 score.
- ▌ Vision Arch Gen Eval · qhjqhj00Evaluates the classification performance of LLM-generated neural network architectures by training each for a single epoch on seven computer vision benchmarks. It measures Top-1 accuracy to assess architectural quality, while also tracking generation efficiency via hash validation speed and duplicate rejection rates. Use when the user wants to benchmark on MNIST, CelebA-Gender, CIFAR-10, CIFAR-100, ImageNette, SVHN, Places365, or asks about evaluating this task. Reports Top-1 accuracy after 1 epoch.
- ▌ Vlm Subtlebench Eval · qhjqhj00This benchmark evaluates vision-language models' ability to perform subtle comparative reasoning between pairs of images. It probes capabilities across ten fine-grained difference types, including spatial, temporal, viewpoint, attribute, and existence changes, requiring models to detect and explain nuanced visual discrepancies that are often missed by standard prompting or simple image concatenation. Use when the user wants to benchmark on VLM-SubtleBench, or asks about evaluating this task. Reports accuracy.
- ▌ Voiceagentbench Eval · qhjqhj00Evaluates speech language models and ASR-LLM pipelines on agentic speech tasks. It probes single/multi-tool orchestration, multi-turn dialogue, and safety refusal capabilities across multiple languages, including English, Hindi, and five Indic languages. Use when the user wants to benchmark on VoiceAgentBench, or asks about evaluating this task. Reports PF (Parameter Filling).
- ▌ Vtab Md Fewshot Eval · qhjqhj00Evaluates few-shot classification performance across diverse visual domains by comparing transfer learning and meta-learning approaches on a unified benchmark combining VTAB and Meta-Dataset. Use when the user wants to benchmark on VTAB+MD, or asks about evaluating this task. Reports accuracy.
- ▌ Wave Unet Musdb Eval · qhjqhj00Evaluates end-to-end time-domain audio source separation models on singing voice and multi-instrument separation tasks. Probes the model's ability to isolate specific audio sources from mixed recordings using raw waveform inputs. Use when the user wants to benchmark on MUSDB, CCMixter, or asks about evaluating this task. Reports MSE.
- ▌ Widspeech Bench Eval · qhjqhj00Evaluates end-to-end speech-to-speech (S2S) language models on real-world conversational tasks, probing their ability to handle diverse query types, paralinguistic features (prosody, disfluencies), and robustness to background noise. Use when the user wants to benchmark on WildSpeech-Bench, or asks about evaluating this task. Reports Score.
- ▌ Wsd Rp Accuracy Eval · qhjqhj00Evaluates a model's ability to disambiguate word senses for both common nouns and proper nouns exhibiting regular polysemy. It probes contextual understanding and the capacity to leverage structured sense glosses and dot-object type classes to select the correct meaning from a candidate inventory. Use when the user wants to benchmark on WSD dataset (CWN 2.0), RP dataset (Revised Mandarin Chinese Dictionary), or asks about evaluating this task. Reports accuracy.
- ▌ X Webagentbench Eval · qhjqhj00Evaluates LLM-based agents' ability to comprehend multilingual shopping instructions and successfully navigate interactive web environments across 14 languages. Use when the user wants to benchmark on X-WebAgentBench, or asks about evaluating this task. Reports Task Score.
- ▌ Xray Report Gen Eval · qhjqhj00Evaluates the capability of vision-language models to generate accurate and clinically relevant radiology reports from chest X-ray images. It probes both linguistic fluency and coverage against reference reports, as well as clinical accuracy in identifying pathological findings. Use when the user wants to benchmark on IU X-ray, MIMIC-CXR, CheXpert Plus, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Yoloe Lvis Coco Eval · qhjqhj00Evaluates open-vocabulary object detection and segmentation capabilities using text, visual, and prompt-free inputs on zero-shot and fine-tuned settings. Use when the user wants to benchmark on LVIS, COCO, or asks about evaluating this task. Reports Fixed AP.
- ▌ Zenbrain Memory Eval · qhjqhj00This evaluation protocol assesses the long-term memory and retrieval capabilities of autonomous AI systems. It measures how well models retain, route, and retrieve information across multiple sessions and varying context lengths, while also evaluating the quality of generated answers using LLM-as-a-judge scoring. Use when the user wants to benchmark on LoCoMo (Real-LoCoMo pool), LongMemEval-S, MemoryAgentBench, MemoryArena, or asks about evaluating this task. Reports NDCG@5.
- ▌ Zipvoice Dialog Eval · qhjqhj00This benchmark evaluates non-autoregressive spoken dialogue generation models on their ability to produce multi-turn conversational audio that matches input text, maintains speaker identity, and accurately handles turn-taking between two speakers. It probes both objective speech quality metrics and subjective human judgments of coherence and similarity. Use when the user wants to benchmark on test-dialog-zh, test-dialog-en, or asks about evaluating this task. Reports cpWER.
- ▌ Zs Xlt News Rec Eval · qhjqhj00Evaluates zero-shot cross-lingual news recommendation by measuring how effectively a model recommends articles in a target language to users who only consume news in a source language. It probes the model's ability to leverage multilingual sentence embeddings and click behavior fusion without task-specific fine-tuning on the target language. Use when the user wants to benchmark on MIND (small) / xMIND (small), or asks about evaluating this task. Reports nDCG@10.
- ▌ 3d LLM Benchmark Eval · qhjqhj00Evaluates whether Vision-Language Models (VLMs) and 3D LLMs genuinely understand 3D spatial reasoning or merely exploit 2D visual priors by rendering point clouds into images. It probes capabilities like object captioning, scene question-answering, and situation understanding across single-view, multi-view, and oracle-viewpoint settings. Use when the user wants to benchmark on 3D MM-Vet, ObjaverseXL-LVIS Caption, ScanQA, SQA3D, or asks about evaluating this task. Reports LLM-eval, EM.
- ▌ 5g Madrl Sumrate Eval · qhjqhj00Evaluates the ability of a multi-agent deep reinforcement learning framework to optimize the 3D placement and trajectory of mobile access points in dynamic 5G networks, balancing sum-rate maximization against user mobility and interference. Use when the user wants to benchmark on Custom 5G Network Simulation, or asks about evaluating this task. Reports sum-rate.
- ▌ Activitynet Comp Eval · qhjqhj00Evaluates fine-grained temporal and compositional alignment in video-text models by testing their ability to distinguish between videos and captions that contain subtle structural disruptions. It probes sensitivity to temporal reordering, action word replacement, and segment-level misalignment, as well as the model's robustness to combined disruptions. Use when the user wants to benchmark on ActivityNet-Comp, YouCook2-Comp, or asks about evaluating this task. Reports binary classification accuracy.
- ▌ Agad Medical Auc Eval · qhjqhj00Evaluates a generative anomaly detection model's ability to distinguish normal from abnormal medical images using pseudo-anomaly generation and self-contrast learning. It probes robustness on fine-grained, real-world medical imaging data with limited anomaly supervision. Use when the user wants to benchmark on Alzheimer's Dataset Dubey (2019), ChestXray Kermany et al. (2018), Lung Histopathology (LC25000 subset), Retinal OCT Kermany et al. (2018), or asks about evaluating this task. Reports AUC.
- ▌ Agentrewardbench Eval · qhjqhj00This benchmark evaluates the effectiveness of LLM-based judges in automatically assessing web agent trajectories. It probes the judges' ability to correctly predict task success, detect side effects, and identify repetitive actions by comparing their outputs against expert human annotations. Use when the user wants to benchmark on AgentRewardBench, or asks about evaluating this task. Reports precision.
- ▌ Aibench Scenario Eval · qhjqhj00Evaluates the end-to-end system-level performance and tail latency of AI-driven online services by simulating real-world user workloads. It probes how cascading interactions between AI and non-AI components affect overall service quality, and tests the validity of statistical queueing models for predicting system latency. Use when the user wants to benchmark on AIBench Scenario (E-commerce & Translation Intelligence), or asks about evaluating this task. Reports latency (avg, p90, p99).
- ▌ Aibench Training Eval · qhjqhj00Evaluates AI training workloads by measuring model complexity, computational cost, convergence rate, and micro-architectural behavior to assess benchmark diversity, repeatability, and cost against industry standards like MLPerf. Use when the user wants to benchmark on AIBench Training, or asks about evaluating this task. Reports convergent_rate.
- ▌ Aj Jaggedness Penalty · qhjqhj00Proposes a diagnostic metric to quantify the discrepancy between macro-averaged benchmark scores and experienced reliability in deployment. It accounts for uneven task coverage by measuring the dispersion of error rates across domains, highlighting how gap-uniform evaluation can misstate actual user experience. Use when the user has predictions and gold and needs to compute CV_d(e_d).
- ▌ Albertgong1 My Metric · qhjqhj00Compute albertgong1/my_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of albertgong1/my_metric.
- ▌ Algerian Dialect Eval · qhjqhj00Evaluates cross-lingual and cross-script transfer performance for sentiment analysis and topic classification on a novel multi-layer Algerian dialect corpus. Probes how script differences (Latin/NArabizi vs. Arabic/Persian/Urdu) and typological similarity impact classification accuracy in code-switched, under-resourced vernaculars. Use when the user wants to benchmark on Algerian Dialect Corpus (NArabizi), or asks about evaluating this task. Reports Macro F1.
- ▌ Amazon Stark Skb Eval · qhjqhj00Evaluates the ability of neural retriever-reranker pipelines to accurately retrieve relevant product entities from semi-structured e-commerce knowledge graphs using natural language queries. It probes semantic matching, cross-encoder reranking effectiveness, and the impact of graph-based augmentation on retrieval precision and recall. Use when the user wants to benchmark on Amazon STaRK SKB, or asks about evaluating this task. Reports Hit@1.
- ▌ Answer Switching Rate · qhjqhj00Measures the causal influence of activation-based interventions (linear directions or multidimensional cones) on an LLM's factual reasoning. It quantifies how effectively steering or ablating specific neural subspaces switches model outputs from truthful to untruthful across a set of propositional prompts. Use when the user has predictions and gold and needs to compute Answer Switching Rate (ASR).
- ▌ Argoverse2 Waymo Eval · qhjqhj00Evaluates the plausibility, diversity, and kinematic consistency of generated multi-agent traffic scene continuations conditioned on 5-second histories. It probes a model's ability to generate realistic, diverse, and physically plausible future trajectories over a 6-second horizon, including out-of-distribution generalization across different autonomous driving datasets. Use when the user wants to benchmark on Argoverse 2 (A2), Waymo (WO), or asks about evaluating this task. Reports minADE.
- ▌ Arkts Codesearch Eval · qhjqhj00Evaluates code embedding models on a semantic code retrieval task where the goal is to find the correct ArkTS function given a natural language docstring or comment. It probes the model's ability to align bilingual documentation with declarative UI and distributed application code semantics. Use when the user wants to benchmark on ArkTS-CodeSearch, or asks about evaluating this task. Reports MRR.
- ▌ Audiomotionbench Eval · qhjqhj00Evaluates large audio-language models' ability to perceive and reason about spatial motion in binaural audio. It probes whether models can correctly infer motion direction and trajectories from interaural cues, rather than relying on linguistic or spectral heuristics. Use when the user wants to benchmark on AudioMotionBench, or asks about evaluating this task. Reports accuracy.
- ▌ Audiosafetybench Eval · qhjqhj00Evaluates audio safety guardrails on detecting both audio-native risks (e.g., harmful sound events, voice attributes) and semantic content risks (e.g., jailbreaks, policy violations). It measures joint accuracy where both risk types must be correctly classified, alongside end-to-end inference latency. Use when the user wants to benchmark on AudioSafetyBench, Jailbreak-AudioBench, Nemotron-Content-Safety-Audio, Omni-SafetyBench, AdvWave, or asks about evaluating this task. Reports accuracy.
- ▌ Av Odyssey Bench Eval · qhjqhj00Evaluates multimodal large language models' ability to perceive, integrate, and reason over interleaved audio and visual inputs. It probes basic auditory perception (e.g., loudness, pitch, duration) and complex cross-modal tasks spanning timbre, tone, melody, spatial reasoning, temporal dynamics, hallucination detection, and intricate reasoning across 10 domains. Use when the user wants to benchmark on AV-Odyssey Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Baleegh Fluency Score · qhjqhj00Compute Baleegh/Fluency_Score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Baleegh/Fluency_Score.
- ▌ Bangla Sentiment Eval · qhjqhj00Evaluates the capability of NLP models to classify sentiment in Bangla text. It compares classical machine learning, CNN, FastText, and transformer-based architectures to determine which model family performs best on low-resource Bangla sentiment tasks. Use when the user wants to benchmark on Multiple publicly available Bangla sentiment datasets, or asks about evaluating this task. Reports accuracy.
- ▌ Bars Recommender Eval · qhjqhj00Evaluates the reproducibility and standardization of evaluation protocols in recommender systems. It probes both candidate item matching (ranking) and click-through rate (CTR) prediction tasks using standardized data splits, hyperparameter configurations, and common industry metrics to ensure fair and comparable model performance. Use when the user wants to benchmark on Criteo, MovieLens, or asks about evaluating this task. Reports NDCG@K, AUC.
- ▌ Binaryconfusionmatrix · qhjqhj00Compute the BinaryConfusionMatrix metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryConfusionMatrix, or asks how to score with BinaryConfusionMatrix.
- ▌ Binaryhammingdistance · qhjqhj00Compute the BinaryHammingDistance metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryHammingDistance, or asks how to score with BinaryHammingDistance.
- ▌ Blmore Challenge Eval · qhjqhj00Evaluates multimodal emotion recognition systems on detecting blended emotions (presence and salience) across unseen actors. It probes the model's ability to generalize actor-invariant emotional semantics from audio, visual, and combined modalities under a strict threshold-based discretization protocol. Use when the user wants to benchmark on BLEMORE, or asks about evaluating this task. Reports Score.
- ▌ Bmdal Regression Eval · qhjqhj00Evaluates the sample efficiency and predictive accuracy of batch-mode deep active learning methods for tabular regression tasks. It probes how well different kernel-based selection strategies reduce prediction error over sequential labeling rounds compared to random sampling. Use when the user wants to benchmark on UCI & OpenML Tabular Regression Benchmark, or asks about evaluating this task. Reports RMSE.
- ▌ Bongard Rwr Plus Eval · qhjqhj00Evaluates vision-language models' ability to perform abstract visual reasoning (AVR) by recognizing fine-grained, abstract visual concepts in Bongard-style matrix problems. It probes capabilities in concept selection, image-to-side classification, and free-form concept description generation. Use when the user wants to benchmark on Bongard-RWR+, or asks about evaluating this task. Reports accuracy.
- ▌ Brats Robustness Eval · qhjqhj00Evaluates the robustness and generalization capability of brain tumor segmentation models when faced with distribution shifts, specifically Gaussian noise perturbations in MRI scans. It probes whether high benchmark accuracy translates to reliable performance on clinically realistic, noisy data rather than just overfitting to clean benchmark distributions. Use when the user wants to benchmark on BraTS2018, or asks about evaluating this task. Reports Dice score.
- ▌ Browsesafe Bench Eval · qhjqhj00Evaluates AI browser agents' ability to detect prompt injection attacks embedded in complex, realistic HTML environments. It probes whether models can distinguish malicious intent from benign distractors across diverse attack types, injection strategies, and linguistic styles. Use when the user wants to benchmark on BrowseSafe-Bench, or asks about evaluating this task. Reports balanced accuracy.
- ▌ Buelfhood Fbeta Score · qhjqhj00Compute buelfhood/fbeta_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of buelfhood/fbeta_score.
- ▌ Calibrated Similarity · qhjqhj00Measures the semantic novelty of LLM-generated text by quantifying its similarity to the closest segment in the model's pretraining corpus. It probes whether models merely reproduce memorized training data or generalize to produce compositionally distinct outputs. Use when the user has predictions and gold and needs to compute calibrated similarity.
- ▌ Calinskiharabaszscore · qhjqhj00Compute the CalinskiHarabaszScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CalinskiHarabaszScore, or asks how to score with CalinskiHarabaszScore.
- ▌ Camelyon17 Wilds Eval · qhjqhj00Evaluates domain generalization capability for metastatic breast cancer detection in histopathology images. The protocol trains models on data from known medical centers and tests them on completely unseen centers with different staining protocols and scanners to measure adaptability. Use when the user wants to benchmark on Camelyon17 WILDS, or asks about evaluating this task. Reports accuracy (%).
- ▌ Cameramotion Vqa Eval · qhjqhj00Evaluates fine-grained camera motion recognition in VideoLLMs using multiple-choice questions and multi-label classification. Probes whether models can distinguish geometric camera movements from object motion and static shots. Use when the user wants to benchmark on CameraMotionVQA, CameraMotionDataset, or asks about evaluating this task. Reports answer accuracy.
- ▌ Capx Adversarial Eval · qhjqhj00Evaluates the robustness of classifiers against universal and individual adversarial perturbations under domain-specific linear constraints. Probes the trade-off between attack success rate and computational efficiency across finance, network security, medical IoT, and cyber-physical systems. Use when the user wants to benchmark on LCLD, IDS, IoMT, SWaT, WADI, or asks about evaluating this task. Reports ASR.
- ▌ Caqa Attribution Eval · qhjqhj00Evaluates the quality and validity of citations in generated answers for complex question answering. It probes whether models can correctly classify attributions as supportive, insufficient, contradictory, or irrelevant, and assesses their ability to handle varying reasoning complexities. Use when the user wants to benchmark on CAQA, ACLE-Manual, or asks about evaluating this task. Reports micro-F1.
- ▌ CD Fer Benchmark Eval · qhjqhj00Evaluates cross-domain facial expression recognition (CD-FER) models by measuring how well they transfer learned features from a labeled source dataset to an unlabeled target dataset. It probes the model's ability to learn domain-invariant representations and adapt to distribution shifts across different facial expression datasets. Use when the user wants to benchmark on RAF-DB, AFE, CK+, JAFFE, SFEW2.0, FER2013, ExpW, or asks about evaluating this task. Reports accuracy.
- ▌ Change Detection Eval · qhjqhj00Evaluates a model's ability to detect semantic building changes between two temporally separated remote sensing images. It probes robustness to weak temporal supervision, label noise, and out-of-domain generalization in large-scale urban environments. Use when the user wants to benchmark on b-FLAIR-test, b-FLAIR-test-spot, LEVIR-CD, WHUCD, S2Looking, or asks about evaluating this task. Reports F1-score (F1), Intersection over Union (IoU).
- ▌ Chestxray14 Bias Eval · qhjqhj00Evaluates the effectiveness of attribute-neutralization models in removing demographic bias (sex and age) from chest X-ray images while preserving diagnostic utility for 15 disease findings. It probes the trade-off between demographic leakage suppression and clinical performance across varying edit intensities. Use when the user wants to benchmark on ChestX-ray14, or asks about evaluating this task. Reports AI-Judge AUC.
- ▌ Chestxray14 Plco Eval · qhjqhj00Evaluates a model's ability to detect and localize multiple pathologies in high-resolution chest X-ray images. It specifically probes the model's robustness to severe class imbalance and its capacity to leverage explicit spatial location information for pathology classification. Use when the user wants to benchmark on ChestX-Ray14, PLCO, or asks about evaluating this task. Reports AUC.
- ▌ Childsafe Safety Eval · qhjqhj00Evaluates LLM safety alignment across four child developmental stages (ages 6–17) using simulated agents grounded in developmental psychology. It probes how models handle sensitive contexts, boundary-testing, and age-specific cognitive limitations in multi-turn interactions. Use when the user wants to benchmark on ChildSafe Dataset, or asks about evaluating this task. Reports semantic_safety_score.
- ▌ Cic Ids2017 Nids Eval · qhjqhj00Evaluates the classification accuracy and adversarial robustness of a Graph Neural Network-based Network Intrusion Detection System (NIDS) on distinguishing benign traffic from various attack types in network flow data. Use when the user wants to benchmark on CIC-IDS2017, or asks about evaluating this task. Reports weighted F1-score.
- ▌ Cifar Robustness Eval · qhjqhj00Evaluates the robustness of adversarially trained neural networks against Projected Gradient Descent (PGD) attacks on CIFAR-10 and CIFAR-100. It measures both clean (natural) classification accuracy and robust accuracy under varying attack strengths (PGD-20 and PGD-100). Use when the user wants to benchmark on CIFAR-10, CIFAR-100, or asks about evaluating this task. Reports PGD-20 accuracy.
- ▌ Clam Adversarial Eval · qhjqhj00Evaluates the robustness of classical neural networks and quantum neural networks against adversarial attacks by measuring performance degradation on a malware classification task after injecting random noise into input features. Use when the user wants to benchmark on ClaMP_Integrated, or asks about evaluating this task. Reports accuracy.
- ▌ Clear Regression Eval · qhjqhj00Evaluates a model's ability to continuously learn from non-stationary data streams without catastrophic forgetting, balancing stability and plasticity in regression tasks. Use when the user wants to benchmark on Artificial periodic dataset, Wind power generation dataset, or asks about evaluating this task. Reports prediction error.
- ▌ Cloud Vm Ranking Eval · qhjqhj00This benchmark evaluates how effectively different Virtual Machines (VMs) can execute specific application workloads by ranking them according to weighted hardware attributes. It probes the capability to map domain-specific application requirements to underlying infrastructure performance characteristics. Use when the user wants to benchmark on Cloud VM Benchmarking Suite, or asks about evaluating this task. Reports S_i.
- ▌ Coco Racial Bias Eval · qhjqhj00This protocol evaluates racial bias in image captioning models by measuring performance disparities between images containing lighter-skinned versus darker-skinned individuals. It probes whether models systematically generate lower-quality captions or exhibit different linguistic patterns for darker-skinned subjects compared to lighter-skinned ones, even when visual content is controlled. Use when the user wants to benchmark on COCO 2014 validation, or asks about evaluating this task. Reports CIDEr.
- ▌ Codeflaws Repair Eval · qhjqhj00This evaluation probes an LLM's ability to automatically detect and fix bugs in C programs by generating correct patches. It measures how effectively the model leverages test feedback, fault localization scores, and iterative reasoning to pass all provided test cases for each buggy submission. Use when the user wants to benchmark on Codeflaws, or asks about evaluating this task. Reports Repair Accuracy.
- ▌ Commitment Audit Eval · qhjqhj00Evaluates the extent to which authors fulfill promises made during peer review rebuttals in their final camera-ready papers, and classifies unfulfilled commitments by severity and difficulty. Use when the user wants to benchmark on ICLR 2025, EMNLP 2024, or asks about evaluating this task. Reports fulfillment rate.
- ▌ Common Voice Asr Eval · qhjqhj00Evaluates multilingual automatic speech recognition (ASR) capabilities, specifically testing speaker generalization and low-resource language adaptation via transfer learning from an English model. It measures how well a model can transcribe audio from diverse, crowdsourced speakers across multiple languages with varying data sizes. Use when the user wants to benchmark on Common Voice, or asks about evaluating this task. Reports character error rate.
- ▌ Connect The Dots Eval · qhjqhj00Evaluates a model's ability to locate and connect dots in sequential order across various visual patterns. It probes precise spatial reasoning and the capacity to generate non-destructive SVG overlays that explain the reasoning process. Use when the user wants to benchmark on Connect-the-Dots, or asks about evaluating this task. Reports Accuracy.