qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Scannet 3d Detection Eval · qhjqhj00Evaluates the ability of 3D object detectors to localize and classify indoor objects using variable-frame sparse RGB-D inputs. It probes generalization across different input modalities (reconstructed point clouds, multi-view RGB-D, monocular RGB-D) and varying numbers of input views. Use when the user wants to benchmark on ScanNet, or asks about evaluating this task. Reports mAP@0.25.
- ▌ Scientific Embedding Eval · qhjqhj00Evaluates static (Word2Vec, FastText) and transformer-based (SciBERT, RoBERTa) embedding models on scientific text using intrinsic (word/sentence similarity) and extrinsic (NER, document classification) tasks. Probes the impact of domain-specific pretraining and sub-word tokenization on representation quality and downstream performance. Use when the user wants to benchmark on UNMSRS, SemEval, Clinical STS 2018, Clinical STS 2019, Conll 2003, CHEMDNER, SciERC, Reuters 12, BioChem 8, or asks about evaluating this task. Reports Pearson, F-Beta.
- ▌ Sdg Gan Oversampling Eval · qhjqhj00Evaluates whether a GAN-based oversampling technique (SDG-GAN) improves binary classification performance on imbalanced tabular data compared to traditional and GAN-based baselines. Use when the user wants to benchmark on Credit Card Fraud Dataset, Pima Diabetes Dataset, Breast Cancer Wisconsin (Diagnostic) Dataset, Gambling Fraud Dataset, or asks about evaluating this task. Reports algorithmic performance.
- ▌ Sdmuse Music Editing Eval · qhjqhj00Evaluates the quality and controllability of a stochastic differential music generation and editing model. It probes the model's ability to generate pop piano music from scratch or conditioned on control signals, and perform fine-grained editing tasks like stroke-based generation, inpainting, and style transfer. Use when the user wants to benchmark on ailabs1k7, or asks about evaluating this task. Reports pitch distribution similarity (PD).
- ▌ Skyscenes Aerial Seg Eval · qhjqhj00Evaluates semantic segmentation models trained on synthetic aerial imagery for their ability to generalize to real-world UAV datasets and adapt to varying environmental conditions like weather, time of day, and camera viewpoint. Use when the user wants to benchmark on SKYSCENES, UAVid, AEROSCAPES, ICG DRONE, SYNDrone, or asks about evaluating this task. Reports mIoU.
- ▌ Soap Note Generation Eval · qhjqhj00Evaluates the ability of audio and text models to generate clinically accurate, well-structured SOAP notes from long-form doctor-patient conversations. Probes long-context audio reasoning, fact-grounding, and clinical documentation quality. Use when the user wants to benchmark on Doctor-Patient SOAP Conversations, or asks about evaluating this task. Reports Faithfulness.
- ▌ Spatial Audio Motion Eval · qhjqhj00Probes a model's ability to detect overlapping audio events in dynamic spatial recordings, estimate their direction of arrival and distance, and perform reasoning about moving sound sources. Use when the user wants to benchmark on STARS23, FOA-MEIR Derived, or asks about evaluating this task. Reports F-score.
- ▌ Spatial Intelligence Eval · qhjqhj00Evaluates a model's ability to perceive, reason about, and reconstruct 3D spatial layouts, multi-view relationships, and perspective-taking from visual inputs. It probes robustness against language shortcuts and tests generalization to longer video sequences and embodied manipulation tasks. Use when the user wants to benchmark on VSI-Bench, MMSI-Bench, MindCube, ViewSpatial-Bench, SITE, MMBench-En, EmbodiedBench (spatial subset), or asks about evaluating this task. Reports accuracy.
- ▌ Speech Drame Realism Eval · qhjqhj00Evaluates a speech foundation model's ability to generate role-play responses grounded in bottom-up, human-grounded realism. It measures fine-grained aspects such as prosodic dynamics, emotional fidelity, character consistency, and contextual fit based on realistic dialogue and media sources. Use when the user wants to benchmark on DRAME-RoleBench (Realism), or asks about evaluating this task. Reports realism_score.
- ▌ Speechparaling Bench Eval · qhjqhj00Evaluates large audio-language models (LALMs) on their ability to generate speech with fine-grained paralinguistic features, including dynamic intra-utterance variation and context-aware adaptation. It probes how well models interpret and modulate tone, pitch, emotion, and non-linguistic vocalizations in response to textual instructions and contextual cues. Use when the user wants to benchmark on SpeechParaling-Bench, or asks about evaluating this task. Reports Judge Score (0-100).
- ▌ Squashing Activation Eval · qhjqhj00Evaluates a novel 'Squashing' activation function against standard alternatives (ReLU, Sigmoid, Tanh) on synthetic 2D classification tasks and the Fashion-MNIST image classification benchmark. It measures how well continuously differentiable logical approximations perform compared to conventional non-linearities in terms of convergence speed and final classification accuracy. Use when the user wants to benchmark on Fashion-MNIST, Synthetic 2D Classification, or asks about evaluating this task. Reports test accuracy.
- ▌ Streamflow Hydrology Eval · qhjqhj00Evaluates hydrological forecasting models on predicting streamflow across 319 US basins. It tests the model's ability to capture long-term temporal dependencies and intermediate hydrological states (soil water, snowpack) under varying data segmentation, training sizes, and noise conditions. Use when the user wants to benchmark on 319 US Basins, or asks about evaluating this task. Reports RMSE.
- ▌ Structured Prompting Eval · qhjqhj00Evaluates how different prompting strategies (baseline, zero-shot, chain-of-thought, and automated optimizers) affect the accuracy, ranking stability, and variance of language model performance across multiple knowledge and reasoning benchmarks. Use when the user wants to benchmark on MMLU-Pro, GSM8K, MedCalc-Bench, GPQA, HeadQA, MedBullets, Medec, or asks about evaluating this task. Reports accuracy.
- ▌ Surge Sequential Rec Eval · qhjqhj00Evaluates a model's ability to predict the next item in a user's interaction sequence by leveraging long-term historical behavior and filtering out noise. It probes the model's capacity to handle varying sequence lengths and efficiently model dynamic user preferences over time. Use when the user wants to benchmark on Taobao, Kuaishou, or asks about evaluating this task. Reports GAUC.
- ▌ Symsearch Omnigibson Eval · qhjqhj00Evaluates an agent's ability to perform open-vocabulary interactive object search in indoor environments using relational semantic reasoning over 3D scene graphs. It probes exploration efficiency, reasoning accuracy, and computational cost compared to embedding-based and LLM-based planners. Use when the user wants to benchmark on SymSearch, OmniGibson, or asks about evaluating this task. Reports Success Rate (SR), Success weighted by Path Length (SPL).
- ▌ Tabular Data Centric Eval · qhjqhj00This evaluation probes the robustness and relative performance of tabular machine learning models when subjected to expert-level, dataset-specific preprocessing pipelines rather than standardized baselines. It specifically measures how feature engineering, hyperparameter optimization, and test-time adaptation shift model rankings and close performance gaps across real-world competition datasets. Use when the user wants to benchmark on Kaggle competition datasets (MBGM, BPCCM, HQC, SCTP, PSSDP, AEAC, OGPCC, SCS, IFD, SVPC, electricity), or asks about evaluating this task. Reports leaderboard rank.
- ▌ Tedigan Text To Face Eval · qhjqhj00Evaluates a model's capability to synthesize high-resolution, diverse, and photorealistic face images conditioned on natural language prompts, and to perform text-guided editing of existing faces while preserving identity and irrelevant attributes. Use when the user wants to benchmark on Multi-Modal CelebA-HQ, or asks about evaluating this task. Reports FID.
- ▌ Temporal Degradation Eval · qhjqhj00Evaluates how temporal misalignment between pretraining/fine-tuning data and evaluation data impacts model performance across classification and summarization benchmarks. Use when the user wants to benchmark on PubCLS, NewSum, TwiERC, AIC, PoliAff, or asks about evaluating this task. Reports Accuracy.
- ▌ Text Image Retrieval Eval · qhjqhj00Evaluates the model's ability to align facial images with their textual descriptions by retrieving the correct image given a text query, and vice versa. It measures how well the model learns cross-modal semantic correspondence for face-centric data. Use when the user wants to benchmark on CelebA-Caption, MM-CelebA, or asks about evaluating this task. Reports R@5, R@10.
- ▌ Text Video Alignment Eval · qhjqhj00This evaluation protocol assesses how well text-to-video generation models align generated content with textual prompts across fine-grained attributes like object counts, colors, actions, and spatial relationships. It measures both semantic alignment and visual/motion quality to determine if refinement techniques successfully correct misalignments without degrading fidelity. Use when the user wants to benchmark on EvalCrafter, T2V-CompBench, or asks about evaluating this task. Reports Text-Video Alignment.
- ▌ Time Series Transfer Eval · qhjqhj00Evaluates the effectiveness of transfer learning on time series data by comparing pre-trained models against models trained from scratch across intra-domain and cross-domain settings. It probes whether shared temporal structures enable knowledge transfer and how dataset size and domain similarity affect predictive performance and training convergence. Use when the user wants to benchmark on LEN-DB, SPEECH, EMG, S&P 500, LOMAX, STEAD, or asks about evaluating this task. Reports MAE, weighted F1 score.
- ▌ Time Moe Forecasting Eval · qhjqhj00Evaluates long-term time series forecasting capabilities of foundation models in both zero-shot (unseen datasets) and in-distribution (fine-tuned) settings across multiple prediction horizons. Use when the user wants to benchmark on ETTh1, ETTh2, ETTm1, ETTm2, Weather, Global Temp, or asks about evaluating this task. Reports MSE.
- ▌ Topic Level Polarity Eval · qhjqhj00Predicts the sentiment polarity associated with a specific topic within a tweet, requiring topic-aware sentiment classification and contextual disambiguation. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports macro-averaged F1.
- ▌ Trace Classification Eval · qhjqhj00Evaluates a model's ability to classify OpenTelemetry workflow traces as benign, suspicious, or malicious, and assesses its knowledge of cybersecurity frameworks via multiple-choice questions. Use when the user wants to benchmark on OpenTelemetry Workflow Traces, or asks about evaluating this task. Reports Overall Accuracy.
- ▌ Trace Gmv Prediction Eval · qhjqhj00This benchmark evaluates a model's ability to predict post-click Gross Merchandise Volume (GMV) under delayed feedback conditions. It specifically probes how well models adapt to rapidly evolving label distributions through online streaming training and whether they can effectively handle the distinct statistical properties of single-purchase versus repurchase transactions. Use when the user wants to benchmark on TRACE, or asks about evaluating this task. Reports AUC.
- ▌ Trec Dl Hole Filling Eval · qhjqhj00Evaluates the ability of LLMs to accurately predict missing relevance judgments (holes) in information retrieval test collections. It measures how well synthetic incomplete judgments can be restored to match ground truth relevance labels across varying hole percentages. Use when the user wants to benchmark on TREC DL 2019/2020/2021, or asks about evaluating this task. Reports Kendall τ.
- ▌ Uniem3m Segmentation Eval · qhjqhj00Evaluates an instance segmentation model's ability to accurately delineate and separate individual microstructural objects in high-resolution electron micrographs, particularly under varying instance densities. Use when the user wants to benchmark on UniEM-3M, or asks about evaluating this task. Reports mAP@0.5.
- ▌ Unified Med Vlm Benchmark · qhjqhj00Evaluates medical vision-language models across a broad capability surface including visual diagnosis, medical imaging, clinical reasoning, text-based QA, report generation, and instruction following. It emphasizes protocol reproducibility, deployment relevance (safety, consistency, faithfulness), and robustness to real-world clinical imagery and OCR conditions. Use when the user wants to benchmark on Public Med-VLM Benchmarks (30+ subsets), Inhouse VQA, Inhouse OCR, Inhouse Caption, or asks about evaluating this task. Reports MCQ/Short QA Accuracy.
- ▌ Urban Tree Detection Eval · qhjqhj00Evaluates the capability of segmentation models to detect and delineate individual tree crowns and canopy coverage from high-resolution aerial imagery across diverse urban and tropical environments. Use when the user wants to benchmark on Zurich Municipal Tree Inventory & Swisstopo Imagery, WeRobotics Open AI Challenge (Tonga), or asks about evaluating this task. Reports Recall.
- ▌ Vevo Voice Imitation Eval · qhjqhj00This benchmark evaluates a model's ability to perform zero-shot voice imitation by disentangling linguistic content, speaker timbre, and vocal style (accent/emotion). It probes the model's capacity to generate high-intelligibility speech that accurately transfers the target speaker's identity and stylistic attributes from a reference clip without task-specific fine-tuning. Use when the user wants to benchmark on Vevo Evaluation Set (AB, CV, ACCENT, EMOTION), or asks about evaluating this task. Reports WER.
- ▌ Video Class Agnostic Eval · qhjqhj00Evaluates a model's ability to segment moving and unknown objects in autonomous driving videos without relying on a closed set of known classes. It probes open-set and motion-based instance segmentation capabilities under varying data distributions and synthetic scenarios. Use when the user wants to benchmark on Cityscapes-VPS, KITTI-MOTS, Carla, or asks about evaluating this task. Reports CAQ, CA-IoU.
- ▌ Vigil Cognitive Bias Eval · qhjqhj00Evaluates a real-time browser extension's ability to detect and mitigate cognitive bias triggers in online text. It probes span-level identification of persuasive rhetoric and bias patterns, alongside system latency and mitigation quality. Use when the user wants to benchmark on SemEval-2020 Task 11, Moralization Corpus, or asks about evaluating this task. Reports micro-F1.
- ▌ Virc Multimodal Math Eval · qhjqhj00Evaluates multimodal mathematical reasoning and high-resolution visual perception capabilities of vision-language models. It probes the model's ability to decompose complex problems into structured reasoning chunks, interleave visual tool calls, and produce accurate final answers across geometric, mathematical, and fine-grained visual benchmarks. Use when the user wants to benchmark on GeoQA, MathVista-Math, MMStar-Math, VisualProbe, V*, HR-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Visualinformationfidelity · qhjqhj00Compute the VisualInformationFidelity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute VisualInformationFidelity, or asks how to score with VisualInformationFidelity.
- ▌ Vla Cross Embodiment Eval · qhjqhj00Evaluates a vision-language-action model's ability to generalize across diverse robotic embodiments, simulation environments, and real-world platforms. It probes cross-embodiment adaptation, parameter-efficient fine-tuning capabilities, and dexterous manipulation performance. Use when the user wants to benchmark on Libero, Simpler, Calvin, VLABench, RoboTwin-2.0, NAVSIM, BridgeData-v2, Soft-Fold, or asks about evaluating this task. Reports success_rate.
- ▌ Vlm Deflection Bench Eval · qhjqhj00This benchmark evaluates the ability of large vision-language models to correctly answer knowledge-based visual questions while properly deferring when evidence is missing or hallucinating when faced with noisy or conflicting retrieval contexts. It disentangles parametric memorization from retrieval robustness across four controlled scenarios. Use when the user wants to benchmark on VLM-DeflectionBench, or asks about evaluating this task. Reports Deflection Rate.
- ▌ Voice Cloning Accent Eval · qhjqhj00Evaluates how commercial voice cloning systems preserve speaker identity and speech intelligibility for standard versus accented Mandarin speakers. It probes the alignment between acoustic embedding distances and human perceptual judgments of similarity and intelligibility across accent conditions. Use when the user wants to benchmark on Standard and Accented Mandarin Speech Dataset, or asks about evaluating this task. Reports Intelligibility gain score.
- ▌ Voicebank Demand Sep Eval · qhjqhj00Evaluates zero-shot language-queried speech enhancement by isolating clean speech from noisy backgrounds. The benchmark tests the model's ability to enhance speech using a fixed text query. Use when the user wants to benchmark on Voicebank-DEMAND, or asks about evaluating this task. Reports PESQ.
- ▌ Votreuth Rehab Depth Eval · qhjqhj00Evaluates marker-less 2D pose estimation models on rehabilitation-specific depth images. It probes the model's ability to generalize from generic adult standing poses to complex clinical postures, including children, and tests robustness against varying subject scales and positions. Use when the user wants to benchmark on ITOP, VtR, VtR-O, or asks about evaluating this task. Reports PCK.
- ▌ Wangchanthaiinstruct Eval · qhjqhj00Evaluates instruction-following capabilities of LLMs in Thai across culture-aware, domain-specific (Medical, Law, Finance, Retail), and multitask settings. Probes factual accuracy, reasoning quality, and fluency in both zero-shot and fine-tuned regimes. Use when the user wants to benchmark on WangchanThaiInstruct, Thai LLM Leaderboard, Thai MT-Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Wenetspeech Wu Bench Eval · qhjqhj00Evaluates speech processing capabilities for the Chinese Wu dialect, including automatic speech recognition (ASR), automatic speech translation (AST), speaker attribute prediction (gender, age), emotion recognition, text-to-speech (TTS), and instruction-following TTS. Use when the user wants to benchmark on WenetSpeech-Wu-Bench, or asks about evaluating this task. Reports CER (%).
- ▌ Wildguard Moderation Eval · qhjqhj00This evaluation protocol assesses the safety moderation capabilities of LLMs and dedicated moderation models. It probes their ability to detect harmful content in user prompts, classify harmful or safe model responses, and identify whether a model appropriately refuses unsafe requests across multiple risk categories. Use when the user wants to benchmark on ToxicChat, OpenAI Mod, AegisSafetyTest, SimpleSafetyTests, Harmbench Prompt, Harmbench Resp, BeaverTails, SafeRLHF, XSTest-Resp, WildGuardTest, or asks about evaluating this task. Reports F1 score.
- ▌ Wingpt 3 0 Benchmark Eval · qhjqhj00Evaluates large language models on comprehensive medical reasoning, clinical calculation, and general cognitive capabilities. It probes domain-specific knowledge application, diagnostic reasoning, and complex problem-solving in real-world clinical and academic settings. Use when the user wants to benchmark on MedCalc, MedReMCQ, CMMLU, MATH-500, MedQA-USMLE, MedMCQA, PubMedQA, or asks about evaluating this task. Reports accuracy.
- ▌ Xai Clip Medical Seg Eval · qhjqhj00Evaluates the computational efficiency and explanation fidelity of an ROI-guided perturbation framework for medical image segmentation. It measures how effectively the method reduces computation while preserving segmentation accuracy and explanation quality compared to full occlusion baselines. Use when the user wants to benchmark on FLARE22, SAROS, CHAOS, or asks about evaluating this task. Reports Dice coefficient.
- ▌ Xas Os Cn Prediction Eval · qhjqhj00Evaluates a machine learning model's ability to predict local chemical descriptors (oxidation state and coordination number) from experimental X-ray absorption spectra, specifically testing how well spectral domain mapping bridges the gap between simulated training data and real experimental measurements. Use when the user wants to benchmark on Combinatorial Zinc Titanate Thin Film XANES, or asks about evaluating this task. Reports OS/CN prediction accuracy.
- ▌ Xwikis Summarisation Eval · qhjqhj00This benchmark evaluates the ability of multilingual and cross-lingual models to generate accurate English summaries from source documents in German, French, and Czech. It probes supervised, zero-shot, and few-shot cross-lingual transfer capabilities, as well as model robustness on out-of-domain news text. Use when the user wants to benchmark on XWikis, D_en→en, Voxeurop, or asks about evaluating this task. Reports ROUGE-L recall.
- ▌ Youtube Implicit Rec Eval · qhjqhj00Evaluates recommender systems on implicit feedback datasets, testing their ability to rank relevant items for users. It probes model versatility across cold-start, offline, and instant recommendation scenarios using side information and sequential context features. Use when the user wants to benchmark on YouTube Implicit Feedback Subset, or asks about evaluating this task. Reports NDCG@100.
- ▌ Yqsong Execution Accuracy · qhjqhj00Compute yqsong/execution_accuracy via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yqsong/execution_accuracy.
- ▌ Tsdown · qhjqhj00 bundleBundle TypeScript and JavaScript libraries with blazing-fast speed powered by Rolldown. Use when building libraries, generating type declarations, bundling for multiple formats, or migrating from tsup.
- ▌ Vitest · qhjqhj00 bundleVitest fast unit testing framework powered by Vite with Jest-compatible API. Use when writing tests, mocking, configuring coverage, or working with test filtering and fixtures.
- ▌ Accented Clinical Asr Eval · qhjqhj00Evaluates ASR models on African-accented clinical speech to measure how well they transcribe medical named entities (MNEs) like drug names, diagnoses, and lab results. It specifically probes the gap between standard word-level accuracy and clinically relevant entity recognition. Use when the user wants to benchmark on AfriSpeech, or asks about evaluating this task. Reports M-WER.
- ▌ Adjusted Mutual Info Score · qhjqhj00Compute the adjusted_mutual_info_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute adjusted_mutual_info_score, or asks how to score with adjusted_mutual_info_score.
- ▌ African LLM Benchmark Eval · qhjqhj00This evaluation probes the cross-lingual reasoning and domain knowledge capabilities of large language models across low-resource African languages. It measures how well models perform on translated benchmarks compared to English, and assesses the impact of cultural appropriateness and fine-tuning data quality on model accuracy. Use when the user wants to benchmark on Winogrande, MMLU (Clinical Sections), Belebele, or asks about evaluating this task. Reports accuracy.
- ▌ AI Writing Assistance Eval · qhjqhj00Evaluates how source disclosure and perceived AI authorship influence human editing behavior and subsequent peer-review acceptance decisions for scientific abstracts. Use when the user wants to benchmark on CS-Conference-Abstracts, or asks about evaluating this task. Reports accept/reject decision.
- ▌ Annbatch Data Loading Eval · qhjqhj00Measures the data loading throughput and epoch iteration time for large-scale biological datasets. It benchmarks how efficiently a loader can fetch and prepare mini-batches from disk compared to existing frameworks. Use when the user wants to benchmark on Tahoe100M, 1000 Genomes GRCh38, Single-cell microscopy images, or asks about evaluating this task. Reports samples/sec.
- ▌ Apptek Callcenter Asr Eval · qhjqhj00This benchmark evaluates automatic speech recognition (ASR) systems on their ability to transcribe long-form, spontaneous call-center dialogues across 14 English accents. It specifically probes robustness to non-standard accents, conversational speech patterns, and sensitivity to audio segmentation strategies. Use when the user wants to benchmark on AppTek Call-Center Dialogues, or asks about evaluating this task. Reports WER.
- ▌ Arbitrary View Action Eval · qhjqhj00Evaluates human action recognition models under varying camera viewpoints and different subjects. It probes the model's ability to generalize across unseen subjects, unseen camera angles, and continuous 360-degree view changes. Use when the user wants to benchmark on Varying-view RGB-D Action Dataset, or asks about evaluating this task. Reports average recognition accuracy.
- ▌ Argoverse2 Trajectory Eval · qhjqhj00Evaluates autonomous driving models on joint trajectory prediction and controllable generation tasks. It probes the model's ability to forecast multi-agent future paths accurately and generate realistic, goal-conditioned trajectories efficiently using diffusion-based sampling. Use when the user wants to benchmark on Argoverse 2, or asks about evaluating this task. Reports avgBrierMinFDE_K.
- ▌ Arima Fraud Detection Eval · qhjqhj00Evaluates unsupervised anomaly detection models on credit card transaction time series to identify fraudulent spending deviations. It probes the ability of models to balance precision and recall in highly imbalanced, real-world financial data without relying on labeled fraud examples. Use when the user wants to benchmark on Credit card transaction time series, or asks about evaluating this task. Reports F-Measure.
- ▌ Audio Spoof Detection Eval · qhjqhj00Evaluates the robustness of audio spoof detection models against real-world audio degradation and manipulation attacks (laundering), including reverberation, additive noise, and re-compression. Use when the user wants to benchmark on ASVspoof 2019 LA, ASVspoof Laundered Database, or asks about evaluating this task. Reports EER.
- ▌ Author Centric Review Eval · qhjqhj00Evaluates an LLM's ability to generate structured, author-centric academic feedback (Summary, Strengths, Weaknesses, Questions) from long research papers. It probes the model's capacity to retrieve salient passages via graph-based retrieval and synthesize constructive pre-submission reviews without relying on full context or multi-agent systems. Use when the user wants to benchmark on ICLR 2024 (ICT), CNT_10, or asks about evaluating this task. Reports human evaluation.
- ▌ Automated Red Teaming Eval · qhjqhj00Evaluates an LLM's capability to generate effective adversarial prompts (red teaming attacks) for arbitrary safety goals. It measures both the success rate of eliciting targeted behaviors and the diversity of the generated attacks across in-domain and out-of-domain objectives. Use when the user wants to benchmark on garak adversarial goals, or asks about evaluating this task. Reports attack success rate.
- ▌ Av Speech Enhancement Eval · qhjqhj00This benchmark evaluates a model's ability to isolate a target speaker's voice from multi-talker audio environments using only lip-region video inputs. It probes audio-visual speech enhancement, testing how well the network predicts magnitude and phase masks to suppress interference and noise while preserving speech intelligibility and perceptual quality. Use when the user wants to benchmark on LRS2, VoxCeleb2, or asks about evaluating this task. Reports PESQ.
- ▌ Average Per Token Log Prob · qhjqhj00Evaluates language models on multiple-choice or candidate-selection downstream tasks by scoring candidate answers based on their likelihood under the model. It measures how well the model assigns high probability to the correct answer among a set of options. Use when the user has predictions and gold and needs to compute average_per_token_log_prob.
- ▌ Bascobasculino Mot Metrics · qhjqhj00Compute bascobasculino/mot-metrics via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bascobasculino/mot-metrics.
- ▌ Bayesian Optical Flow Eval · qhjqhj00Evaluates a Bayesian statistical inversion method for estimating optical flow fields and quantifying their uncertainty from image pairs, compared against deterministic baselines. Use when the user wants to benchmark on Synthetic benchmark flow fields, Middlebury dataset, or asks about evaluating this task. Reports reconstruction accuracy.
- ▌ Bayesian Optimization Eval · qhjqhj00Evaluates the ability of Bayesian optimization methods to efficiently search discrete spaces (molecules, arithmetic expressions) by maximizing or minimizing a black-box objective function over a limited budget of oracle calls. It probes how well a model aligns its latent representation with the objective landscape to guide search. Use when the user wants to benchmark on Guacamol, TDC DRD3, Arithmetic Expression, or asks about evaluating this task. Reports objective value.
- ▌ Bayling2 Multilingual Eval · qhjqhj00Evaluates multilingual translation quality and cross-lingual reasoning across high-resource and low-resource languages. It probes the model's ability to align languages and transfer capabilities from high-resource to low-resource settings without extensive low-resource instruction data. Use when the user wants to benchmark on Flores-101, WMT22, Belebele, XNLI, GSM8K, or asks about evaluating this task. Reports BLEU (sacrebleu), COMET.
- ▌ Behavioral Prediction Eval · qhjqhj00Predicts individual strategic decisions by conditioning on structured psychometric trait profiles (e.g., Big Five personality traits). It probes a model's ability to map high-dimensional psychological embeddings to discrete behavioral outcomes in unseen situational contexts. Use when the user wants to benchmark on Strategic Scenario Dataset, or asks about evaluating this task. Reports balanced accuracy, macro-F1.
- ▌ Bert Noise Robustness Eval · qhjqhj00Evaluates BERT's robustness to synthetic character-level noise across sentiment classification and textual similarity tasks. It probes how spelling mistakes and typos disrupt subword tokenization and degrade contextual embeddings under varying noise intensities. Use when the user wants to benchmark on IMDB, SST-2, STS-B, or asks about evaluating this task. Reports F1 score.
- ▌ Bimr Interpretability Eval · qhjqhj00This benchmark evaluates the reasoning interpretability of knowledge graph completion models by measuring how well their generated multi-hop paths or rules can be understood and validated. It probes whether models produce semantically reasonable explanations rather than just statistically valid paths, highlighting the gap between link prediction accuracy and actual explainability. Use when the user wants to benchmark on WD15K, FB15K-237, or asks about evaluating this task. Reports GI (Global Interpretability).
- ▌ Binaryprecisionrecallcurve · qhjqhj00Compute the BinaryPrecisionRecallCurve metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryPrecisionRecallCurve, or asks how to score with BinaryPrecisionRecallCurve.
- ▌ Bing Chat Open Domain Eval · qhjqhj00Evaluates an LLM's ability to perform zero-shot dialogue segmentation and joint state tracking on real-world, open-domain human-LLM conversations. It probes the model's capacity to identify topic shifts, assign intent and domain labels, and maintain context over long multi-turn interactions without hallucination. Use when the user wants to benchmark on Bing Chat (Internal Human-LLM Dialogue Dataset), or asks about evaluating this task. Reports JGA (I/D).
- ▌ Black Box Attribution Eval · qhjqhj00Evaluates the faithfulness of black-box attribution methods by measuring how well identified input regions align with the model's decision-making process. It tests the ability of explanation algorithms to pinpoint critical features that drive correct predictions or cause errors. Use when the user wants to benchmark on ImageNet, CUB-200-2011, CelebA, VGG-Face2, LC25000 (Lung), VGG-Sound, or asks about evaluating this task. Reports Deletion AUC, Insertion AUC.
- ▌ Blas Performance Benchmark · qhjqhj00Evaluates the runtime performance and throughput of a C++ expression template library (SALT) against optimized BLAS implementations (Intel MKL) and other template libraries (Eigen) for standard vector operations. Use when the user has predictions and gold and needs to compute performance (GFLOPS).
- ▌ Borehole Segmentation Eval · qhjqhj00This evaluation probes a model's ability to perform weakly supervised multimodal segmentation of acoustic borehole images by refining threshold-guided pseudo-labels using depth-aligned well logs. It measures how well the predicted segmentation aligns with a provisional target map, testing spatial coherence and multimodal feature fusion rather than absolute geological accuracy. Use when the user wants to benchmark on Antilope25 & Botorosa47 borehole intervals, or asks about evaluating this task. Reports permutation-invariant agreement.
- ▌ Brain Tumor Detection Eval · qhjqhj00Evaluates a deep learning model's ability to classify brain MRI images into four tumor categories (glioma, meningioma, no tumor, pituitary). It probes multi-class image classification performance, generalization to unseen medical scans, and the model's capacity to balance precision and recall across classes. Use when the user wants to benchmark on Public MRI dataset (unspecified), or asks about evaluating this task. Reports accuracy.
- ▌ Brats 2023 Meningioma Eval · qhjqhj00Evaluates 3D medical image segmentation models on intracranial meningioma MRI scans. It probes volumetric accuracy and boundary sharpness across three tumor subregions (enhancing tumor, tumor core, whole tumor) under varying contrast and lesion size conditions. Use when the user wants to benchmark on BraTS 2023 Intracranial Meningioma Challenge, or asks about evaluating this task. Reports DSC.
- ▌ Bucc Bitext Retrieval Eval · qhjqhj00Evaluates the capability of sentence embeddings to retrieve exact translation pairs (bitexts) from large monolingual corpora across different languages. It tests the model's ability to distinguish parallel sentences from semantically similar but non-parallel ones. Use when the user wants to benchmark on BUCC bitext mining task, or asks about evaluating this task. Reports F1 score.
- ▌ C Mteb Sentence Embed Eval · qhjqhj00Evaluates the zero-shot quality of sentence and word embeddings from decoder-only LLMs on downstream tasks without additional training. Probes capabilities in text classification, clustering, retrieval, semantic similarity, and polysemous word disambiguation. Use when the user wants to benchmark on C-MTEB, SLPWC (C-SEM), WSD, or asks about evaluating this task. Reports Average C-MTEB normalized score.
- ▌ Cbr Encrypted Traffic Eval · qhjqhj00Evaluates an ANN-based adaptive classifier's ability to classify encrypted network traffic into known categories such as malware families, operating systems, browsers, and applications. It specifically probes the model's capacity to dynamically adapt to new or out-of-distribution classes without retraining, while measuring any performance degradation on existing classes compared to traditional baselines. Use when the user wants to benchmark on BOA, MTA, or asks about evaluating this task. Reports classification performance.
- ▌ Chain Of Instructions Eval · qhjqhj00Evaluates LLMs on multi-step compositional instruction following, generalization to hard single-step tasks, and multilingual summarization. It probes the model's ability to chain subtask outputs as inputs for subsequent steps and maintain coherence across multiple instructions. Use when the user wants to benchmark on CoI2-test, CoI3-test, BIG-Bench Hard (BBH), Multilingual Summarization, or asks about evaluating this task. Reports Rouge-L.
- ▌ Chime4 Asr Robustness Eval · qhjqhj00Evaluates the robustness of automatic speech recognition systems in noisy environments by measuring word error rates on enhanced speech from speaker extraction models across matched and mismatched noisy conditions. Use when the user wants to benchmark on CHiME-4, VoiceBank-DEMAND, WHAM!, or asks about evaluating this task. Reports WER.
- ▌ Cifar100 Miniimagenet Eval · qhjqhj00This evaluation protocol probes a model's ability to perform standard supervised image classification and adaptive few-shot episodic learning. It measures how well the model generalizes to unseen classes under limited supervision by averaging accuracy over multiple sampled episodes. Use when the user wants to benchmark on CIFAR-100, Mini-ImageNet, or asks about evaluating this task. Reports accuracy.
- ▌ Classactionprediction Eval · qhjqhj00Predicts the outcome (win or lose) of U.S. class action lawsuits based on plaintiff complaint texts. It probes a model's ability to extract legally relevant allegations from long-form, unverified legal documents and make binary judgment predictions. Use when the user wants to benchmark on ClassActionPrediction, or asks about evaluating this task. Reports accuracy.
- ▌ Climate Finance Bench Eval · qhjqhj00Evaluates Retrieval-Augmented Generation (RAG) systems on climate-finance question answering using expert-validated Q&A pairs from corporate sustainability reports. It measures answer correctness across different retrieval strategies and LLMs, while also quantifying the environmental footprint (GHG emissions) of each configuration. Use when the user wants to benchmark on Climate Finance Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Clinical Note Scoring Eval · qhjqhj00Evaluates the capability of transformer-based models to automatically assign numerical scores to clinical patient notes. It probes how masked language modeling pretraining and pseudo-labeling strategies improve scoring performance across different model architectures. Use when the user wants to benchmark on Unspecified clinical patient notes dataset, or asks about evaluating this task. Reports CV Score.
- ▌ Clipimagequalityassessment · qhjqhj00Compute the CLIPImageQualityAssessment metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CLIPImageQualityAssessment, or asks how to score with CLIPImageQualityAssessment.
- ▌ Cloth Unfolding Grasp Eval · qhjqhj00Evaluates a robot's ability to select effective grasp poses on hanging, folded cloth to maximize unfolding coverage. It probes material-aware perception, geometric reasoning, and policy-based grasp generation under complex folding configurations. Use when the user wants to benchmark on ICRA 2024 Cloth Competition dataset, or asks about evaluating this task. Reports relative coverage.
- ▌ Cnn Edge Optimization Eval · qhjqhj00Evaluates the trade-offs between predictive accuracy, model compression, and dynamic inference efficiency of CNN optimization techniques (pruning, quantization, early-exit) for edge deployment. It probes how different architectures handle static compression versus input-adaptive latency reduction under hardware-constrained conditions. Use when the user wants to benchmark on Unspecified classification dataset, or asks about evaluating this task. Reports accuracy (%).
- ▌ Cointegrated Blaser 2 0 Qe · qhjqhj00Compute cointegrated/blaser_2_0_qe via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of cointegrated/blaser_2_0_qe.
- ▌ Coliee Task4 Legal QA Eval · qhjqhj00Evaluates large language models' ability to perform legal textual entailment and question answering in monolingual and cross-lingual settings. It probes how well models handle linguistic and structural disparities between English and Japanese legal contexts and questions. Use when the user wants to benchmark on COLIEE Task 4, or asks about evaluating this task. Reports accuracy.
- ▌ Commonsense Retrieval Eval · qhjqhj00This benchmark probes the commonsense reasoning capabilities of vision-language models by evaluating their ability to match images to text riddles (or vice versa) where the subject entity is replaced with a demonstrative pronoun. It specifically tests relational knowledge retrieval and generalization to unseen knowledge triples. Use when the user wants to benchmark on DANCE Diagnostic Set, or asks about evaluating this task. Reports Acc@50.
- ▌ Compression Benchmark Eval · qhjqhj00Evaluates the trade-offs between model accuracy and resource efficiency when applying various DNN compression techniques on mobile hardware. It measures how different compression methods affect inference speed, energy consumption, and storage footprint across standard vision and audio datasets. Use when the user wants to benchmark on CIFAR-10, MNIST, CIFAR-100, ImageNet, UbiSound, Har, or asks about evaluating this task. Reports accuracy.
- ▌ Comtail Translation Metric · qhjqhj00Evaluates the ability of neural models to predict human-assigned translation quality scores for Indian language pairs, measuring alignment with crowd-sourced DA+SQM ratings. Use when the user has predictions and gold and needs to compute Pearson correlation, Spearman correlation.
- ▌ Construction Site 10k Eval · qhjqhj00Evaluates vision-language models on construction site safety inspection tasks, including image captioning, safety rule violation detection, reasoning, and visual grounding of specific objects. Use when the user wants to benchmark on ConstructionSite 10k, or asks about evaluating this task. Reports IoU.
- ▌ Continual Learning Metrics · qhjqhj00Evaluates a model's ability to retain knowledge from previously learned tasks while continuously training on new ones, and measures how past knowledge facilitates learning new tasks and improves performance on old ones. Use when the user has predictions and gold and needs to compute Average Performance (AP).
- ▌ Conv Search Rewriting Eval · qhjqhj00This evaluation probes a model's ability to rewrite conversational search queries to maximize retrieval effectiveness. It measures how well reformulated questions help both sparse and dense retrievers locate relevant passages across different dialogue contexts, including initial turns and topic shifts. Use when the user wants to benchmark on QReCC, TopiOCQA, or asks about evaluating this task. Reports MRR.
- ▌ Covid19 Cxr Detection Eval · qhjqhj00Evaluates a deep CNN's ability to classify chest X-ray images into COVID-19 positive and negative/healthy categories using region-edge and channel-boosted features. Use when the user wants to benchmark on Three datasets (names not specified in section), or asks about evaluating this task. Reports 95% Confidence Interval (CI).
- ▌ Crowd Pose Estimation Eval · qhjqhj00Evaluates the ability of pose estimation models to accurately predict 2D keypoints for humans and animals in crowded, occluded, and multi-instance scenarios. It probes robustness to detection ambiguity, overlapping instances, and the transferability of conditional pose inputs from bottom-up detectors to top-down refiners. Use when the user wants to benchmark on CrowdPose, OCHuman, COCO, Multi-Animal (SchoolingFish, Marmosets, Tri-Mouse), or asks about evaluating this task. Reports AP.