qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Apres Paper Revision Eval · qhjqhj00This protocol evaluates an LLM's ability to predict a paper's future scientific impact based on its text and peer reviews, and its ability to iteratively revise the manuscript to maximize that predicted impact. It probes the model's capacity for rubric discovery, agentic text editing, and alignment with human expert preferences. Use when the user wants to benchmark on ICLR & NeurIPS Peer Review Dataset, or asks about evaluating this task. Reports MAE, Improvement Score ($\Delta S$).
- ▌ Ares Android Testing Eval · qhjqhj00Evaluates the ability of automated black-box testing tools and reinforcement learning agents to explore Android applications effectively. It probes how well algorithms navigate complex UI states, maximize code/activity coverage, and trigger unique application crashes within a fixed time budget. Use when the user wants to benchmark on F-Droid top starred apps, AndroTest, Synthetic FATE models, or asks about evaluating this task. Reports AUC.
- ▌ Asr Noise Robustness Eval · qhjqhj00Evaluates the robustness of end-to-end automatic speech recognition models to real-world acoustic distortions, including far-field reverberation, mixed sampling rates, low-bitrate codecs, and background noise at varying signal-to-noise ratios. Use when the user wants to benchmark on LibriSpeech, BUT ReverbDB, Hub5 Switchboard & CallHome, AISHELL-2, or asks about evaluating this task. Reports greedy WER (%).
- ▌ Atc Asr Domain Shift Eval · qhjqhj00This benchmark evaluates the robustness of self-supervised speech recognition models under domain shift in air traffic control communications. It probes few-shot fine-tuning capabilities, sensitivity to audio quality and accents, and potential gender bias in transcription performance. Use when the user wants to benchmark on NATS, ISAVIA, LiveATC-Test, ATCO2-Test, LDC-ATCC, UWB-ATCC, ATCOSIM, or asks about evaluating this task. Reports WER.
- ▌ Av Speech Separation Eval · qhjqhj00Evaluates a model's ability to separate target speaker speech from audio mixtures (noise or other speakers) using synchronized visual face cues. It probes speaker-independent audio-visual fusion and robustness to varying numbers of speakers and background noise. Use when the user wants to benchmark on AVSpeech, AudioSet, CHiME-2, Mandarin, TCD-TIMIT, CUAVE, or asks about evaluating this task. Reports SDR improvement.
- ▌ Backbone Fine Tuning Eval · qhjqhj00Evaluates the fine-tuning performance of lightweight, pre-trained CNN and attention-based backbones across diverse image classification domains, including natural images, remote sensing, medical histopathology, and plant imaging. It probes how well different architectures generalize under data-scarce conditions and whether ImageNet pre-training accuracy correlates with downstream task performance. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny ImageNet, Stanford Dogs, Flowers102, CUB200, Stanford Cars, DTD, UC Merced Land Use, EuroSAT, PlantVillage, PlantCLEF, Galaxy10, BreakHis, RSNA, Food-101, or asks about evaluating this task. Reports Top-1 classification accuracy.
- ▌ Bangla Math Olympiad Eval · qhjqhj00Evaluates large language models' ability to solve mathematical Olympiad problems in Bangla and English. It probes multilingual reasoning, step-by-step problem solving, and the impact of retrieval-augmented generation and fine-tuning on low-resource language math tasks. Use when the user wants to benchmark on BDMO dataset, Test dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Banglabook Sentiment Eval · qhjqhj00Evaluates the ability of models to classify Bangla book reviews into three sentiment categories (Positive, Neutral, Negative). It probes product-specific sentiment analysis in a low-resource language, testing both contextual understanding and robustness to class imbalance and lexical overlap. Use when the user wants to benchmark on BANGLABOOK, or asks about evaluating this task. Reports weighted average F1-score.
- ▌ Blimp Glue Superglue Eval · qhjqhj00Evaluates language understanding, linguistic acceptability, sentiment analysis, natural language inference, and factual reasoning under low-resource fine-tuning conditions. Use when the user wants to benchmark on BLiMP, GLUE (subset), SuperGLUE (subset), or asks about evaluating this task. Reports accuracy.
- ▌ Blink Vision Centric Eval · qhjqhj00This evaluation probes a model's ability to leverage raw visual representations for vision-centric tasks without relying on language priors or domain expertise. It tests pixel-level matching, depth perception, 3D object awareness, and art style recognition across multiple-choice and regression-style tasks. Use when the user wants to benchmark on CV-Bench (Depth Order), SPair-71k, FunKPoint, HPatches, MOCHI, WikiArt (BLINK Art Style), or asks about evaluating this task. Reports multiple-choice VQA.
- ▌ Bomjin Code Eval Octopack · qhjqhj00Compute bomjin/code_eval_octopack via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bomjin/code_eval_octopack.
- ▌ Brats Tcga Tumor Seg Eval · qhjqhj00Evaluates the capability of deep learning models to segment brain tumors from multi-modal MRI scans. It probes volumetric overlap accuracy and boundary localization precision across distinct tumor sub-regions (enhancing tumor, tumor core, whole tumor). Use when the user wants to benchmark on BraTS-Glioma (BraTS 2020), TCGA LGG, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
- ▌ Brats20 Segmentation Eval · qhjqhj00Evaluates the capability of 2D and 3D convolutional neural networks to segment brain tumor sub-regions (enhancing tumor, whole tumor, tumor core) from multi-modal volumetric MRI scans. It specifically probes the effectiveness of ImageNet pretraining and architectural extensions on segmentation accuracy and robustness across benchmark and private clinical data. Use when the user wants to benchmark on BraTS 2020, Syrian-Lebanese Hospital Clinical Dataset, or asks about evaluating this task. Reports Dice score.
- ▌ Breakout Determinism Eval · qhjqhj00Measures the sensitivity of deep Q-learning performance to various sources of nondeterminism (GPU operations, environment stochasticity, exploration seeds, weight initialization, minibatch sampling) by comparing performance variance across controlled experimental groups. Use when the user wants to benchmark on Atari BREAKOUT, or asks about evaluating this task. Reports mean score.
- ▌ Breast Mass Severity Eval · qhjqhj00This evaluation probes a model's ability to predict breast cancer severity (benign vs. malignant) using clinical and radiological features. It measures classification performance across multiple metrics to assess diagnostic reliability and clinical utility. Use when the user wants to benchmark on Mammographic mass dataset, or asks about evaluating this task. Reports Accuracy.
- ▌ Budget AI Researcher Eval · qhjqhj00Evaluates a retrieval-augmented generation framework's ability to synthesize novel, feasible, and interesting research abstracts by combining distant topics from AI conference literature. It probes long-range concept recombination and grounded ideation capabilities. Use when the user wants to benchmark on AI Conference Papers (ICLR, NeurIPS, ICML, ACL, ECCV), or asks about evaluating this task. Reports Novelty.
- ▌ Carla Counterfactual Eval · qhjqhj00Evaluates the quality and feasibility of model-agnostic counterfactual explanations generated for tabular data. It probes whether generated counterfactuals successfully flip classifier predictions while maintaining sparsity, proximity, actionability (immutable constraints), and plausibility across multiple binary classification tasks. Use when the user wants to benchmark on adult, COMPAS, Give Me Some Credit, HELOC, Irish, Saheart, Titanic, Wine, or asks about evaluating this task. Reports Success rate.
- ▌ Caser Sequential Rec Eval · qhjqhj00Evaluates a model's ability to capture sequential user behavior patterns for personalized top-N item recommendation. It probes the model's capacity to model temporal dependencies, skip behaviors, and union-level sequential patterns from historical interactions to predict future items. Use when the user wants to benchmark on MovieLens, Gowalla, Foursquare, Tmall, or asks about evaluating this task. Reports MAP.
- ▌ Cbm Concept Accuracy Eval · qhjqhj00Evaluates whether Concept Bottleneck Models learn semantically meaningful concept representations from input images under varying annotation granularity and concept correlation structures. Measures how well the model predicts intermediate concepts and downstream tasks compared to standard neural networks. Use when the user wants to benchmark on Playing cards, CheXpert, or asks about evaluating this task. Reports concept accuracy.
- ▌ Chexpert Atelectasis Eval · qhjqhj00Evaluates the diagnostic accuracy of predictive algorithms and human radiologists on chest X-ray images for detecting atelectasis. It specifically probes whether human expertise provides actionable, non-redundant information on input subsets where algorithms are algorithmically indistinguishable. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
- ▌ Clickbait Mitigation Eval · qhjqhj00This evaluation probes a recommender system's ability to mitigate clickbait by measuring performance exclusively on user interactions that result in positive post-click feedback (likes), rather than raw click-through rates. Use when the user wants to benchmark on Unspecified in provided section, or asks about evaluating this task. Reports post-click satisfaction (likes).
- ▌ Climate Segmentation Eval · qhjqhj00This evaluation probes pixel-level weather pattern segmentation (atmospheric rivers and tropical cyclones) from multi-channel climate data. It measures both segmentation accuracy and exascale training throughput/scaling efficiency across different network architectures and hardware configurations. Use when the user wants to benchmark on Climate weather pattern dataset, or asks about evaluating this task. Reports IoU.
- ▌ Clinical Turing Test Eval · qhjqhj00Evaluates the physiological realism and clinical fidelity of synthetic 12-lead ECGs by measuring how often expert clinicians can correctly distinguish them from real clinical recordings, and how accurately they can diagnose specific pathologies in both synthetic and real signals. Use when the user wants to benchmark on MedalCare-XL, or asks about evaluating this task. Reports accuracy.
- ▌ Closp Crisislandmark Eval · qhjqhj00Evaluates cross-modal retrieval and zero-shot classification capabilities for remote sensing imagery (SAR and multispectral optical) paired with text descriptions. It probes how well unified semantic embeddings align heterogeneous geospatial data with natural language for crisis event and land cover analysis. Use when the user wants to benchmark on CrisisLandMark, or asks about evaluating this task. Reports nDCG@1000.
- ▌ Cloudops Forecasting Eval · qhjqhj00Evaluates time series forecasting models, particularly pre-trained Transformers, on cloud operations data. It probes zero-shot generalization, architectural efficiency, and scaling behavior against classical and deep learning baselines. Use when the user wants to benchmark on azure2017, borg2011, ali2018, or asks about evaluating this task. Reports sMAPE.
- ▌ Conformal Prediction Eval · qhjqhj00Evaluates the ability of conformal prediction frameworks to produce statistically valid prediction sets with instance-level uncertainty quantification for encoder-only transformers, measuring both classification accuracy and calibration efficiency across standard NLP benchmarks. Use when the user wants to benchmark on GLUE, SuperGLUE, or asks about evaluating this task. Reports Test Accuracy.
- ▌ Continual Multimodal Eval · qhjqhj00This benchmark evaluates a model's ability to sequentially learn a mix of visual understanding and generation tasks without catastrophically forgetting previously acquired knowledge. It specifically probes intra-modal retention (maintaining performance on earlier tasks) and inter-modal stability (preventing updates for one modality from degrading the other). Use when the user wants to benchmark on ScienceQA, TextVQA, GQA, VizWiz, ImageNet, CustomConcept101, or asks about evaluating this task. Reports Average Accuracy (ACC).
- ▌ Counterfactual Chaos Eval · qhjqhj00Evaluates the reliability of counterfactual trajectory estimation in chaotic versus non-chaotic dynamical systems under parameter uncertainty and observational noise. It probes whether Bayesian filtering and particle-based smoothing can accurately recover 'what-if' scenarios when small initial perturbations lead to divergent outcomes. Use when the user wants to benchmark on Lorenz System, Rössler System, Logistic Growth, or asks about evaluating this task. Reports RMSE_t.
- ▌ Cross Domain Meta Dl Eval · qhjqhj00Probes few-shot image classification generalization across diverse domains and highly variable task regimes (2–20 ways, 1–20 shots) without relying on pre-trained backbones. Use when the user wants to benchmark on Meta-Album, or asks about evaluating this task. Reports accuracy.
- ▌ Cross Lingual F5 Tts Eval · qhjqhj00Evaluates the intelligibility, speaker similarity, and naturalness of synthesized speech in cross-lingual voice cloning and TTS scenarios. It also measures the accuracy of a language-agnostic speaking rate predictor for duration modeling across multiple languages. Use when the user wants to benchmark on Emilia, Seed-TTS-eval, LibriSpeech-PC test-clean, FLEURS, or asks about evaluating this task. Reports WER.
- ▌ Csrec Sequential Rec Eval · qhjqhj00This evaluation protocol assesses the ranking performance and robustness of sequential recommendation models trained with confident soft labels. It measures whether predicted item sequences align with actual user interactions and verifies if recommendations correspond to genuinely positive user preferences using explicit rating thresholds. Use when the user wants to benchmark on Last.FM, Yelp, Amazon Electronics, Amazon Movies and TV, or asks about evaluating this task. Reports Recall@n, NDCG@n.
- ▌ Cultural Positioning Eval · qhjqhj00Evaluates whether an LLM's value profile aligns with specific cultural norms using World Values Survey data, and tests the model's steerability when provided with diverse cultural contexts. It probes the extent to which constitutional AI codifies dominant cultural biases and resists prompt-based cultural adaptation. Use when the user wants to benchmark on World Values Survey (WVS) Wave 7, or asks about evaluating this task. Reports Pearson correlation.
- ▌ Cvr Ctcvr Estimation Eval · qhjqhj00Evaluates the ranking performance of models for click-through rate (CTR) and post-click conversion rate (CVR) estimation in recommendation systems. It probes the model's ability to correctly rank items by their predicted probability of conversion, while mitigating sample selection bias and false independence assumptions between clicks and conversions. Use when the user wants to benchmark on Industrial Benchmark, Ali-CCP, or asks about evaluating this task. Reports AUC.
- ▌ D2sac Asp Scheduling Eval · qhjqhj00Evaluates a diffusion-based reinforcement learning agent's capability to dynamically assign AI-generated content tasks to edge service providers under stochastic workloads, optimizing for user utility while preventing system crashes and minimizing training time. Use when the user wants to benchmark on Custom AIGC Edge Simulation Environment, Gym Benchmark Tasks, or asks about evaluating this task. Reports Cumulative Reward.
- ▌ Dctracks Track Recon Eval · qhjqhj00Evaluates the performance of machine learning and traditional algorithms for reconstructing particle tracks in drift chamber detectors. It probes hit-level matching accuracy, track-level reconstruction efficiency, charge identification correctness, and momentum resolution under realistic detector conditions. Use when the user wants to benchmark on DCTracks, or asks about evaluating this task. Reports track efficiency.
- ▌ Deepfake Speech Auth Eval · qhjqhj00Evaluates the robustness of audio-based biometric authentication systems against deepfake speech synthesis attacks. It measures how easily voice cloning models can bypass speaker verification and how effectively anti-spoofing detectors can distinguish genuine from synthetic speech. Use when the user wants to benchmark on AISHELL-3, or asks about evaluating this task. Reports Bypass Rate.
- ▌ Deepurban Trajectory Eval · qhjqhj00Evaluates trajectory prediction and planning capabilities in high-density urban environments with significant vehicle-to-vulnerable-road-user interactions. It measures prediction accuracy and safety compliance using displacement errors and collision scores. Use when the user wants to benchmark on DeepUrban, or asks about evaluating this task. Reports ADE.
- ▌ Sm3 Text To Query Eval · qhjqhj00Evaluates text-to-query systems across relational, document, and graph database models using four query languages (SQL, MQL, Cypher, SPARQL). It probes the ability of models to translate natural language medical questions into correct, executable database queries using standardized SNOMED-CT aligned synthetic patient data. Use when the user wants to benchmark on SM3-Text-to-Query, or asks about evaluating this task. Reports correctness.
- ▌ Social Media Bias Eval · qhjqhj00This benchmark evaluates the ability of models to automatically detect multiple dimensions of media bias (e.g., hate speech, racial, gender, political, linguistic, and text-level context bias) in social media posts across different topic domains. It probes a model's robustness to domain shift and severe class imbalance in multi-label bias identification tasks. Use when the user wants to benchmark on Social Media Bias Dataset (YouTube & Reddit), or asks about evaluating this task. Reports weighted average F1 score.
- ▌ Soft Pairwise Accuracy · qhjqhj00Evaluates the reliability and discriminative power of automatic machine translation metrics by comparing their statistical significance against human MQM judgments. It measures how well a metric's pairwise system rankings align with human preferences using permutation-based p-values rather than hard binary decisions. Use when the user has predictions and gold and needs to compute Soft Pairwise Accuracy (SPA).
- ▌ Sparrow Alignment Eval · qhjqhj00Evaluates the alignment, factual grounding, and rule-following capabilities of dialogue agents through human preference comparisons. It also measures resilience to adversarial probing for specific harm rules and the quality of evidence-supported responses. Use when the user wants to benchmark on ELI5 + Free Dialogue Test Set, or asks about evaluating this task. Reports Three-model preference rate.
- ▌ Sparse Gpu Kernel Eval · qhjqhj00Evaluates the performance and efficiency of custom sparse GPU kernels for SpMM and SDDMM operations against standard libraries like cuSPARSE on deep learning workloads. It measures computational throughput, memory usage, and end-to-end speedups across various model architectures and batch sizes. Use when the user wants to benchmark on Sparse Matrix Dataset from DNNs, or asks about evaluating this task. Reports Geometric mean speedup.
- ▌ Spatial Reasoning Eval · qhjqhj00Evaluates vision-language models' ability to count objects and reason about spatial relationships (depth, distance, relative position) in images. It probes segmentation capabilities, attention alignment, and robustness to linguistic variations (out-of-distribution shifts). Use when the user wants to benchmark on CLEVR_CoGenT_ValB, CVBench, Pixmo-Count, Static Spatial Reasoning (SAT), VSR, VC Bench, or asks about evaluating this task. Reports Accuracy.
- ▌ Spatialdistortionindex · qhjqhj00Compute the SpatialDistortionIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpatialDistortionIndex, or asks how to score with SpatialDistortionIndex.
- ▌ Speech Likelihood Eval · qhjqhj00Evaluates the ability of generative latent variable models and autoregressive baselines to model speech audio distributions at varying temporal resolutions. It measures how well models capture intra-frame and inter-frame correlations in audio waveforms by optimizing likelihood objectives. Use when the user wants to benchmark on TIMIT, LibriSpeech, or asks about evaluating this task. Reports bits per frame (bpf).
- ▌ Speech Separation Eval · qhjqhj00Evaluates speech separation models' ability to isolate individual speaker signals from multi-speaker mixtures under various acoustic conditions, including moving sources, environmental noise, and musical noise. It measures both objective signal quality and subjective perceptual metrics to assess generalization from synthetic to real-world dynamic scenarios. Use when the user wants to benchmark on SonicSet, RealSEP, HumanSEP, LRS2-2Mix, Libri2Mix, or asks about evaluating this task. Reports SI-SNR.
- ▌ Speechmentalmanip Eval · qhjqhj00Binary classification of spoken multi-speaker dialogues to detect the presence of mental manipulation tactics. It probes an audio-language model's ability to identify subtle manipulative cues in synthetic speech without relying on text transcripts. Use when the user wants to benchmark on SpeechMentalManip, or asks about evaluating this task. Reports accuracy.
- ▌ Stackoverflow Ner Eval · qhjqhj00Evaluates named entity recognition capabilities on software programming texts. It specifically probes the model's ability to identify fine-grained code-related entities like variable names, libraries, and data structures in StackOverflow posts. Use when the user wants to benchmark on StackOverflow NER corpus, or asks about evaluating this task. Reports F1.
- ▌ Stella Living Lab Eval · qhjqhj00Evaluates academic search and recommendation systems in live production environments using A/B testing and user interaction logs, bridging the gap between offline test collections and real-world performance. Use when the user wants to benchmark on LIVIVO, GESIS Search, or asks about evaluating this task. Reports click-paths.
- ▌ Stochastic Ackley Eval · qhjqhj00Evaluates the ability of uncertainty-aware deep neural networks to approximate a highly irregular, multi-extremum function and quantify predictive uncertainty. It specifically probes how well the models handle in-distribution versus out-of-distribution parameter regimes. Use when the user wants to benchmark on Stochastic Ackley Function, or asks about evaluating this task. Reports Relative Error (RE).
- ▌ Success Rate 3d Policy · qhjqhj00Evaluates the robustness and generalization of 3D policy learning models for robotic manipulation across varying environmental conditions, temporal horizons, and real-world interference. It probes spatial understanding, fine-grained pose control, and resilience to domain randomization and lighting changes. Use when the user wants to benchmark on RoboTwin 2.0, ManiSkill2, Real-World Manipulation, or asks about evaluating this task. Reports Success Rate (%).
- ▌ Superb Downstream Eval · qhjqhj00Evaluates pre-trained speech models on downstream spoken language understanding tasks. It probes the model's ability to classify spoken intents, fill semantic slots in transcriptions, and detect specific keywords in audio. Use when the user wants to benchmark on SUPERB, or asks about evaluating this task. Reports test accuracy.
- ▌ Svld Points Ratio Eval · qhjqhj00Probes a model's ability to predict social engagement (upvote ratio) from multimodal inputs (images, videos, and text). It evaluates cross-modal fusion and regression capabilities on socially grounded, context-rich data. Use when the user wants to benchmark on SVLD, or asks about evaluating this task. Reports Mean L1point ratio prediction error.
- ▌ Synthetic Geology Eval · qhjqhj00Evaluates a flow matching generative model's ability to produce realistic 3D subsurface geological models, both unconditionally and conditioned on sparse borehole data. It probes the model's capacity for geological interpolation, structural feature reconstruction, and probabilistic uncertainty estimation. Use when the user wants to benchmark on Synthetic Geology / StructuralGeo Dataset, or asks about evaluating this task. Reports probabilistic confidence intervals.
- ▌ Synthetic4relight Eval · qhjqhj00Evaluates a 3D scene representation's capability for novel view synthesis, relighting, and inverse rendering (estimating diffuse albedo and roughness) from posed RGB images. Use when the user wants to benchmark on Synthetic4Relight, or asks about evaluating this task. Reports PSNR.
- ▌ Tabular Benchmark Eval · qhjqhj00Evaluates the predictive performance and stability of 32 deep learning and tree-based tabular models across a large collection of diverse tabular datasets. It probes how well different architectures handle classification and regression tasks, and how dataset characteristics influence method rankings. Use when the user wants to benchmark on LAMDA-TALENT Benchmark, or asks about evaluating this task. Reports average_rank.
- ▌ Tally Gpu Sharing Eval · qhjqhj00Evaluates the performance isolation and resource sharing capabilities of GPU scheduling systems for concurrent deep learning workloads. It probes how well a system maintains tail latency for high-priority inference tasks while maximizing throughput for best-effort training tasks under varying traffic loads and workload combinations. Use when the user wants to benchmark on Tally Benchmark Suite, or asks about evaluating this task. Reports 99th-percentile latency.
- ▌ Tamper Resistance Eval · qhjqhj00This benchmark evaluates the robustness of LLM safety safeguards against fine-tuning-based tampering attacks. It measures whether a model can maintain low accuracy on weaponized knowledge (forget) and high benign capabilities (retain) after undergoing various supervised fine-tuning attacks. Use when the user wants to benchmark on WMDP, MMLU, HarmBench, MT-Bench, or asks about evaluating this task. Reports Post-Attack Forget accuracy, Attack Success Rate (ASR).
- ▌ Tb Classification Eval · qhjqhj00Evaluates deep learning models' ability to classify chest X-rays as tuberculosis or normal, comparing whole-image vs. lung-segmented inputs. Use when the user wants to benchmark on Kaggle CXR images and lung mask dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Theory Of Mind QA Eval · qhjqhj00Evaluates a model's ability to track first-order and second-order false beliefs, distinguishing an agent's mental state from physical reality and memory. It probes whether systems can maintain consistent world-state representations when agents hold incorrect beliefs about object locations or events. Use when the user wants to benchmark on Sally-Anne & Icecream Van Tasks, or asks about evaluating this task. Reports accuracy.
- ▌ Thinkjepa Ego Dex Eval · qhjqhj00Evaluates a model's ability to forecast future 3D hand/joint trajectories and latent video representations from egocentric video inputs. It probes long-horizon temporal consistency and physical plausibility in dexterous manipulation scenarios. Use when the user wants to benchmark on EgoDex, EgoExo4D, or asks about evaluating this task. Reports ADE.
- ▌ Thyme Scene Graph Eval · qhjqhj00Evaluates a model's ability to generate dynamic video scene graphs by predicting inter-object relationships and attributes across multiple temporal frames. It specifically probes temporal consistency, handling of occlusions, and modeling of long-range dependencies in both ground-level and aerial video footage. Use when the user wants to benchmark on ASPIRe, AeroEye-v1.0, or asks about evaluating this task. Reports R@20, mR@20.
- ▌ Timer Time Series Eval · qhjqhj00Evaluates a large decoder-only Transformer model on standard time series benchmarks for forecasting, imputation, and anomaly detection. It probes the model's few-shot generalization, scalability, and robustness in data-scarce scenarios compared to encoder-only baselines. Use when the user wants to benchmark on ETT, ECL, Traffic, Weather, PEMS, UCR Anomaly Archive, or asks about evaluating this task. Reports MSE.
- ▌ Topofair Fairness Eval · qhjqhj00Evaluates fairness-aware link prediction models on synthetic graphs with controlled topological biases. It probes how structural properties like assortativity, heterogeneity, and class imbalance impact fairness metrics (SP, EO) and predictive accuracy (Hit@10, AUC). Use when the user wants to benchmark on Opinion use case, Friendship use case, Collab use case, Real datasets (Collab, Polblogs, Facebook), or asks about evaluating this task. Reports Statistical Parity (SP), Equalized Odds (EO).
- ▌ Toxicity Analysis Eval · qhjqhj00Measures the toxicity of text generated by language models conditioned on specific prompts, evaluating how alignment techniques like prompting and context distillation affect harmful content generation. Use when the user wants to benchmark on RealToxicityPrompts, or asks about evaluating this task. Reports mean toxicity score.
- ▌ Tracknet Tracking Eval · qhjqhj00Evaluates the ability of deep learning models to detect and track high-speed, tiny objects (tennis and badminton balls) in broadcast sports videos. It probes robustness to motion blur, occlusion, and domain shifts by comparing single-frame vs. multi-frame tracking and transfer learning across different sports. Use when the user wants to benchmark on Tennis, Badminton, or asks about evaluating this task. Reports F1-measure.
- ▌ Train O Matic Wsd Eval · qhjqhj00Evaluates the quality of automatically generated multilingual word sense disambiguation (WSD) training corpora by training a supervised WSD system (IMS) on them and measuring performance on standard WSD benchmark datasets. It probes whether synthetic sense-annotated data can match or exceed manually annotated corpora, particularly for low-resource languages. Use when the user wants to benchmark on Senseval-2, Senseval-3, SemEval-2007, SemEval-2013, SemEval-2015, or asks about evaluating this task. Reports F1.
- ▌ Trilemma Of Truth Eval · qhjqhj00Evaluates large language models' ability to distinguish factually true statements from factually false and unverifiable ('neither') statements. It probes both prompt-based output probabilities and internal hidden activations to measure veracity classification accuracy and uncertainty quantification. Use when the user wants to benchmark on Trilemma of Truth Datasets, or asks about evaluating this task. Reports MCC.
- ▌ Tucano2 Portfolio Eval · qhjqhj00This evaluation protocol probes the language understanding, reasoning, and instruction-following capabilities of Portuguese LLMs across diverse domains including academic exams, natural language inference, physical commonsense, and code generation. It is specifically designed to provide reliable training signals during pretraining and assess post-training alignment. Use when the user wants to benchmark on ARC Challenge, Calame, Global PIQA, HellaSwag, LAMBADA, ENEM, BLUEX, OAB, Belebele, MMLU, IFEval-PT, GSM8K-PT, RULER-PT, HumanEval, or asks about evaluating this task. Reports accuracy (log-likelihood selection).
- ▌ United Medasr Asr Eval · qhjqhj00This evaluation measures the transcription accuracy of a fine-tuned automatic speech recognition model across four diverse speech benchmarks. It specifically probes the model's robustness to different speaking styles, accents, and linguistic contexts after applying a noise reduction step and a BART-based semantic correction pipeline. Use when the user wants to benchmark on LibriSpeech, Europarl-ASR, TED-LIUM, FLEURS, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Unlearning Recsys Eval · qhjqhj00Evaluates the ability of recommender systems to efficiently remove specific user interactions or sensitive items (unlearning) while preserving recommendation utility. It probes real-world operational constraints, including handling sequential small-batch deletion requests, domain-specific triggers, and low-latency execution across collaborative filtering, session-based, and next-basket recommendation tasks. Use when the user wants to benchmark on TaFeng, Dunnhumby, Instacart, RSC15, DIGI, NOWP, Goodreads, MovieLens, Amazon Reviews, or asks about evaluating this task. Reports Recall, PHR, nDCG.
- ▌ Urban Pathfinding Eval · qhjqhj00Evaluates real-time urban pathfinding algorithms under dynamic traffic and weather conditions. It measures how well traditional graph search methods and deep learning models predict optimal routes and minimize travel time in a simulated Berlin city environment. Use when the user wants to benchmark on Berlin Urban Simulation, or asks about evaluating this task. Reports Average Travel Time (s).
- ▌ Usb Summarization Eval · qhjqhj00Evaluates multiple text summarization capabilities including extractive/abstractive generation, factuality verification, factual error correction, topic-constrained generation, sentence compression, evidence extraction, and unsupported span detection across diverse domains. Use when the user wants to benchmark on Extractive Summarization (EXT), Abstractive Summarization (ABS), Factuality Classification (FAC), Fixing Factuality (FIX), Topic-based Summarization (TOPIC), Multi-sentence Compression (COMP), Evidence Extraction (EVEXT), Unsupported Span Prediction (UNSUP), or asks about evaluating this task. Reports ROUGE.
- ▌ Video Outpainting Eval · qhjqhj00Evaluates a model's ability to generate spatially and temporally consistent video content outside the original frame boundaries (video outpainting), while preserving source structure and visual realism. Use when the user wants to benchmark on DAVIS 2017, YouTube-VOS, or asks about evaluating this task. Reports FVD.
- ▌ Videogameqa Bench Eval · qhjqhj00Evaluates vision-language models on video game quality assurance tasks, including glitch detection, temporal reasoning, and bug reporting. It probes the model's ability to process sampled video frames, identify visual anomalies, and generate structured or descriptive reports about game glitches. Use when the user wants to benchmark on VideoGameQA-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Viona Fuzzy Reordering · qhjqhj00Compute Viona/fuzzy_reordering via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Viona/fuzzy_reordering.
- ▌ Visual Genome Sgg Eval · qhjqhj00Evaluates fine-grained scene graph generation by predicting subject-predicate-object triplets from images. It measures recall and F1 scores across head, body, and tail predicate classes to assess performance on long-tailed distributions and missing annotations. Use when the user wants to benchmark on Visual Genome, or asks about evaluating this task. Reports mR@K, F@K.
- ▌ Visual Sycophancy Eval · qhjqhj00This evaluation probes how Vision-Language Models ground their responses in visual input versus relying on language priors or user bias. It measures perceptual awareness, visual dependency, and alignment conflicts by comparing model behavior across original, blank, noisy, and semantically conflicting images. Use when the user wants to benchmark on GQA, VQAv2, A-OKVQA, POPE, or asks about evaluating this task. Reports VNS.
- ▌ Vln Task Planning Eval · qhjqhj00Evaluates an agent's ability to decompose coarse-grained natural language navigation instructions into executable subtasks and navigate through simulated environments to reach target locations or interact with objects. It probes task planning, visual-language grounding, and dynamic error recovery in continuous or discrete navigation spaces. Use when the user wants to benchmark on R2R, REVERIE, ALFRED, or asks about evaluating this task. Reports Success Rate (SR).
- ▌ Vos Language Referring · qhjqhj00Evaluates a model's ability to perform pixel-level video object segmentation guided by natural language referring expressions, testing both language grounding and temporal consistency in dynamic scenes. Use when the user wants to benchmark on DAVIS-16, DAVIS-17, or asks about evaluating this task. Reports performance score.
- ▌ Vqa Cot Reasoning Eval · qhjqhj00Probes the ability of vision-language models to perform multi-step chain-of-thought reasoning on visual inputs across diverse domains like charts, documents, science diagrams, and math. It measures both direct answer accuracy and structured reasoning accuracy. Use when the user wants to benchmark on A-OKVQA, ChartQA, DocVQA, InfoVQA, TextVQA, AI2D, ScienceQA, MathVista, OCRBench, MMStar, MMMU, or asks about evaluating this task. Reports accuracy.
- ▌ Whisper Zero Shot Eval · qhjqhj00Evaluates the zero-shot generalization capability of a speech recognition model across diverse English and multilingual domains. It measures robustness to out-of-distribution audio, varying noise levels, and translation tasks without any dataset-specific fine-tuning. Use when the user wants to benchmark on LibriSpeech, Common Voice, Fleurs, CoVoST2, Multilingual LibriSpeech (MLS), VoxPopuli, or asks about evaluating this task. Reports WER.
- ▌ Widget Captioning Eval · qhjqhj00This benchmark evaluates a model's ability to generate natural language descriptions for individual mobile UI elements using multimodal inputs. It probes the capability to fuse visual appearance and structural hierarchy data to produce accurate, context-aware captions for accessibility and UI understanding tasks. Use when the user wants to benchmark on Widget Captioning Dataset, or asks about evaluating this task. Reports CIDEr.
- ▌ Window Wise Complexity · qhjqhj00Quantifies the intrinsic predictability of time series data by measuring window-wise pattern complexity in the frequency domain. It establishes a data-driven performance lower bound for forecasting models and identifies whether standard benchmarks have reached saturation. Use when the user has predictions and gold and needs to compute window-wise complexity.
- ▌ Zebrafish Sctrans Eval · qhjqhj00Evaluates the ability of topological data analysis methods to capture developmental transitions and cell lineage dynamics in single-cell RNA sequencing time-series data. It specifically tests whether higher-order simplicial complexity can outperform conventional topological invariants like Betti numbers in identifying critical biological stages. Use when the user wants to benchmark on Farrell et al. (2018) zebrafish scRNA-seq, or asks about evaluating this task. Reports normalized simplicial complexity.
- ▌ 3d 2d Vl Grounding Eval · qhjqhj00Evaluates a unified vision-language model's ability to ground natural language instructions to 3D objects and 2D regions, as well as answer 3D visual questions. It probes spatial reasoning, cross-modal alignment, and robustness to different 3D input representations (mesh-sampled vs. sensor RGB-D point clouds). Use when the user wants to benchmark on SR3D, NR3D, ScanRefer, RefCOCO, RefCOCO+, RefCOCOg, ScanQA, SQA3D, or asks about evaluating this task. Reports top-1 accuracy (Acc@25/50/75).
- ▌ 3d Shape Retrieval Eval · qhjqhj00Evaluates algorithms for retrieving geometrically and topologically similar 3D shapes from a database given a query shape. Probes pose invariance, shape descriptor robustness, and retrieval ranking accuracy. Use when the user wants to benchmark on NIST shape benchmark, or asks about evaluating this task. Reports precision-recall.
- ▌ Ade20k Scene Parse Eval · qhjqhj00Evaluates a model's ability to perform dense pixel-wise semantic segmentation across 150 common scene categories, including both discrete objects and amorphous 'stuff' classes. It probes fine-grained scene understanding and the model's capacity to handle class imbalance and varying object scales. Use when the user wants to benchmark on SceneParse150, or asks about evaluating this task. Reports Mean IoU.
- ▌ Adjustedmutualinfoscore · qhjqhj00Compute the AdjustedMutualInfoScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute AdjustedMutualInfoScore, or asks how to score with AdjustedMutualInfoScore.
- ▌ Amp Classification Eval · qhjqhj00Evaluates the ability of reprogrammed language models to classify antimicrobial peptide (AMP) sequences into binary categories (toxic vs. non-toxic, or AMP vs. non-AMP) using limited labeled data. Use when the user wants to benchmark on AMP Dataset, or asks about evaluating this task. Reports Test Accuracy.
- ▌ Amp Motion Control Eval · qhjqhj00Evaluates a physics-based character's ability to learn stylized locomotion and complex task execution (e.g., navigating targets, avoiding obstacles) by imitating unstructured motion datasets. It probes the model's capacity to compose disparate skills, generalize across gaits, and maintain high-fidelity motion tracking without manual motion planning. Use when the user wants to benchmark on AMP Motion Datasets, or asks about evaluating this task. Reports normalized task return.
- ▌ Antibody Domainbed Eval · qhjqhj00Evaluates out-of-distribution generalization of protein language models and sequence CNNs for therapeutic antibody design across different antigen targets and generative models. It probes robustness to covariate shifts, label shifts, and assay biases in molecular sequence data. Use when the user wants to benchmark on Antibody DomainBed, or asks about evaluating this task. Reports accuracy.
- ▌ Average Precision Score · qhjqhj00Compute the average_precision_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute average_precision_score, or asks how to score with average_precision_score.
- ▌ Balanced Accuracy Score · qhjqhj00Compute the balanced_accuracy_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute balanced_accuracy_score, or asks how to score with balanced_accuracy_score.
- ▌ Bayes Factor Odds Ratio · qhjqhj00Evaluates the likelihood of different binary black hole formation channels (CEE, CHE, SMT) given gravitational wave strain data by comparing Bayesian evidence and prior odds. Use when the user has predictions and gold and needs to compute Bayes factor ($\mathcal{B}$), Odds ratio ($\mathcal{O}$).
- ▌ Benchmark Accuracy Eval · qhjqhj00Evaluates zero-shot language model performance across a suite of 10 standard NLP benchmarks covering commonsense reasoning, science QA, and language modeling. It measures task accuracy and correlates it with word-level statistical overlap metrics to assess distributional alignment between pre-training data and evaluation sets. Use when the user wants to benchmark on ARC Easy, ARC Challenge, Hellaswag, MMLU, SciQ, OpenBookQA, PIQA, lambada, SocialIQA, SWAG, or asks about evaluating this task. Reports accuracy.
- ▌ Biosage Scientific Eval · qhjqhj00Evaluates a compound AI architecture's ability to retrieve, synthesize, and reason across cross-disciplinary scientific knowledge. It probes performance on established single-domain science benchmarks and a novel benchmark specifically designed for bio-AI cross-domain synthesis and reasoning. Use when the user wants to benchmark on LitQA2, GPQA, WMDP, HLE-Bio, BioSage Cross-Disciplinary Benchmark, or asks about evaluating this task. Reports accuracy.
- ▌ Bowdbeg Matching Series · qhjqhj00Compute bowdbeg/matching_series via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bowdbeg/matching_series.
- ▌ Brats Segmentation Eval · qhjqhj00Evaluates the precision of brain tumor segmentation models on multi-modal MRI scans across three clinically relevant regions (Whole Tumor, Tumor Core, Enhancing Tumor). It probes the model's ability to accurately delineate heterogeneous tumor boundaries and correctly identify positive tumor voxels in medical imaging data. Use when the user wants to benchmark on BraTS2019/2020, or asks about evaluating this task. Reports Dice coefficient.
- ▌ Brisc Segmentation Eval · qhjqhj00Evaluates deep learning models on multi-planar brain tumor segmentation from contrast-enhanced T1-weighted MRI scans. It probes multi-scale feature integration, cross-view generalization, and robustness to class imbalance across glioma, meningioma, and pituitary tumor types. Use when the user wants to benchmark on BRISC, or asks about evaluating this task. Reports mIoU.