all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 25 of 76

  1. ▌
    Apres Paper Revision Eval · qhjqhj00
    This protocol evaluates an LLM's ability to predict a paper's future scientific impact based on its text and peer reviews, and its ability to iteratively revise the manuscript to maximize that predicted impact. It probes the model's capacity for rubric discovery, agentic text editing, and alignment with human expert preferences. Use when the user wants to benchmark on ICLR & NeurIPS Peer Review Dataset, or asks about evaluating this task. Reports MAE, Improvement Score ($\Delta S$).
    3 repo stars
  2. ▌
    Ares Android Testing Eval · qhjqhj00
    Evaluates the ability of automated black-box testing tools and reinforcement learning agents to explore Android applications effectively. It probes how well algorithms navigate complex UI states, maximize code/activity coverage, and trigger unique application crashes within a fixed time budget. Use when the user wants to benchmark on F-Droid top starred apps, AndroTest, Synthetic FATE models, or asks about evaluating this task. Reports AUC.
    3 repo stars
  3. ▌
    Asr Noise Robustness Eval · qhjqhj00
    Evaluates the robustness of end-to-end automatic speech recognition models to real-world acoustic distortions, including far-field reverberation, mixed sampling rates, low-bitrate codecs, and background noise at varying signal-to-noise ratios. Use when the user wants to benchmark on LibriSpeech, BUT ReverbDB, Hub5 Switchboard & CallHome, AISHELL-2, or asks about evaluating this task. Reports greedy WER (%).
    3 repo stars
  4. ▌
    Atc Asr Domain Shift Eval · qhjqhj00
    This benchmark evaluates the robustness of self-supervised speech recognition models under domain shift in air traffic control communications. It probes few-shot fine-tuning capabilities, sensitivity to audio quality and accents, and potential gender bias in transcription performance. Use when the user wants to benchmark on NATS, ISAVIA, LiveATC-Test, ATCO2-Test, LDC-ATCC, UWB-ATCC, ATCOSIM, or asks about evaluating this task. Reports WER.
    3 repo stars
  5. ▌
    Av Speech Separation Eval · qhjqhj00
    Evaluates a model's ability to separate target speaker speech from audio mixtures (noise or other speakers) using synchronized visual face cues. It probes speaker-independent audio-visual fusion and robustness to varying numbers of speakers and background noise. Use when the user wants to benchmark on AVSpeech, AudioSet, CHiME-2, Mandarin, TCD-TIMIT, CUAVE, or asks about evaluating this task. Reports SDR improvement.
    3 repo stars
  6. ▌
    Backbone Fine Tuning Eval · qhjqhj00
    Evaluates the fine-tuning performance of lightweight, pre-trained CNN and attention-based backbones across diverse image classification domains, including natural images, remote sensing, medical histopathology, and plant imaging. It probes how well different architectures generalize under data-scarce conditions and whether ImageNet pre-training accuracy correlates with downstream task performance. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny ImageNet, Stanford Dogs, Flowers102, CUB200, Stanford Cars, DTD, UC Merced Land Use, EuroSAT, PlantVillage, PlantCLEF, Galaxy10, BreakHis, RSNA, Food-101, or asks about evaluating this task. Reports Top-1 classification accuracy.
    3 repo stars
  7. ▌
    Bangla Math Olympiad Eval · qhjqhj00
    Evaluates large language models' ability to solve mathematical Olympiad problems in Bangla and English. It probes multilingual reasoning, step-by-step problem solving, and the impact of retrieval-augmented generation and fine-tuning on low-resource language math tasks. Use when the user wants to benchmark on BDMO dataset, Test dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  8. ▌
    Banglabook Sentiment Eval · qhjqhj00
    Evaluates the ability of models to classify Bangla book reviews into three sentiment categories (Positive, Neutral, Negative). It probes product-specific sentiment analysis in a low-resource language, testing both contextual understanding and robustness to class imbalance and lexical overlap. Use when the user wants to benchmark on BANGLABOOK, or asks about evaluating this task. Reports weighted average F1-score.
    3 repo stars
  9. ▌
    Blimp Glue Superglue Eval · qhjqhj00
    Evaluates language understanding, linguistic acceptability, sentiment analysis, natural language inference, and factual reasoning under low-resource fine-tuning conditions. Use when the user wants to benchmark on BLiMP, GLUE (subset), SuperGLUE (subset), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Blink Vision Centric Eval · qhjqhj00
    This evaluation probes a model's ability to leverage raw visual representations for vision-centric tasks without relying on language priors or domain expertise. It tests pixel-level matching, depth perception, 3D object awareness, and art style recognition across multiple-choice and regression-style tasks. Use when the user wants to benchmark on CV-Bench (Depth Order), SPair-71k, FunKPoint, HPatches, MOCHI, WikiArt (BLINK Art Style), or asks about evaluating this task. Reports multiple-choice VQA.
    3 repo stars
  11. ▌
    Bomjin Code Eval Octopack · qhjqhj00
    Compute bomjin/code_eval_octopack via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bomjin/code_eval_octopack.
    3 repo stars
  12. ▌
    Brats Tcga Tumor Seg Eval · qhjqhj00
    Evaluates the capability of deep learning models to segment brain tumors from multi-modal MRI scans. It probes volumetric overlap accuracy and boundary localization precision across distinct tumor sub-regions (enhancing tumor, tumor core, whole tumor). Use when the user wants to benchmark on BraTS-Glioma (BraTS 2020), TCGA LGG, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
    3 repo stars
  13. ▌
    Brats20 Segmentation Eval · qhjqhj00
    Evaluates the capability of 2D and 3D convolutional neural networks to segment brain tumor sub-regions (enhancing tumor, whole tumor, tumor core) from multi-modal volumetric MRI scans. It specifically probes the effectiveness of ImageNet pretraining and architectural extensions on segmentation accuracy and robustness across benchmark and private clinical data. Use when the user wants to benchmark on BraTS 2020, Syrian-Lebanese Hospital Clinical Dataset, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  14. ▌
    Breakout Determinism Eval · qhjqhj00
    Measures the sensitivity of deep Q-learning performance to various sources of nondeterminism (GPU operations, environment stochasticity, exploration seeds, weight initialization, minibatch sampling) by comparing performance variance across controlled experimental groups. Use when the user wants to benchmark on Atari BREAKOUT, or asks about evaluating this task. Reports mean score.
    3 repo stars
  15. ▌
    Breast Mass Severity Eval · qhjqhj00
    This evaluation probes a model's ability to predict breast cancer severity (benign vs. malignant) using clinical and radiological features. It measures classification performance across multiple metrics to assess diagnostic reliability and clinical utility. Use when the user wants to benchmark on Mammographic mass dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  16. ▌
    Budget AI Researcher Eval · qhjqhj00
    Evaluates a retrieval-augmented generation framework's ability to synthesize novel, feasible, and interesting research abstracts by combining distant topics from AI conference literature. It probes long-range concept recombination and grounded ideation capabilities. Use when the user wants to benchmark on AI Conference Papers (ICLR, NeurIPS, ICML, ACL, ECCV), or asks about evaluating this task. Reports Novelty.
    3 repo stars
  17. ▌
    Carla Counterfactual Eval · qhjqhj00
    Evaluates the quality and feasibility of model-agnostic counterfactual explanations generated for tabular data. It probes whether generated counterfactuals successfully flip classifier predictions while maintaining sparsity, proximity, actionability (immutable constraints), and plausibility across multiple binary classification tasks. Use when the user wants to benchmark on adult, COMPAS, Give Me Some Credit, HELOC, Irish, Saheart, Titanic, Wine, or asks about evaluating this task. Reports Success rate.
    3 repo stars
  18. ▌
    Caser Sequential Rec Eval · qhjqhj00
    Evaluates a model's ability to capture sequential user behavior patterns for personalized top-N item recommendation. It probes the model's capacity to model temporal dependencies, skip behaviors, and union-level sequential patterns from historical interactions to predict future items. Use when the user wants to benchmark on MovieLens, Gowalla, Foursquare, Tmall, or asks about evaluating this task. Reports MAP.
    3 repo stars
  19. ▌
    Cbm Concept Accuracy Eval · qhjqhj00
    Evaluates whether Concept Bottleneck Models learn semantically meaningful concept representations from input images under varying annotation granularity and concept correlation structures. Measures how well the model predicts intermediate concepts and downstream tasks compared to standard neural networks. Use when the user wants to benchmark on Playing cards, CheXpert, or asks about evaluating this task. Reports concept accuracy.
    3 repo stars
  20. ▌
    Chexpert Atelectasis Eval · qhjqhj00
    Evaluates the diagnostic accuracy of predictive algorithms and human radiologists on chest X-ray images for detecting atelectasis. It specifically probes whether human expertise provides actionable, non-redundant information on input subsets where algorithms are algorithmically indistinguishable. Use when the user wants to benchmark on CheXpert, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
    3 repo stars
  21. ▌
    Clickbait Mitigation Eval · qhjqhj00
    This evaluation probes a recommender system's ability to mitigate clickbait by measuring performance exclusively on user interactions that result in positive post-click feedback (likes), rather than raw click-through rates. Use when the user wants to benchmark on Unspecified in provided section, or asks about evaluating this task. Reports post-click satisfaction (likes).
    3 repo stars
  22. ▌
    Climate Segmentation Eval · qhjqhj00
    This evaluation probes pixel-level weather pattern segmentation (atmospheric rivers and tropical cyclones) from multi-channel climate data. It measures both segmentation accuracy and exascale training throughput/scaling efficiency across different network architectures and hardware configurations. Use when the user wants to benchmark on Climate weather pattern dataset, or asks about evaluating this task. Reports IoU.
    3 repo stars
  23. ▌
    Clinical Turing Test Eval · qhjqhj00
    Evaluates the physiological realism and clinical fidelity of synthetic 12-lead ECGs by measuring how often expert clinicians can correctly distinguish them from real clinical recordings, and how accurately they can diagnose specific pathologies in both synthetic and real signals. Use when the user wants to benchmark on MedalCare-XL, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  24. ▌
    Closp Crisislandmark Eval · qhjqhj00
    Evaluates cross-modal retrieval and zero-shot classification capabilities for remote sensing imagery (SAR and multispectral optical) paired with text descriptions. It probes how well unified semantic embeddings align heterogeneous geospatial data with natural language for crisis event and land cover analysis. Use when the user wants to benchmark on CrisisLandMark, or asks about evaluating this task. Reports nDCG@1000.
    3 repo stars
  25. ▌
    Cloudops Forecasting Eval · qhjqhj00
    Evaluates time series forecasting models, particularly pre-trained Transformers, on cloud operations data. It probes zero-shot generalization, architectural efficiency, and scaling behavior against classical and deep learning baselines. Use when the user wants to benchmark on azure2017, borg2011, ali2018, or asks about evaluating this task. Reports sMAPE.
    3 repo stars
  26. ▌
    Conformal Prediction Eval · qhjqhj00
    Evaluates the ability of conformal prediction frameworks to produce statistically valid prediction sets with instance-level uncertainty quantification for encoder-only transformers, measuring both classification accuracy and calibration efficiency across standard NLP benchmarks. Use when the user wants to benchmark on GLUE, SuperGLUE, or asks about evaluating this task. Reports Test Accuracy.
    3 repo stars
  27. ▌
    Continual Multimodal Eval · qhjqhj00
    This benchmark evaluates a model's ability to sequentially learn a mix of visual understanding and generation tasks without catastrophically forgetting previously acquired knowledge. It specifically probes intra-modal retention (maintaining performance on earlier tasks) and inter-modal stability (preventing updates for one modality from degrading the other). Use when the user wants to benchmark on ScienceQA, TextVQA, GQA, VizWiz, ImageNet, CustomConcept101, or asks about evaluating this task. Reports Average Accuracy (ACC).
    3 repo stars
  28. ▌
    Counterfactual Chaos Eval · qhjqhj00
    Evaluates the reliability of counterfactual trajectory estimation in chaotic versus non-chaotic dynamical systems under parameter uncertainty and observational noise. It probes whether Bayesian filtering and particle-based smoothing can accurately recover 'what-if' scenarios when small initial perturbations lead to divergent outcomes. Use when the user wants to benchmark on Lorenz System, Rössler System, Logistic Growth, or asks about evaluating this task. Reports RMSE_t.
    3 repo stars
  29. ▌
    Cross Domain Meta Dl Eval · qhjqhj00
    Probes few-shot image classification generalization across diverse domains and highly variable task regimes (2–20 ways, 1–20 shots) without relying on pre-trained backbones. Use when the user wants to benchmark on Meta-Album, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  30. ▌
    Cross Lingual F5 Tts Eval · qhjqhj00
    Evaluates the intelligibility, speaker similarity, and naturalness of synthesized speech in cross-lingual voice cloning and TTS scenarios. It also measures the accuracy of a language-agnostic speaking rate predictor for duration modeling across multiple languages. Use when the user wants to benchmark on Emilia, Seed-TTS-eval, LibriSpeech-PC test-clean, FLEURS, or asks about evaluating this task. Reports WER.
    3 repo stars
  31. ▌
    Csrec Sequential Rec Eval · qhjqhj00
    This evaluation protocol assesses the ranking performance and robustness of sequential recommendation models trained with confident soft labels. It measures whether predicted item sequences align with actual user interactions and verifies if recommendations correspond to genuinely positive user preferences using explicit rating thresholds. Use when the user wants to benchmark on Last.FM, Yelp, Amazon Electronics, Amazon Movies and TV, or asks about evaluating this task. Reports Recall@n, NDCG@n.
    3 repo stars
  32. ▌
    Cultural Positioning Eval · qhjqhj00
    Evaluates whether an LLM's value profile aligns with specific cultural norms using World Values Survey data, and tests the model's steerability when provided with diverse cultural contexts. It probes the extent to which constitutional AI codifies dominant cultural biases and resists prompt-based cultural adaptation. Use when the user wants to benchmark on World Values Survey (WVS) Wave 7, or asks about evaluating this task. Reports Pearson correlation.
    3 repo stars
  33. ▌
    Cvr Ctcvr Estimation Eval · qhjqhj00
    Evaluates the ranking performance of models for click-through rate (CTR) and post-click conversion rate (CVR) estimation in recommendation systems. It probes the model's ability to correctly rank items by their predicted probability of conversion, while mitigating sample selection bias and false independence assumptions between clicks and conversions. Use when the user wants to benchmark on Industrial Benchmark, Ali-CCP, or asks about evaluating this task. Reports AUC.
    3 repo stars
  34. ▌
    D2sac Asp Scheduling Eval · qhjqhj00
    Evaluates a diffusion-based reinforcement learning agent's capability to dynamically assign AI-generated content tasks to edge service providers under stochastic workloads, optimizing for user utility while preventing system crashes and minimizing training time. Use when the user wants to benchmark on Custom AIGC Edge Simulation Environment, Gym Benchmark Tasks, or asks about evaluating this task. Reports Cumulative Reward.
    3 repo stars
  35. ▌
    Dctracks Track Recon Eval · qhjqhj00
    Evaluates the performance of machine learning and traditional algorithms for reconstructing particle tracks in drift chamber detectors. It probes hit-level matching accuracy, track-level reconstruction efficiency, charge identification correctness, and momentum resolution under realistic detector conditions. Use when the user wants to benchmark on DCTracks, or asks about evaluating this task. Reports track efficiency.
    3 repo stars
  36. ▌
    Deepfake Speech Auth Eval · qhjqhj00
    Evaluates the robustness of audio-based biometric authentication systems against deepfake speech synthesis attacks. It measures how easily voice cloning models can bypass speaker verification and how effectively anti-spoofing detectors can distinguish genuine from synthetic speech. Use when the user wants to benchmark on AISHELL-3, or asks about evaluating this task. Reports Bypass Rate.
    3 repo stars
  37. ▌
    Deepurban Trajectory Eval · qhjqhj00
    Evaluates trajectory prediction and planning capabilities in high-density urban environments with significant vehicle-to-vulnerable-road-user interactions. It measures prediction accuracy and safety compliance using displacement errors and collision scores. Use when the user wants to benchmark on DeepUrban, or asks about evaluating this task. Reports ADE.
    3 repo stars
  38. ▌
    Sm3 Text To Query Eval · qhjqhj00
    Evaluates text-to-query systems across relational, document, and graph database models using four query languages (SQL, MQL, Cypher, SPARQL). It probes the ability of models to translate natural language medical questions into correct, executable database queries using standardized SNOMED-CT aligned synthetic patient data. Use when the user wants to benchmark on SM3-Text-to-Query, or asks about evaluating this task. Reports correctness.
    3 repo stars
  39. ▌
    Social Media Bias Eval · qhjqhj00
    This benchmark evaluates the ability of models to automatically detect multiple dimensions of media bias (e.g., hate speech, racial, gender, political, linguistic, and text-level context bias) in social media posts across different topic domains. It probes a model's robustness to domain shift and severe class imbalance in multi-label bias identification tasks. Use when the user wants to benchmark on Social Media Bias Dataset (YouTube & Reddit), or asks about evaluating this task. Reports weighted average F1 score.
    3 repo stars
  40. ▌
    Soft Pairwise Accuracy · qhjqhj00
    Evaluates the reliability and discriminative power of automatic machine translation metrics by comparing their statistical significance against human MQM judgments. It measures how well a metric's pairwise system rankings align with human preferences using permutation-based p-values rather than hard binary decisions. Use when the user has predictions and gold and needs to compute Soft Pairwise Accuracy (SPA).
    3 repo stars
  41. ▌
    Sparrow Alignment Eval · qhjqhj00
    Evaluates the alignment, factual grounding, and rule-following capabilities of dialogue agents through human preference comparisons. It also measures resilience to adversarial probing for specific harm rules and the quality of evidence-supported responses. Use when the user wants to benchmark on ELI5 + Free Dialogue Test Set, or asks about evaluating this task. Reports Three-model preference rate.
    3 repo stars
  42. ▌
    Sparse Gpu Kernel Eval · qhjqhj00
    Evaluates the performance and efficiency of custom sparse GPU kernels for SpMM and SDDMM operations against standard libraries like cuSPARSE on deep learning workloads. It measures computational throughput, memory usage, and end-to-end speedups across various model architectures and batch sizes. Use when the user wants to benchmark on Sparse Matrix Dataset from DNNs, or asks about evaluating this task. Reports Geometric mean speedup.
    3 repo stars
  43. ▌
    Spatial Reasoning Eval · qhjqhj00
    Evaluates vision-language models' ability to count objects and reason about spatial relationships (depth, distance, relative position) in images. It probes segmentation capabilities, attention alignment, and robustness to linguistic variations (out-of-distribution shifts). Use when the user wants to benchmark on CLEVR_CoGenT_ValB, CVBench, Pixmo-Count, Static Spatial Reasoning (SAT), VSR, VC Bench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  44. ▌
    Spatialdistortionindex · qhjqhj00
    Compute the SpatialDistortionIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpatialDistortionIndex, or asks how to score with SpatialDistortionIndex.
    3 repo stars
  45. ▌
    Speech Likelihood Eval · qhjqhj00
    Evaluates the ability of generative latent variable models and autoregressive baselines to model speech audio distributions at varying temporal resolutions. It measures how well models capture intra-frame and inter-frame correlations in audio waveforms by optimizing likelihood objectives. Use when the user wants to benchmark on TIMIT, LibriSpeech, or asks about evaluating this task. Reports bits per frame (bpf).
    3 repo stars
  46. ▌
    Speech Separation Eval · qhjqhj00
    Evaluates speech separation models' ability to isolate individual speaker signals from multi-speaker mixtures under various acoustic conditions, including moving sources, environmental noise, and musical noise. It measures both objective signal quality and subjective perceptual metrics to assess generalization from synthetic to real-world dynamic scenarios. Use when the user wants to benchmark on SonicSet, RealSEP, HumanSEP, LRS2-2Mix, Libri2Mix, or asks about evaluating this task. Reports SI-SNR.
    3 repo stars
  47. ▌
    Speechmentalmanip Eval · qhjqhj00
    Binary classification of spoken multi-speaker dialogues to detect the presence of mental manipulation tactics. It probes an audio-language model's ability to identify subtle manipulative cues in synthetic speech without relying on text transcripts. Use when the user wants to benchmark on SpeechMentalManip, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  48. ▌
    Stackoverflow Ner Eval · qhjqhj00
    Evaluates named entity recognition capabilities on software programming texts. It specifically probes the model's ability to identify fine-grained code-related entities like variable names, libraries, and data structures in StackOverflow posts. Use when the user wants to benchmark on StackOverflow NER corpus, or asks about evaluating this task. Reports F1.
    3 repo stars
  49. ▌
    Stella Living Lab Eval · qhjqhj00
    Evaluates academic search and recommendation systems in live production environments using A/B testing and user interaction logs, bridging the gap between offline test collections and real-world performance. Use when the user wants to benchmark on LIVIVO, GESIS Search, or asks about evaluating this task. Reports click-paths.
    3 repo stars
  50. ▌
    Stochastic Ackley Eval · qhjqhj00
    Evaluates the ability of uncertainty-aware deep neural networks to approximate a highly irregular, multi-extremum function and quantify predictive uncertainty. It specifically probes how well the models handle in-distribution versus out-of-distribution parameter regimes. Use when the user wants to benchmark on Stochastic Ackley Function, or asks about evaluating this task. Reports Relative Error (RE).
    3 repo stars
  51. ▌
    Success Rate 3d Policy · qhjqhj00
    Evaluates the robustness and generalization of 3D policy learning models for robotic manipulation across varying environmental conditions, temporal horizons, and real-world interference. It probes spatial understanding, fine-grained pose control, and resilience to domain randomization and lighting changes. Use when the user wants to benchmark on RoboTwin 2.0, ManiSkill2, Real-World Manipulation, or asks about evaluating this task. Reports Success Rate (%).
    3 repo stars
  52. ▌
    Superb Downstream Eval · qhjqhj00
    Evaluates pre-trained speech models on downstream spoken language understanding tasks. It probes the model's ability to classify spoken intents, fill semantic slots in transcriptions, and detect specific keywords in audio. Use when the user wants to benchmark on SUPERB, or asks about evaluating this task. Reports test accuracy.
    3 repo stars
  53. ▌
    Svld Points Ratio Eval · qhjqhj00
    Probes a model's ability to predict social engagement (upvote ratio) from multimodal inputs (images, videos, and text). It evaluates cross-modal fusion and regression capabilities on socially grounded, context-rich data. Use when the user wants to benchmark on SVLD, or asks about evaluating this task. Reports Mean L1point ratio prediction error.
    3 repo stars
  54. ▌
    Synthetic Geology Eval · qhjqhj00
    Evaluates a flow matching generative model's ability to produce realistic 3D subsurface geological models, both unconditionally and conditioned on sparse borehole data. It probes the model's capacity for geological interpolation, structural feature reconstruction, and probabilistic uncertainty estimation. Use when the user wants to benchmark on Synthetic Geology / StructuralGeo Dataset, or asks about evaluating this task. Reports probabilistic confidence intervals.
    3 repo stars
  55. ▌
    Synthetic4relight Eval · qhjqhj00
    Evaluates a 3D scene representation's capability for novel view synthesis, relighting, and inverse rendering (estimating diffuse albedo and roughness) from posed RGB images. Use when the user wants to benchmark on Synthetic4Relight, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  56. ▌
    Tabular Benchmark Eval · qhjqhj00
    Evaluates the predictive performance and stability of 32 deep learning and tree-based tabular models across a large collection of diverse tabular datasets. It probes how well different architectures handle classification and regression tasks, and how dataset characteristics influence method rankings. Use when the user wants to benchmark on LAMDA-TALENT Benchmark, or asks about evaluating this task. Reports average_rank.
    3 repo stars
  57. ▌
    Tally Gpu Sharing Eval · qhjqhj00
    Evaluates the performance isolation and resource sharing capabilities of GPU scheduling systems for concurrent deep learning workloads. It probes how well a system maintains tail latency for high-priority inference tasks while maximizing throughput for best-effort training tasks under varying traffic loads and workload combinations. Use when the user wants to benchmark on Tally Benchmark Suite, or asks about evaluating this task. Reports 99th-percentile latency.
    3 repo stars
  58. ▌
    Tamper Resistance Eval · qhjqhj00
    This benchmark evaluates the robustness of LLM safety safeguards against fine-tuning-based tampering attacks. It measures whether a model can maintain low accuracy on weaponized knowledge (forget) and high benign capabilities (retain) after undergoing various supervised fine-tuning attacks. Use when the user wants to benchmark on WMDP, MMLU, HarmBench, MT-Bench, or asks about evaluating this task. Reports Post-Attack Forget accuracy, Attack Success Rate (ASR).
    3 repo stars
  59. ▌
    Tb Classification Eval · qhjqhj00
    Evaluates deep learning models' ability to classify chest X-rays as tuberculosis or normal, comparing whole-image vs. lung-segmented inputs. Use when the user wants to benchmark on Kaggle CXR images and lung mask dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  60. ▌
    Theory Of Mind QA Eval · qhjqhj00
    Evaluates a model's ability to track first-order and second-order false beliefs, distinguishing an agent's mental state from physical reality and memory. It probes whether systems can maintain consistent world-state representations when agents hold incorrect beliefs about object locations or events. Use when the user wants to benchmark on Sally-Anne & Icecream Van Tasks, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Thinkjepa Ego Dex Eval · qhjqhj00
    Evaluates a model's ability to forecast future 3D hand/joint trajectories and latent video representations from egocentric video inputs. It probes long-horizon temporal consistency and physical plausibility in dexterous manipulation scenarios. Use when the user wants to benchmark on EgoDex, EgoExo4D, or asks about evaluating this task. Reports ADE.
    3 repo stars
  62. ▌
    Thyme Scene Graph Eval · qhjqhj00
    Evaluates a model's ability to generate dynamic video scene graphs by predicting inter-object relationships and attributes across multiple temporal frames. It specifically probes temporal consistency, handling of occlusions, and modeling of long-range dependencies in both ground-level and aerial video footage. Use when the user wants to benchmark on ASPIRe, AeroEye-v1.0, or asks about evaluating this task. Reports R@20, mR@20.
    3 repo stars
  63. ▌
    Timer Time Series Eval · qhjqhj00
    Evaluates a large decoder-only Transformer model on standard time series benchmarks for forecasting, imputation, and anomaly detection. It probes the model's few-shot generalization, scalability, and robustness in data-scarce scenarios compared to encoder-only baselines. Use when the user wants to benchmark on ETT, ECL, Traffic, Weather, PEMS, UCR Anomaly Archive, or asks about evaluating this task. Reports MSE.
    3 repo stars
  64. ▌
    Topofair Fairness Eval · qhjqhj00
    Evaluates fairness-aware link prediction models on synthetic graphs with controlled topological biases. It probes how structural properties like assortativity, heterogeneity, and class imbalance impact fairness metrics (SP, EO) and predictive accuracy (Hit@10, AUC). Use when the user wants to benchmark on Opinion use case, Friendship use case, Collab use case, Real datasets (Collab, Polblogs, Facebook), or asks about evaluating this task. Reports Statistical Parity (SP), Equalized Odds (EO).
    3 repo stars
  65. ▌
    Toxicity Analysis Eval · qhjqhj00
    Measures the toxicity of text generated by language models conditioned on specific prompts, evaluating how alignment techniques like prompting and context distillation affect harmful content generation. Use when the user wants to benchmark on RealToxicityPrompts, or asks about evaluating this task. Reports mean toxicity score.
    3 repo stars
  66. ▌
    Tracknet Tracking Eval · qhjqhj00
    Evaluates the ability of deep learning models to detect and track high-speed, tiny objects (tennis and badminton balls) in broadcast sports videos. It probes robustness to motion blur, occlusion, and domain shifts by comparing single-frame vs. multi-frame tracking and transfer learning across different sports. Use when the user wants to benchmark on Tennis, Badminton, or asks about evaluating this task. Reports F1-measure.
    3 repo stars
  67. ▌
    Train O Matic Wsd Eval · qhjqhj00
    Evaluates the quality of automatically generated multilingual word sense disambiguation (WSD) training corpora by training a supervised WSD system (IMS) on them and measuring performance on standard WSD benchmark datasets. It probes whether synthetic sense-annotated data can match or exceed manually annotated corpora, particularly for low-resource languages. Use when the user wants to benchmark on Senseval-2, Senseval-3, SemEval-2007, SemEval-2013, SemEval-2015, or asks about evaluating this task. Reports F1.
    3 repo stars
  68. ▌
    Trilemma Of Truth Eval · qhjqhj00
    Evaluates large language models' ability to distinguish factually true statements from factually false and unverifiable ('neither') statements. It probes both prompt-based output probabilities and internal hidden activations to measure veracity classification accuracy and uncertainty quantification. Use when the user wants to benchmark on Trilemma of Truth Datasets, or asks about evaluating this task. Reports MCC.
    3 repo stars
  69. ▌
    Tucano2 Portfolio Eval · qhjqhj00
    This evaluation protocol probes the language understanding, reasoning, and instruction-following capabilities of Portuguese LLMs across diverse domains including academic exams, natural language inference, physical commonsense, and code generation. It is specifically designed to provide reliable training signals during pretraining and assess post-training alignment. Use when the user wants to benchmark on ARC Challenge, Calame, Global PIQA, HellaSwag, LAMBADA, ENEM, BLUEX, OAB, Belebele, MMLU, IFEval-PT, GSM8K-PT, RULER-PT, HumanEval, or asks about evaluating this task. Reports accuracy (log-likelihood selection).
    3 repo stars
  70. ▌
    United Medasr Asr Eval · qhjqhj00
    This evaluation measures the transcription accuracy of a fine-tuned automatic speech recognition model across four diverse speech benchmarks. It specifically probes the model's robustness to different speaking styles, accents, and linguistic contexts after applying a noise reduction step and a BART-based semantic correction pipeline. Use when the user wants to benchmark on LibriSpeech, Europarl-ASR, TED-LIUM, FLEURS, or asks about evaluating this task. Reports Word Error Rate (WER).
    3 repo stars
  71. ▌
    Unlearning Recsys Eval · qhjqhj00
    Evaluates the ability of recommender systems to efficiently remove specific user interactions or sensitive items (unlearning) while preserving recommendation utility. It probes real-world operational constraints, including handling sequential small-batch deletion requests, domain-specific triggers, and low-latency execution across collaborative filtering, session-based, and next-basket recommendation tasks. Use when the user wants to benchmark on TaFeng, Dunnhumby, Instacart, RSC15, DIGI, NOWP, Goodreads, MovieLens, Amazon Reviews, or asks about evaluating this task. Reports Recall, PHR, nDCG.
    3 repo stars
  72. ▌
    Urban Pathfinding Eval · qhjqhj00
    Evaluates real-time urban pathfinding algorithms under dynamic traffic and weather conditions. It measures how well traditional graph search methods and deep learning models predict optimal routes and minimize travel time in a simulated Berlin city environment. Use when the user wants to benchmark on Berlin Urban Simulation, or asks about evaluating this task. Reports Average Travel Time (s).
    3 repo stars
  73. ▌
    Usb Summarization Eval · qhjqhj00
    Evaluates multiple text summarization capabilities including extractive/abstractive generation, factuality verification, factual error correction, topic-constrained generation, sentence compression, evidence extraction, and unsupported span detection across diverse domains. Use when the user wants to benchmark on Extractive Summarization (EXT), Abstractive Summarization (ABS), Factuality Classification (FAC), Fixing Factuality (FIX), Topic-based Summarization (TOPIC), Multi-sentence Compression (COMP), Evidence Extraction (EVEXT), Unsupported Span Prediction (UNSUP), or asks about evaluating this task. Reports ROUGE.
    3 repo stars
  74. ▌
    Video Outpainting Eval · qhjqhj00
    Evaluates a model's ability to generate spatially and temporally consistent video content outside the original frame boundaries (video outpainting), while preserving source structure and visual realism. Use when the user wants to benchmark on DAVIS 2017, YouTube-VOS, or asks about evaluating this task. Reports FVD.
    3 repo stars
  75. ▌
    Videogameqa Bench Eval · qhjqhj00
    Evaluates vision-language models on video game quality assurance tasks, including glitch detection, temporal reasoning, and bug reporting. It probes the model's ability to process sampled video frames, identify visual anomalies, and generate structured or descriptive reports about game glitches. Use when the user wants to benchmark on VideoGameQA-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  76. ▌
    Viona Fuzzy Reordering · qhjqhj00
    Compute Viona/fuzzy_reordering via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Viona/fuzzy_reordering.
    3 repo stars
  77. ▌
    Visual Genome Sgg Eval · qhjqhj00
    Evaluates fine-grained scene graph generation by predicting subject-predicate-object triplets from images. It measures recall and F1 scores across head, body, and tail predicate classes to assess performance on long-tailed distributions and missing annotations. Use when the user wants to benchmark on Visual Genome, or asks about evaluating this task. Reports mR@K, F@K.
    3 repo stars
  78. ▌
    Visual Sycophancy Eval · qhjqhj00
    This evaluation probes how Vision-Language Models ground their responses in visual input versus relying on language priors or user bias. It measures perceptual awareness, visual dependency, and alignment conflicts by comparing model behavior across original, blank, noisy, and semantically conflicting images. Use when the user wants to benchmark on GQA, VQAv2, A-OKVQA, POPE, or asks about evaluating this task. Reports VNS.
    3 repo stars
  79. ▌
    Vln Task Planning Eval · qhjqhj00
    Evaluates an agent's ability to decompose coarse-grained natural language navigation instructions into executable subtasks and navigate through simulated environments to reach target locations or interact with objects. It probes task planning, visual-language grounding, and dynamic error recovery in continuous or discrete navigation spaces. Use when the user wants to benchmark on R2R, REVERIE, ALFRED, or asks about evaluating this task. Reports Success Rate (SR).
    3 repo stars
  80. ▌
    Vos Language Referring · qhjqhj00
    Evaluates a model's ability to perform pixel-level video object segmentation guided by natural language referring expressions, testing both language grounding and temporal consistency in dynamic scenes. Use when the user wants to benchmark on DAVIS-16, DAVIS-17, or asks about evaluating this task. Reports performance score.
    3 repo stars
  81. ▌
    Vqa Cot Reasoning Eval · qhjqhj00
    Probes the ability of vision-language models to perform multi-step chain-of-thought reasoning on visual inputs across diverse domains like charts, documents, science diagrams, and math. It measures both direct answer accuracy and structured reasoning accuracy. Use when the user wants to benchmark on A-OKVQA, ChartQA, DocVQA, InfoVQA, TextVQA, AI2D, ScienceQA, MathVista, OCRBench, MMStar, MMMU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  82. ▌
    Whisper Zero Shot Eval · qhjqhj00
    Evaluates the zero-shot generalization capability of a speech recognition model across diverse English and multilingual domains. It measures robustness to out-of-distribution audio, varying noise levels, and translation tasks without any dataset-specific fine-tuning. Use when the user wants to benchmark on LibriSpeech, Common Voice, Fleurs, CoVoST2, Multilingual LibriSpeech (MLS), VoxPopuli, or asks about evaluating this task. Reports WER.
    3 repo stars
  83. ▌
    Widget Captioning Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate natural language descriptions for individual mobile UI elements using multimodal inputs. It probes the capability to fuse visual appearance and structural hierarchy data to produce accurate, context-aware captions for accessibility and UI understanding tasks. Use when the user wants to benchmark on Widget Captioning Dataset, or asks about evaluating this task. Reports CIDEr.
    3 repo stars
  84. ▌
    Window Wise Complexity · qhjqhj00
    Quantifies the intrinsic predictability of time series data by measuring window-wise pattern complexity in the frequency domain. It establishes a data-driven performance lower bound for forecasting models and identifies whether standard benchmarks have reached saturation. Use when the user has predictions and gold and needs to compute window-wise complexity.
    3 repo stars
  85. ▌
    Zebrafish Sctrans Eval · qhjqhj00
    Evaluates the ability of topological data analysis methods to capture developmental transitions and cell lineage dynamics in single-cell RNA sequencing time-series data. It specifically tests whether higher-order simplicial complexity can outperform conventional topological invariants like Betti numbers in identifying critical biological stages. Use when the user wants to benchmark on Farrell et al. (2018) zebrafish scRNA-seq, or asks about evaluating this task. Reports normalized simplicial complexity.
    3 repo stars
  86. ▌
    3d 2d Vl Grounding Eval · qhjqhj00
    Evaluates a unified vision-language model's ability to ground natural language instructions to 3D objects and 2D regions, as well as answer 3D visual questions. It probes spatial reasoning, cross-modal alignment, and robustness to different 3D input representations (mesh-sampled vs. sensor RGB-D point clouds). Use when the user wants to benchmark on SR3D, NR3D, ScanRefer, RefCOCO, RefCOCO+, RefCOCOg, ScanQA, SQA3D, or asks about evaluating this task. Reports top-1 accuracy (Acc@25/50/75).
    3 repo stars
  87. ▌
    3d Shape Retrieval Eval · qhjqhj00
    Evaluates algorithms for retrieving geometrically and topologically similar 3D shapes from a database given a query shape. Probes pose invariance, shape descriptor robustness, and retrieval ranking accuracy. Use when the user wants to benchmark on NIST shape benchmark, or asks about evaluating this task. Reports precision-recall.
    3 repo stars
  88. ▌
    Ade20k Scene Parse Eval · qhjqhj00
    Evaluates a model's ability to perform dense pixel-wise semantic segmentation across 150 common scene categories, including both discrete objects and amorphous 'stuff' classes. It probes fine-grained scene understanding and the model's capacity to handle class imbalance and varying object scales. Use when the user wants to benchmark on SceneParse150, or asks about evaluating this task. Reports Mean IoU.
    3 repo stars
  89. ▌
    Adjustedmutualinfoscore · qhjqhj00
    Compute the AdjustedMutualInfoScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute AdjustedMutualInfoScore, or asks how to score with AdjustedMutualInfoScore.
    3 repo stars
  90. ▌
    Amp Classification Eval · qhjqhj00
    Evaluates the ability of reprogrammed language models to classify antimicrobial peptide (AMP) sequences into binary categories (toxic vs. non-toxic, or AMP vs. non-AMP) using limited labeled data. Use when the user wants to benchmark on AMP Dataset, or asks about evaluating this task. Reports Test Accuracy.
    3 repo stars
  91. ▌
    Amp Motion Control Eval · qhjqhj00
    Evaluates a physics-based character's ability to learn stylized locomotion and complex task execution (e.g., navigating targets, avoiding obstacles) by imitating unstructured motion datasets. It probes the model's capacity to compose disparate skills, generalize across gaits, and maintain high-fidelity motion tracking without manual motion planning. Use when the user wants to benchmark on AMP Motion Datasets, or asks about evaluating this task. Reports normalized task return.
    3 repo stars
  92. ▌
    Antibody Domainbed Eval · qhjqhj00
    Evaluates out-of-distribution generalization of protein language models and sequence CNNs for therapeutic antibody design across different antigen targets and generative models. It probes robustness to covariate shifts, label shifts, and assay biases in molecular sequence data. Use when the user wants to benchmark on Antibody DomainBed, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  93. ▌
    Average Precision Score · qhjqhj00
    Compute the average_precision_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute average_precision_score, or asks how to score with average_precision_score.
    3 repo stars
  94. ▌
    Balanced Accuracy Score · qhjqhj00
    Compute the balanced_accuracy_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute balanced_accuracy_score, or asks how to score with balanced_accuracy_score.
    3 repo stars
  95. ▌
    Bayes Factor Odds Ratio · qhjqhj00
    Evaluates the likelihood of different binary black hole formation channels (CEE, CHE, SMT) given gravitational wave strain data by comparing Bayesian evidence and prior odds. Use when the user has predictions and gold and needs to compute Bayes factor ($\mathcal{B}$), Odds ratio ($\mathcal{O}$).
    3 repo stars
  96. ▌
    Benchmark Accuracy Eval · qhjqhj00
    Evaluates zero-shot language model performance across a suite of 10 standard NLP benchmarks covering commonsense reasoning, science QA, and language modeling. It measures task accuracy and correlates it with word-level statistical overlap metrics to assess distributional alignment between pre-training data and evaluation sets. Use when the user wants to benchmark on ARC Easy, ARC Challenge, Hellaswag, MMLU, SciQ, OpenBookQA, PIQA, lambada, SocialIQA, SWAG, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  97. ▌
    Biosage Scientific Eval · qhjqhj00
    Evaluates a compound AI architecture's ability to retrieve, synthesize, and reason across cross-disciplinary scientific knowledge. It probes performance on established single-domain science benchmarks and a novel benchmark specifically designed for bio-AI cross-domain synthesis and reasoning. Use when the user wants to benchmark on LitQA2, GPQA, WMDP, HLE-Bio, BioSage Cross-Disciplinary Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  98. ▌
    Bowdbeg Matching Series · qhjqhj00
    Compute bowdbeg/matching_series via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of bowdbeg/matching_series.
    3 repo stars
  99. ▌
    Brats Segmentation Eval · qhjqhj00
    Evaluates the precision of brain tumor segmentation models on multi-modal MRI scans across three clinically relevant regions (Whole Tumor, Tumor Core, Enhancing Tumor). It probes the model's ability to accurately delineate heterogeneous tumor boundaries and correctly identify positive tumor voxels in medical imaging data. Use when the user wants to benchmark on BraTS2019/2020, or asks about evaluating this task. Reports Dice coefficient.
    3 repo stars
  100. ▌
    Brisc Segmentation Eval · qhjqhj00
    Evaluates deep learning models on multi-planar brain tumor segmentation from contrast-enhanced T1-weighted MRI scans. It probes multi-scale feature integration, cross-view generalization, and robustness to class imbalance across glioma, meningioma, and pituitary tumor types. Use when the user wants to benchmark on BRISC, or asks about evaluating this task. Reports mIoU.
    3 repo stars