all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 30 of 76

  1. ▌
    Binarycalibrationerror · qhjqhj00
    Compute the BinaryCalibrationError metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryCalibrationError, or asks how to score with BinaryCalibrationError.
    3 repo stars
  2. ▌
    Binarymatthewscorrcoef · qhjqhj00
    Compute the BinaryMatthewsCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryMatthewsCorrCoef, or asks how to score with BinaryMatthewsCorrCoef.
    3 repo stars
  3. ▌
    Biological Visual Eval · qhjqhj00
    Evaluates frozen visual embedding extractors on ecological trait alignment, fine-grained intra-species variation preservation, and zero-shot/few-shot transfer learning across diverse biological domains. Use when the user wants to benchmark on FishNet, NeWT, AwA2, Herb, PlantDoc, Life stage-Diff/Align, Sex-Diff/Align, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  4. ▌
    Biomedical Cypher Eval · qhjqhj00
    Probes an LLM's ability to generate syntactically and semantically correct Cypher queries for a biomedical knowledge graph, and execute them to answer domain-specific questions without hallucination. Use when the user wants to benchmark on Custom Biomedical QA Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  5. ▌
    Blast Forecasting Eval · qhjqhj00
    Evaluates the zero-shot forecasting capability of universal time series models across multiple domains and prediction horizons. It probes how well pre-trained models generalize to unseen datasets and measures predictive accuracy using standard error metrics. Use when the user wants to benchmark on ETTh1, ETTh2, ETTm1, ETTm2, Weather, GlobalTemp, GIFT-Eval, or asks about evaluating this task. Reports MSE, MAE.
    3 repo stars
  6. ▌
    Bnn Nvm Benchmark Eval · qhjqhj00
    This benchmark evaluates the inference accuracy and hardware performance of binary neural networks (BNNs) deployed on non-volatile memory crossbar architectures. It probes how hardware constraints like ADC resolution and first-layer input precision affect model accuracy, latency, energy efficiency, and chip area. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    Breast Mri Lr Seg Eval · qhjqhj00
    Evaluates the ability of 3D medical image segmentation models to accurately delineate left and right breast tissues in MRI scans. It probes anatomical partitioning robustness and generalization across diverse clinical sources. Use when the user wants to benchmark on Breast MRI Left-Right Segmentation Dataset, or asks about evaluating this task. Reports Dice Similarity Coefficient (DSC).
    3 repo stars
  8. ▌
    Browser Inference Eval · qhjqhj00
    Evaluates the performance overhead and latency characteristics of running deep learning inference directly in web browsers compared to native environments. It probes the impact of WebAssembly runtime inefficiencies, SIMD limitations, and WebGL GPU abstraction on prediction, warmup, and setup phases across various models and hardware configurations. Use when the user wants to benchmark on Inference Benchmark (ResNet50, VGG16, MobileNetV2), or asks about evaluating this task. Reports prediction_latency.
    3 repo stars
  9. ▌
    Calicalcausalrank Eval · qhjqhj00
    Evaluates multi-objective ad ranking models on click-through rate (CTR) and conversion rate (CVR) prediction tasks, focusing on ranking quality, score calibration, and counterfactual utility estimation under position and selection bias. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.
    3 repo stars
  10. ▌
    Canmt Contrastive Eval · qhjqhj00
    Evaluates context-aware neural machine translation models on their ability to correctly translate context-dependent discourse phenomena (e.g., anaphoric pronouns, deixis, ellipsis) across sentence boundaries in concatenated input windows. Use when the user wants to benchmark on En→Ru movie subtitles (Voita et al., 2019), En→De TED talk subtitles (IWSLT17), Voita contrastive set, ContraPro, or asks about evaluating this task. Reports Contrastive accuracy.
    3 repo stars
  11. ▌
    Carla Leaderboard Eval · qhjqhj00
    Evaluates autonomous driving agents on their ability to navigate complex urban environments while balancing route completion, safety, and rule compliance. It probes how well models handle dynamic traffic interactions, edge-case scenarios, and physical constraints in a closed-loop simulation. Use when the user wants to benchmark on CARLA Leaderboard 2.0 scenarios, CARLA 42 Routes, Town05, or asks about evaluating this task. Reports Driving Score (DS).
    3 repo stars
  12. ▌
    Cfbenchmark Basic Eval · qhjqhj00
    Evaluates Chinese large language models on financial text processing capabilities, specifically entity recognition, text classification, and content generation within the financial domain. It tests the models' adaptability using zero-shot and few-shot (3 examples) prompting strategies across eight distinct tasks. Use when the user wants to benchmark on CFBenchmark-Basic, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  13. ▌
    Chaos Forecasting Eval · qhjqhj00
    Evaluates the ability of time series forecasting models to predict trajectories of low-dimensional chaotic dynamical systems. It probes how well models capture underlying deterministic chaos, smoothness, and multi-scale temporal dependencies without explicit trend or seasonality signals. Use when the user wants to benchmark on Chaotic Dynamical Systems Benchmark, or asks about evaluating this task. Reports sMAPE.
    3 repo stars
  14. ▌
    Chatgpt Multitask Eval · qhjqhj00
    Evaluates ChatGPT's zero-shot multitask capabilities across summarization, machine translation, sentiment analysis, question answering, dialogue, and misinformation detection. It probes the model's generalization, reasoning, multilingual understanding, and task-specific performance without fine-tuning. Use when the user wants to benchmark on CNN/DM, SAMSum, FLoRes-200, NusaX, bAbI, EntailmentBank, CLUTRR, StepGame, Pep-3k, COVID-Social, COVID-Scientific, MultiWOZ2.2, OpenDialKG, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  15. ▌
    Chime2 Robust Asr Eval · qhjqhj00
    Evaluates the effectiveness of a time-domain speech enhancement frontend in improving automatic speech recognition performance on noisy and reverberant speech. It probes the model's ability to enhance speech without introducing distortion that degrades downstream ASR accuracy. Use when the user wants to benchmark on CHiME-2, or asks about evaluating this task. Reports WER.
    3 repo stars
  16. ▌
    Chinese LLM Bench Eval · qhjqhj00
    Evaluates Chinese large language models' world knowledge, academic understanding, and multi-dimensional alignment after pretraining or instruction fine-tuning. It probes the model's ability to follow instructions, reason across domains, and maintain safety and helpfulness standards in Chinese. Use when the user wants to benchmark on C-Eval, CMMLU, Alignbench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    Cimi4d Annotation Eval · qhjqhj00
    Probes the accuracy of automatically generated 3D human pose and translation annotations for rock climbing motions. It evaluates how well a LiDAR-IMU fusion and blending optimization pipeline reconstructs off-ground climbing poses compared to manual ground truth. Use when the user wants to benchmark on CIMI4D, or asks about evaluating this task. Reports PMPJPE.
    3 repo stars
  18. ▌
    Clip Cxr Fairness Eval · qhjqhj00
    Evaluates zero-shot classification performance of CLIP-based vision-language models on chest X-rays, assessing fairness across demographic subgroups (age, sex, race) and robustness to spurious correlations (presence of chest drains in pneumothorax cases). Use when the user wants to benchmark on MIMIC-CXR, or asks about evaluating this task. Reports AUPRCadj.
    3 repo stars
  19. ▌
    Cnn Music Tagging Eval · qhjqhj00
    This evaluation benchmarks CNN-based models on automatic music tagging, measuring their ability to predict multiple genre, instrument, and mood labels from audio spectrograms. It assesses both standard classification performance and robustness to audio transformations like time-stretching and pitch shifting. Use when the user wants to benchmark on MagnaTagATune, Million Song Dataset, MTG-Jamendo, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  20. ▌
    Codeparrot Apps Metric · qhjqhj00
    Compute codeparrot/apps_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of codeparrot/apps_metric.
    3 repo stars
  21. ▌
    Colour Mnist Bias Eval · qhjqhj00
    Evaluates how well a classifier maintains performance on a biased dataset when trained on different coreset selection strategies. It probes the model's robustness to dataset bias and measures the effectiveness of data frugality techniques in mitigating bias while preserving accuracy across varying data budgets. Use when the user wants to benchmark on Colour-MNIST, or asks about evaluating this task. Reports classifier performance.
    3 repo stars
  22. ▌
    Comet Thermal Sim Eval · qhjqhj00
    Evaluates the accuracy and overhead of an integrated thermal simulation toolchain (CoMeT) for modeling processor-memory thermal dynamics across 2D, 2.5D, and 3D architectures. It probes the tool's ability to capture thermal coupling, leakage power effects, and DVFS/DTM interactions under diverse compute and memory-intensive workloads. Use when the user wants to benchmark on PARSEC 2.1, SPLASH-2, SPEC CPU2017, or asks about evaluating this task. Reports Temperature.
    3 repo stars
  23. ▌
    Compositional Arc Eval · qhjqhj00
    This benchmark evaluates systematic generalization in abstract spatial reasoning by testing whether models can infer and compose geometric transformations (e.g., translation, rotation, reflection) from limited few-shot examples. It specifically probes out-of-distribution compositionality by training on known transformation primitives and level-1 compositions, then testing on novel level-2 compositions. Use when the user wants to benchmark on Compositional-ARC, or asks about evaluating this task. Reports exact match accuracy.
    3 repo stars
  24. ▌
    Courtguard Safety Eval · qhjqhj00
    Evaluates LLM safety guardrails and policy-adaptation frameworks on their ability to correctly identify harmful, toxic, or policy-violating content across diverse attack vectors. It probes robustness against automated jailbreaks, over-refusal in benign contexts, and zero-shot adaptability to out-of-domain policy enforcement. Use when the user wants to benchmark on AdvBenchM, WildGuard, HarmBench, JailJudge, PKU-SafeRLHF, ToxicChat, BeaverTails, XSTest, PAN Wikipedia Vandalism Corpus 2010, Human-Verified Attack Suite Dataset, or asks about evaluating this task. Reports Accuracy, F1-Score (macro-averaged).
    3 repo stars
  25. ▌
    Covid Drug Design Eval · qhjqhj00
    This benchmark evaluates deep graph generative models (JT-VAE and DQN) for their ability to design novel molecular structures optimized for high predicted potency against the SARS-CoV-2 3CL-protease, while balancing drug-likeness, lipophilicity, and synthesizability. It also assesses structural novelty relative to known antivirals and predicted binding affinity using in silico classifiers. Use when the user wants to benchmark on ChEMBL/BindingDB/ToxCat pharmacology dataset, or asks about evaluating this task. Reports pIC50.
    3 repo stars
  26. ▌
    Ct Rate Zero Shot Eval · qhjqhj00
    This benchmark evaluates the zero-shot multi-abnormality detection capability of a visual-language foundation model on 3D chest CT volumes. It probes the model's ability to generalize to unseen data distributions and classify multiple pathologies simultaneously without task-specific supervised training. Use when the user wants to benchmark on CT-RATE, RAD-ChestCT, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  27. ▌
    Cultural Aware Mt Eval · qhjqhj00
    Evaluates machine translation systems on culturally specific items (CSIs) to measure how well they preserve cultural nuances, entities, and meanings compared to reference translations. It probes both automated lexical/semantic alignment and human judgment on translation accuracy for culturally grounded content. Use when the user wants to benchmark on Wikipedia Cultural Parallel Corpus, or asks about evaluating this task. Reports CSI-Match.
    3 repo stars
  28. ▌
    Curiosity Redteam Eval · qhjqhj00
    Evaluates automated red-teaming methods on their ability to generate diverse and effective prompts that elicit toxic responses from target LLMs. It probes both the effectiveness (toxicity elicitation rate) and diversity (textual and semantic variation) of generated test cases across text continuation and instruction-following tasks. Use when the user wants to benchmark on IMDb review dataset, Alpaca dataset, Databricks dataset, or asks about evaluating this task. Reports toxic response rate.
    3 repo stars
  29. ▌
    Curlora Continual Eval · qhjqhj00
    Tests a model's ability to learn sequentially across multiple NLP tasks while retaining prior knowledge. It specifically probes catastrophic forgetting mitigation during continual fine-tuning by measuring performance drops on earlier tasks after learning new ones. Use when the user wants to benchmark on GLUE-MRPC, GLUE-SST-2, Sentiment140, WikiText-2, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  30. ▌
    Cxmind Chest Xray Eval · qhjqhj00
    Evaluates multimodal large language models on chest X-ray diagnosis across visual understanding, text generation, spatiotemporal alignment, and foundational medical language capabilities. It probes the model's ability to interpret radiological images, generate clinical reports, localize anomalies, and reason over medical text. Use when the user wants to benchmark on MIMIC-CXR & CheXpert, OpenI, Language Datasets (CHIP-CDN, CMeEE, IMCS-V2-MRG, DDx-basic, MedSafety, MedHG, Med-Exam), MS-CXR, RSNA, CXR-AL14, or asks about evaluating this task. Reports Accuracy (Acc).
    3 repo stars
  31. ▌
    Depth Anything Ac Eval · qhjqhj00
    Evaluates zero-shot monocular relative depth estimation robustness under complex environmental conditions such as low light, adverse weather (rain, fog, snow), and synthetic noise. It probes the model's ability to recover fine-grained spatial relationships and object boundaries from degraded inputs without fine-tuning. Use when the user wants to benchmark on DA-2K (multi-condition), NuScenes-night, Robotcar-night, Driving-Stereo, KITTI-C, KITTI, NYU-D, Sintel, ETH3D, DIODE, or asks about evaluating this task. Reports AbsRel, δ1.
    3 repo stars
  32. ▌
    Detailverifybench Eval · qhjqhj00
    Evaluates multimodal large language models' ability to pinpoint erroneous content at the token level within long-form image captions. It probes whether models can distinguish between visually grounded facts and hallucinated details by localizing specific tokens that contradict the input image. Use when the user wants to benchmark on DetailVerifyBench, or asks about evaluating this task. Reports token-level F1.
    3 repo stars
  33. ▌
    Dexcanvas Success Eval · qhjqhj00
    This evaluation probes a robot policy's ability to successfully reproduce human-demonstrated dexterous manipulation trajectories in physics simulation. It measures robustness by testing performance under nominal conditions and under controlled initial pose perturbations. Use when the user wants to benchmark on DexCanvas, or asks about evaluating this task. Reports success rate.
    3 repo stars
  34. ▌
    Dino V2 Radiology Eval · qhjqhj00
    Evaluates the cross-task generalizability of the DINOv2 vision foundation model on medical image analysis tasks, specifically disease classification and organ segmentation across X-ray, CT, and MRI modalities. Use when the user wants to benchmark on NIH Chest X-ray, CheXpert, SARS-CoV-2, Brain Tumor, Montgomery County (MC), AMOS, MSD Heart, MSD Hipp, MSD Spleen, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  35. ▌
    Disrouter Routing Eval · qhjqhj00
    Evaluates a distributed multi-agent routing system's ability to autonomously assign queries to the most appropriate LLM based on intrinsic self-assessment. It probes the trade-off between routing accuracy and inference cost across diverse mathematical, commonsense, and reading comprehension benchmarks. Use when the user wants to benchmark on GSM8K, ARC, MMLU, RACE_HIGH, OpenbookQA, DROP, CosmosQA, SQuAD, HellaSwag, HeadQA, or asks about evaluating this task. Reports utility.
    3 repo stars
  36. ▌
    Document Haystack Eval · qhjqhj00
    This benchmark evaluates Vision Language Models' ability to retrieve specific textual or multimodal "needles" embedded within long documents ranging from 5 to 200 pages. It probes visual-text alignment, long-context retrieval capabilities, and performance degradation as document length and token consumption increase. Use when the user wants to benchmark on Document Haystack, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  37. ▌
    Drug Pair Scoring Eval · qhjqhj00
    This evaluation benchmarks deep learning architectures on predicting drug-drug interactions, polypharmacy side effects, and drug synergy. It measures how well models encode molecular graphs and combine them to score pairwise biological outcomes across multiple pharmacological domains. Use when the user wants to benchmark on TWOSIDES, Drugbank DDI, DrugComb, DrugCombDB, OncolyPharm, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  38. ▌
    Drugbank Hetionet Eval · qhjqhj00
    Evaluates the ability of matrix completion algorithms to predict missing biological interactions (drug-target or compound-disease) using sparse association matrices and side information. Use when the user wants to benchmark on DrugBank, Hetionet (Drug Repurposing), or asks about evaluating this task. Reports AUPR.
    3 repo stars
  39. ▌
    Drv Code Security Eval · qhjqhj00
    Evaluates the effectiveness of a Detect-Repair-Verify (DRV) workflow for fixing security vulnerabilities in LLM-generated code across different programming languages and granularity scopes (project, requirement, file). It measures how well iterative repair converges to a state that is both functionally correct and secure. Use when the user wants to benchmark on Custom LLM-generated code artifacts (JS, PHP, Python), or asks about evaluating this task. Reports S\C Yield Rate.
    3 repo stars
  40. ▌
    Dualrec Movielens Eval · qhjqhj00
    Evaluates a hybrid sequential and LLM-based framework for next-item movie recommendation. It probes the model's ability to capture temporal user preferences and semantic genre consistency to predict the next movie a user will watch. Use when the user wants to benchmark on MovieLens-1M, or asks about evaluating this task. Reports NDCG@5.
    3 repo stars
  41. ▌
    Ebible Benchmarks Eval · qhjqhj00
    Machine translation performance on low-resource languages using verse-aligned Bible texts. It probes model robustness across different biblical book genres (Gospels, Epistles, OT books) and the utility of related language data for translation. Use when the user wants to benchmark on eBible Corpus, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  42. ▌
    Ecg Preprocessing Eval · qhjqhj00
    Evaluates how ECG signal pre-processing techniques, particularly down-sampling rates, affect the performance of multi-label time-series classification models for diagnosing heart conditions. It probes the trade-off between signal fidelity, computational cost, and diagnostic accuracy across varying sampling frequencies. Use when the user wants to benchmark on Unspecified multi-label ECG datasets, or asks about evaluating this task. Reports MRR.
    3 repo stars
  43. ▌
    Edge Lm Inference Eval · qhjqhj00
    Evaluates the feasibility and performance trade-offs of running small generative language models on edge hardware. It probes memory constraints, inference latency, token throughput, and energy efficiency across different quantization schemes and system configurations. Use when the user wants to benchmark on None (system-level inference benchmark), or asks about evaluating this task. Reports generation_throughput.
    3 repo stars
  44. ▌
    Egoscreen Emotion Eval · qhjqhj00
    Evaluates a model's ability to predict human emotional responses to movie scenes from an egocentric, first-person screen-view perspective. It probes multimodal long-context reasoning by combining visual frames, audio cues, and narrative summaries to handle domain shifts from cinematic to realistic viewing conditions. Use when the user wants to benchmark on EgoScreen-Emotion (ESE), or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  45. ▌
    Embodied Reasoner Eval · qhjqhj00
    Evaluates an agent's ability to perform long-horizon embodied interactive tasks by synergizing visual search, reasoning, and action. It probes spatial reasoning, self-reflection, and planning capabilities in both simulated and real-world environments. Use when the user wants to benchmark on Unspecified (Simulated & Real-world tasks), or asks about evaluating this task. Reports success_rate.
    3 repo stars
  46. ▌
    Energy First Arch Eval · qhjqhj00
    Evaluates the classification accuracy and training energy efficiency of biologically-inspired and physics-guided neural architectures against conventional baselines across diverse data modalities. It probes whether action-principle regularization yields modality-specific performance gains and reduced internal activation energy without accuracy loss. Use when the user wants to benchmark on Fashion-MNIST, CIFAR-10, DVS Gesture, SHD, SSC, WESAD, DREAMER, SEED-IV, 20newsgroups, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  47. ▌
    Ethos Hate Speech Eval · qhjqhj00
    Evaluates the ability of NLP models to detect and classify hate speech in social media comments. It probes both binary hate/non-hate classification and multi-label categorization across specific demographic/identity-based hate categories. The benchmark emphasizes handling overlapping labels and class imbalance typical of real-world user-generated text. Use when the user wants to benchmark on ETHOS, or asks about evaluating this task. Reports F1-score (macro).
    3 repo stars
  48. ▌
    Ezsql SQL To Text Eval · qhjqhj00
    Evaluates a model's capability to generate fluent natural language descriptions from SQL queries (SQL-to-text) and measures how well the generated text can augment training data for Text-to-SQL parsers. Use when the user wants to benchmark on WikiSQL, Spider, or asks about evaluating this task. Reports BLEU-4.
    3 repo stars
  49. ▌
    Fedgraphnn System Eval · qhjqhj00
    Evaluates the computational efficiency and security overhead of federated graph neural network training. Measures how system-level metrics like training time, FLOPs, and parameter counts scale across diverse graph datasets under non-IID data partitioning and secure aggregation protocols. Use when the user wants to benchmark on SIDER, BACE, Clintox, BBBP, Tox21, FreeSolv, ESOL, Lipo, hERG, QM9, Ciao, Epinions, CORA, Citeseer, DBLP, PubMed, or asks about evaluating this task. Reports Wall-clock Time.
    3 repo stars
  50. ▌
    Fewshot Ft Vs Icl Eval · qhjqhj00
    Evaluates the in-domain and out-of-domain generalization capabilities of large language models adapted via few-shot fine-tuning versus in-context learning across standard natural language inference and paraphrase detection benchmarks. Use when the user wants to benchmark on MNLI, RTE, QQP, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  51. ▌
    Ffdl Platform Overhead · qhjqhj00
    Evaluates the runtime performance overhead and scalability of a containerized deep learning platform (FfDL) compared to bare-metal and specialized hardware, using standard image classification benchmarks. It measures how much throughput degrades when running DL training jobs in a Kubernetes-based multi-tenant environment versus direct execution. Use when the user has predictions and gold and needs to compute performance_overhead_pct.
    3 repo stars
  52. ▌
    Financial Exam QA Eval · qhjqhj00
    Evaluates whether LLMs can perform domain-level conceptual understanding and precise financial reasoning across multilingual professional certification exams. It probes analytical rigor and regulatory knowledge integration rather than simple factual recall. Use when the user wants to benchmark on EFPA, GRFinQA, CFA, CPA, BBF, SAHM, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  53. ▌
    Finteam Financial Eval · qhjqhj00
    Evaluates a multi-agent financial system's ability to answer real-world investor inquiries across macroeconomic, industry, and company analysis scenarios. Probes accuracy, thoroughness, clarity, and professional financial reasoning through automated LLM-judging and human preference testing. Use when the user wants to benchmark on NGA Grand Era Investor Inquiries, or asks about evaluating this task. Reports Overall Score.
    3 repo stars
  54. ▌
    Flood Forecasting Eval · qhjqhj00
    Evaluates a model's ability to predict river water levels and forecast floods using spatiotemporal radar precipitation data. It probes the model's accuracy across multiple forecasting lead times (2h to 12h) and its robustness in capturing extreme hydrological events compared to baseline and deep learning models. Use when the user wants to benchmark on Goslar, Göttingen, or asks about evaluating this task. Reports NSE.
    3 repo stars
  55. ▌
    Framework Latency Eval · qhjqhj00
    Evaluates the computational efficiency and hardware utilization of five deep learning frameworks (Caffe, Neon, TensorFlow, Theano, Torch) across standard neural network architectures on CPU and GPU hardware. Use when the user wants to benchmark on MNIST, ImageNet, IMDB, or asks about evaluating this task. Reports forward pass time (ms).
    3 repo stars
  56. ▌
    Freshwiki Article Eval · qhjqhj00
    Evaluates the ability of LLMs to generate comprehensive, well-organized, and verifiable Wikipedia-like articles from a given topic. It probes outline planning, factual coverage, structural coherence, and source grounding. Use when the user wants to benchmark on FreshWiki, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  57. ▌
    Full Duplex Bench Eval · qhjqhj00
    Evaluates real-time interactive behaviors in full-duplex spoken dialogue models. It specifically probes turn-taking, pause handling, backchanneling, and interruption management capabilities without relying on human studies. Use when the user wants to benchmark on Full-Duplex-Bench, or asks about evaluating this task. Reports descriptive metrics.
    3 repo stars
  58. ▌
    Gazebo Simulation Eval · qhjqhj00
    Evaluates an MPC-based autonomous driving controller's ability to perform collision avoidance and lane maneuvers (overtaking, merging, following) in dynamic environments using a physics-based simulator. It tests the controller's real-time feasibility and trajectory smoothness under varying traffic densities. Use when the user wants to benchmark on Gazebo Simulation, or asks about evaluating this task. Reports computation time.
    3 repo stars
  59. ▌
    Gcai Constitution Eval · qhjqhj00
    Evaluates the moral grounding, coherence, fairness, and real-world applicability of AI alignment constitutions through human surveys, alongside the downstream safety alignment and general capabilities of fine-tuned language models. Use when the user wants to benchmark on BABELSCAPE/ALERT, MMLU, Social Bias BBQ, or asks about evaluating this task. Reports 5-point Likert rating.
    3 repo stars
  60. ▌
    Globo Session Rec Eval · qhjqhj00
    This benchmark evaluates session-based news recommendation systems by predicting the next article a user will click based on their recent interaction history. It probes a model's ability to capture short-term user intent, handle temporal dynamics, and leverage both content and contextual features in a streaming environment. Use when the user wants to benchmark on Globo.com, or asks about evaluating this task. Reports HR@5.
    3 repo stars
  61. ▌
    Glue Low Resource Eval · qhjqhj00
    Evaluates the generalization and stability of finetuned pretrained language models on low-resource NLP tasks. It probes how well models adapt to sentiment classification, natural language inference, paraphrasing, similarity assessment, and linguistic acceptability when trained on severely limited data (300–1000 examples). Use when the user wants to benchmark on GLUE, or asks about evaluating this task. Reports averaged evaluation metrics.
    3 repo stars
  62. ▌
    Gmtkn55 Molecular Eval · qhjqhj00
    Evaluates the accuracy of density functionals in predicting broad molecular properties, including atomization energies, barrier heights, and noncovalent interactions, relative to high-level quantum chemistry reference data. Use when the user wants to benchmark on GMTKN55 (specifically GMTKN53 subset), or asks about evaluating this task. Reports WTMAD1.
    3 repo stars
  63. ▌
    Gnn Vs Dnn Lhc Bb Eval · qhjqhj00
    Compares graph neural networks against deep fully-connected feedforward networks for binary classification of b-quark pairs in top-quark-antiquark collisions at the LHC. It probes whether explicit relational inductive biases in GNNs outperform permutation-invariant DNNs when provided with equivalent kinematic and relational features. Use when the user wants to benchmark on LHC t\bar{t} b\bar{b} event classification, or asks about evaluating this task. Reports mean ROC-AUC ($\mu_{\mathrm{AUC}}$).
    3 repo stars
  64. ▌
    Golden Touchstone Eval · qhjqhj00
    This benchmark evaluates the capability of large language models to perform a wide range of financial natural language processing tasks in both English and Chinese. It probes domain-specific understanding, information extraction, reasoning, and generation across sentiment analysis, classification, entity/relation extraction, summarization, question answering, and stock movement prediction. Use when the user wants to benchmark on FPB, Fiqa-SA, Headlines, FOMC, lendingclub, NER, FinRE, CFA, EDTSUM, Finqa, Convfinqa, DJIA, FinFe-CN, FinNL-CN, FinESE-CN, FinRE-CN, FinQa-CN, or asks about evaluating this task. Reports Weighted-F1.
    3 repo stars
  65. ▌
    Gpt4 Based Exact Match · qhjqhj00
    Evaluates whether a language model's final numerical answer to a grade-school math word problem matches the ground truth. It uses an external LLM to extract the final answer from the model's generated solution and compares it against the gold answer. Use when the user has predictions and gold and needs to compute GPT4-based-Exact-Match.
    3 repo stars
  66. ▌
    Gravity Inversion Eval · qhjqhj00
    Evaluates a 3D gravity inversion framework's ability to recover subsurface density structures from gravity data, measuring both computational efficiency (scaling and speedup) and geological accuracy (recovery of synthetic anomalies and fit to field observations). Use when the user wants to benchmark on Synthetic and Field Gravity Inversion Benchmarks, or asks about evaluating this task. Reports speedup.
    3 repo stars
  67. ▌
    Guicourse Gui Nav Eval · qhjqhj00
    Evaluates vision-language models' ability to navigate graphical user interfaces by predicting correct action types and precise screen coordinates. It probes OCR, pixel-level grounding, and multi-step task planning across web and mobile environments. Use when the user wants to benchmark on GUIAct, Mind2Web, AITW, or asks about evaluating this task. Reports StepSR.
    3 repo stars
  68. ▌
    Gw Reconstruction Eval · qhjqhj00
    Evaluates the fidelity of model-agnostic Bayesian waveform reconstruction for unbound binary black hole fly-bys across different detector noise environments and frame function parameterizations. It also quantifies the astrophysical detection sensitivity and expected event rates for current and next-generation interferometers. Use when the user wants to benchmark on Simulated Hyperbolic BBH Encounters, or asks about evaluating this task. Reports overlap.
    3 repo stars
  69. ▌
    H Sinn Turbulence Eval · qhjqhj00
    Evaluates a Convolutional Autoencoder's ability to compress and reconstruct geophysical turbulence fields while preserving high-order statistical moments. It specifically probes the model's capacity to capture non-Gaussian, intermittent structures like extreme vertical drafts without degrading point-wise accuracy. Use when the user wants to benchmark on Stratified turbulence simulation data, or asks about evaluating this task. Reports MAPE on kurtosis ($K_w$).
    3 repo stars
  70. ▌
    Habitat Benchmark Eval · qhjqhj00
    Evaluates a robot's ability to perform long-horizon mobile manipulation and object rearrangement tasks in simulated environments. It probes hierarchical planning, whole-body continuous control, and robust recovery from failures across multi-step subtask sequences. Use when the user wants to benchmark on Habitat Benchmark, or asks about evaluating this task. Reports completion rate.
    3 repo stars
  71. ▌
    Habitat Predictor Eval · qhjqhj00
    Probes the ability of a runtime-based predictor to accurately estimate GPU training iteration execution times and cost-normalized throughput across different DNN architectures and GPU generations without requiring full training runs. Use when the user wants to benchmark on ImageNet, WMT'16, LSUN, or asks about evaluating this task. Reports average prediction error.
    3 repo stars
  72. ▌
    Hallucination Tax Eval · qhjqhj00
    Evaluates whether reinforcement finetuned language models appropriately refuse to answer unanswerable or ambiguous questions, and measures their accuracy on standard solvable math benchmarks to ensure performance is not degraded by the refusal training. Use when the user wants to benchmark on UWMP, SelfAware, Synthetic Unanswerable Math (SUM), GSM8K, Minerva, MATH-500, OlympiadBench, AMC23, or asks about evaluating this task. Reports refusal_rate.
    3 repo stars
  73. ▌
    He Xingwei Sari Metric · qhjqhj00
    Compute He-Xingwei/sari_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of He-Xingwei/sari_metric.
    3 repo stars
  74. ▌
    Health Orsc Bench Eval · qhjqhj00
    This benchmark evaluates large language models' tendency to over-refuse benign health-related queries and their ability to provide safe, helpful completions in medical contexts. It specifically probes the trade-off between safety alignment and utility by measuring refusal rates on carefully curated boundary prompts across varying difficulty levels. Use when the user wants to benchmark on Health-ORSC-Bench, or asks about evaluating this task. Reports Over-Refusal Rate (ORR).
    3 repo stars
  75. ▌
    Helms Exploration Eval · qhjqhj00
    Evaluates the exploration efficiency and adaptability of a robot planner in unknown environments. It measures how well the system covers free space while minimizing travel distance and adapting to natural language preferences. Use when the user wants to benchmark on Simulated dungeon environments, Indoor office environment (130m x 100m), or asks about evaluating this task. Reports Travel Distance.
    3 repo stars
  76. ▌
    Hinmix Robust Cmt Eval · qhjqhj00
    Evaluates machine translation models on code-mixed and noisy Hindi-English and Bengali-English text, measuring robustness to script variations, romanization, and synthetic noise. The protocol tests both in-domain performance on the HINMIX corpus and out-of-domain generalizability on LinCE, SpokenTutorial, and IITB Hi-En. It also assesses zero-shot transfer to unseen code-mixed Bengali-English translation. Use when the user wants to benchmark on HINMIX, or asks about evaluating this task. Reports SacreBLEU.
    3 repo stars
  77. ▌
    Hintel Alignbench Eval · qhjqhj00
    Evaluates multimodal large language models on vision-language tasks in Hindi and Telugu, measuring performance regression when transitioning from English to these Indian languages. It probes native-language visual question answering, mathematical reasoning, and multiple-choice comprehension across STEM and cultural domains. Use when the user wants to benchmark on HinTel-AlignBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  78. ▌
    Histoires Morales Eval · qhjqhj00
    Evaluates the linguistic quality and cultural appropriateness of the Histoires Morales dataset. It measures reference-free translation accuracy and assesses whether moral norms and actions align with native French speakers' cultural values. Use when the user wants to benchmark on Histoires Morales, or asks about evaluating this task. Reports CometKiwi22.
    3 repo stars
  79. ▌
    Holistic Motion2d Eval · qhjqhj00
    Evaluates the quality, text-motion alignment, and diversity of generated 2D whole-body human motion sequences conditioned on text prompts. It probes the model's ability to capture fine-grained spatial-temporal dynamics and handle occlusions or noisy 2D pose data. Use when the user wants to benchmark on Holistic-Motion2D, or asks about evaluating this task. Reports FID.
    3 repo stars
  80. ▌
    Hopkins Rfo Bench Eval · qhjqhj00
    Evaluates object detection models on chest X-rays for identifying critical retained foreign objects (RFOs) like sponges and needles. It probes both classification accuracy and precise localization of rare medical anomalies under data-scarce conditions. Use when the user wants to benchmark on Hopkins RFOs Bench, or asks about evaluating this task. Reports ACC.
    3 repo stars
  81. ▌
    Hula Antispoofing Eval · qhjqhj00
    This benchmark evaluates anti-spoofing systems for synthetic speech detection, with a specific focus on prosody-awareness, emotional/expressive spoofing, and cross-lingual robustness. It tests whether models can distinguish real speech from TTS, VC, and adversarial attacks across diverse channel conditions and languages. Use when the user wants to benchmark on ASVspoof 2019 (LA), ASVspoof 2021 (LA), ASVspoof 2024 (Track 1), EmoFake, Mixed Emotions, ADD 2022 (Track 1), HABLA, or asks about evaluating this task. Reports EER%.
    3 repo stars
  82. ▌
    Humanoid Everyday Eval · qhjqhj00
    Evaluates imitation learning and vision-language-action policies on open-world humanoid manipulation. It probes robustness to high-dimensional action spaces, multimodal sensor fusion, and fine-grained visuospatial perception across locomotion, tool use, and precise manipulation tasks. Use when the user wants to benchmark on Humanoid Everyday, or asks about evaluating this task. Reports success rate.
    3 repo stars
  83. ▌
    Imagenet Transfer Eval · qhjqhj00
    Evaluates the transferability of adversarially robust ImageNet pretraining to downstream classification tasks. It probes whether robustness induces more generalizable and discriminative feature representations compared to standard training, measured under both fixed-feature and full-network fine-tuning settings. Use when the user wants to benchmark on Birdsnap, Caltech-101, Caltech-256, CIFAR-10, CIFAR-100, Describable Textures (DTD), FGVC Aircraft, Food-101, Oxford 102 Flowers, Oxford-IIIT Pets, SUN397, Stanford Cars, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  84. ▌
    Infinite Dsprites Eval · qhjqhj00
    Evaluates continual learning methods on a procedurally generated benchmark of 500 shape classification tasks. It probes a model's ability to learn incrementally over a long horizon without catastrophic forgetting, while maintaining open-set recognition and one-shot generalization capabilities on unseen shapes. Use when the user wants to benchmark on Infinite dSprites (idSprites), or asks about evaluating this task. Reports average test accuracy.
    3 repo stars
  85. ▌
    Infinity Instruct Eval · qhjqhj00
    Evaluates the conversational and foundational capabilities of LLMs fine-tuned on instruction datasets, comparing them against proprietary and open-source baselines across multiple standard benchmarks. Use when the user wants to benchmark on AlpacaEval 2.0, Arena-Hard, MT-Bench, MATH, GSM-8K, HumanEval, MBPP, MMLU, LUC-EVAL, or asks about evaluating this task. Reports Overall*.
    3 repo stars
  86. ▌
    Integrated Brier Score · qhjqhj00
    Evaluates the calibration and predictive accuracy of ensemble forecasting methods for time-to-event outcomes in meteorology. It compares how well different combination techniques predict the timing of events like the first hard freeze. Use when the user has predictions and gold and needs to compute Mean Integrated Brier Score (IBS).
    3 repo stars
  87. ▌
    Interactive Audio Eval · qhjqhj00
    Evaluates Large Audio Models (LAMs) on real-world, task-oriented voice assistant interactions by capturing user preferences through open-ended pairwise comparisons. It measures how well models align with actual user needs and preferences in an interactive setting, rather than relying on static reference-based benchmarks. Use when the user wants to benchmark on TalkArena Interactive User Preferences, or asks about evaluating this task. Reports Bradley-Terry model score.
    3 repo stars
  88. ▌
    Interpreter Bench Eval · qhjqhj00
    Evaluates the runtime performance and memory efficiency of a speculatively staged Python interpreter against standard baselines like CPython and PyPy. It probes the interpreter's ability to eliminate dynamic type-checking overhead and optimize instruction dispatch through compile-time specialization. Use when the user wants to benchmark on Computer Language Benchmarks Game, or asks about evaluating this task. Reports speedup.
    3 repo stars
  89. ▌
    Inverse Rendering Eval · qhjqhj00
    Evaluates a model's capability to perform novel view synthesis, decompose scene properties (albedo, normals, roughness), and relight scenes under new lighting conditions using Gaussian surfels. It specifically probes the model's ability to model indirect illumination and inter-reflections without relying on pre-trained novel view synthesis data. Use when the user wants to benchmark on TensoIR*, Synthetic4Relight*, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  90. ▌
    Jialinsong Apps Metric · qhjqhj00
    Compute jialinsong/apps_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of jialinsong/apps_metric.
    3 repo stars
  91. ▌
    Kbmlcoding Apps Metric · qhjqhj00
    Compute kbmlcoding/apps_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of kbmlcoding/apps_metric.
    3 repo stars
  92. ▌
    Kitti Eigen Depth Eval · qhjqhj00
    Evaluates the accuracy of self-supervised monocular depth estimation models on urban driving scenes. It probes the model's ability to predict per-pixel depth from a single image or video sequence, handling occlusions, moving objects, and scale ambiguity. Use when the user wants to benchmark on KITTI 2015 (Eigen split), Make3D, or asks about evaluating this task. Reports Abs Rel, δ < 1.25.
    3 repo stars
  93. ▌
    Kv Cache Eviction Eval · qhjqhj00
    Evaluates the quality of learned key-value (KV) cache eviction policies in preserving long-context reasoning and generation capabilities under strict memory constraints. It measures how well different compression strategies retain critical tokens without access to query-specific attention scores during the compression phase. Use when the user wants to benchmark on RULER-4k, OASST2-4k, BoolQ, ARC-Challenge, MMLU, HellaSwag, GovReport, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  94. ▌
    Lana I Geothermal Eval · qhjqhj00
    Evaluates a multimodal machine learning workflow's ability to predict 3D subsurface geological, hydrogeological, and geophysical features from sparse, heterogeneous field data. It tests the model's generalization capability using transductive learning and mutual information maximization across five cross-validation splits. Use when the user wants to benchmark on Lana'i 3D Subsurface Grid, or asks about evaluating this task. Reports R-squared.
    3 repo stars
  95. ▌
    Language Transfer Eval · qhjqhj00
    Evaluates multilingual language transfer by measuring how well models adapt to German and Bulgarian while preserving source language (English) capabilities. Probes catastrophic forgetting and cross-lingual generalization across reasoning, math, reading comprehension, and commonsense tasks. Use when the user wants to benchmark on Multilingual Language Transfer Benchmarks (EN/DE/BG), or asks about evaluating this task. Reports normalized accuracy.
    3 repo stars
  96. ▌
    Lexicon Grounding Eval · qhjqhj00
    Evaluates how well a model learns word meanings and general language modeling performance when trained with lexicon-level contrastive visual grounding. Probes concrete vs. abstract word acquisition, verb relation learning, and next-token prediction accuracy on held-out text. Use when the user wants to benchmark on Word Relatedness, Semantic Feature Prediction, Context Understanding, Lexical Relation Prediction, SimVerb-3500, or asks about evaluating this task. Reports Perplexity.
    3 repo stars
  97. ▌
    Libero Liberoplus Eval · qhjqhj00
    Evaluates a Vision-Language-Action model's ability to perform precise robotic manipulation and maintain robustness under environmental perturbations. It probes spatial understanding, object manipulation, instruction following, and long-horizon task execution in both standard and perturbed simulation environments. Use when the user wants to benchmark on LIBERO, LIBERO-Plus, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  98. ▌
    LLM Decision Bias Eval · qhjqhj00
    This evaluation probes whether explicitly unbiased LLMs exhibit automatic, stereotype-driven preferences in decision-making scenarios. It measures implicit bias by comparing model agreement rates across stereotypical versus counter-stereotypical social categories (e.g., gender-career, race-health) using a single-choice decision prompt rather than a relative comparison. Use when the user wants to benchmark on LLM Decision Bias (Absolute Variant), or asks about evaluating this task. Reports normalized yes-to-no ratio.
    3 repo stars
  99. ▌
    Lod Ood Detection Eval · qhjqhj00
    Evaluates a model's ability to distinguish in-distribution (ID) samples from out-of-distribution (OOD) samples using a threshold-free loss-difference clustering approach. It measures detection performance across standard benchmarks with diverse natural images and hard benchmarks where OOD classes share the same source dataset as ID. Use when the user wants to benchmark on CIFAR100, SVHN, Places, LSUN-Crop, LSUN-Resize, Textures, CIFAR10, TinyImageNet, or asks about evaluating this task. Reports FPR95.
    3 repo stars
  100. ▌
    Lvwerra Accuracy Score · qhjqhj00
    Compute lvwerra/accuracy_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of lvwerra/accuracy_score.
    3 repo stars