all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 21 of 76

  1. ▌
    Multilingual Transfer Eval · qhjqhj00
    Probes cross-lingual transfer and multilingual representation learning across natural language inference, named entity recognition, question answering, and English text classification. The protocol evaluates how well a model trained on English data generalizes to 14 other languages, while also measuring per-language and multilingual fine-tuning performance. Use when the user wants to benchmark on XNLI, CoNLL-2002/2003, MLQA, GLUE, or asks about evaluating this task. Reports F1 score, Accuracy.
    3 repo stars
  2. ▌
    Multimodal Benchmarks Eval · qhjqhj00
    Evaluates multimodal perception and reasoning capabilities across image, video, and audio understanding tasks. It probes the model's ability to process heterogeneous modalities and answer complex questions or transcribe speech accurately. Use when the user wants to benchmark on AI2D, MMMU, MMStar, OCRBench, MMVet, Mathvista, LongVideoBench, DiDeMo, AVQA, MVBench, Video-MME, Aishell1, LibriSpeech, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  3. ▌
    Multioff Hateful Meme Eval · qhjqhj00
    Binary classification of memes as offensive or non-offensive. It probes a model's ability to detect hate speech in multimodal content by leveraging serialized scene graphs and knowledge graph entities alongside raw text. Use when the user wants to benchmark on MultiOFF, or asks about evaluating this task. Reports F1 score (offensive class).
    3 repo stars
  4. ▌
    Mushroom Segmentation Eval · qhjqhj00
    This benchmark evaluates the zero-shot instance segmentation capability of models trained on synthetic data when applied to real-world agricultural scenes. It probes the model's ability to generalize across domain gaps, handling varying lighting, mushroom densities, and developmental stages without fine-tuning on real annotations. Use when the user wants to benchmark on Real-world on-field dataset, M18K, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  5. ▌
    Mv Adapter Multi View Eval · qhjqhj00
    Evaluates a diffusion adapter's ability to generate geometrically consistent multi-view images conditioned on text prompts or reference images with camera parameters. It measures visual fidelity, image-text alignment, and multi-view structural similarity against ground-truth 3D scans. Use when the user wants to benchmark on Objaverse, Google Scanned Objects (GSO), or asks about evaluating this task. Reports FID.
    3 repo stars
  6. ▌
    N2c2 Concept Relation Eval · qhjqhj00
    Evaluates clinical NLP models on extracting medical concepts and their relations from clinical notes. It probes the model's ability to handle nested/overlapped concepts and assesses cross-institutional generalization across different benchmark years. Use when the user wants to benchmark on n2c2 2018, n2c2 2022, n2c2 cross-institution (MIMIC-train/UW-test), or asks about evaluating this task. Reports strict micro-averaged F1-score.
    3 repo stars
  7. ▌
    Neural News Detection Eval · qhjqhj00
    Evaluates the ability of classifiers and LLMs to detect machine-generated news headlines across four languages. It probes cross-lingual generalization, robustness to zero-shot vs fine-tuned generators, and the effectiveness of linguistic vs transformer-based features for authenticity verification. Use when the user wants to benchmark on Multilingual Neural News Detection Benchmark, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  8. ▌
    Neuroncap Closed Loop Eval · qhjqhj00
    Evaluates closed-loop autonomous driving performance in simulated real-world scenarios, measuring safety and planning efficiency through collision avoidance and impact speed mitigation. Use when the user wants to benchmark on nuScenes, or asks about evaluating this task. Reports NeuroNCAP Score (NNS).
    3 repo stars
  9. ▌
    Newcicids Adversarial Eval · qhjqhj00
    Evaluates the adversarial robustness of tree ensemble models (RF, XGB, LGBM, EBM) on enterprise network intrusion detection using the corrected NewCICIDS dataset. It measures how well models maintain detection performance on benign and malicious traffic when subjected to constrained adversarial perturbations of time-series traffic features. Use when the user wants to benchmark on NewCICIDS, or asks about evaluating this task. Reports F1S.
    3 repo stars
  10. ▌
    Norwegian Transformer Eval · qhjqhj00
    Evaluates the cross-lingual transfer and domain adaptation capabilities of a Norwegian BERT model against multilingual and monolingual baselines on token-level (NER/POS) and sequence-level (sentiment/political affiliation) classification tasks. Use when the user wants to benchmark on NorNE, CoNLL-2003, NoReC, Norwegian Parliament Speeches, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  11. ▌
    Nsynth Reconstruction Eval · qhjqhj00
    Evaluates the ability of audio generative models to reconstruct raw musical note waveforms and interpolate timbre and pitch in a learned latent space. Probes phase preservation, harmonic structure modeling, and the decoupling of pitch and timbre information. Use when the user wants to benchmark on NSynth, or asks about evaluating this task. Reports Classification accuracy.
    3 repo stars
  12. ▌
    Nuscenes 3d Detection Eval · qhjqhj00
    Evaluates the capability of multi-view 3D object detection models to accurately localize and classify objects in autonomous driving scenes using camera inputs. It measures detection accuracy alongside computational efficiency and inference latency to assess real-time deployment feasibility. Use when the user wants to benchmark on NuScenes, or asks about evaluating this task. Reports NDS.
    3 repo stars
  13. ▌
    Nyu Breast Cancer Seg Eval · qhjqhj00
    Evaluates weakly-supervised segmentation and classification performance on high-resolution mammography images for detecting malignant and benign breast lesions. Use when the user wants to benchmark on NYU Breast Cancer Screening Dataset v1.0, or asks about evaluating this task. Reports Dice similarity coefficient.
    3 repo stars
  14. ▌
    Olive Branch Learning Eval · qhjqhj00
    Evaluates a topology-aware federated learning framework (OBL) with a satellite-assignment algorithm (CNASA) for space-air-ground integrated networks. It probes the trade-off between model accuracy and training latency under strict non-IID data distributions and varying network topologies. Use when the user wants to benchmark on MNIST, Fashion-MNIST, CIFAR-10, or asks about evaluating this task. Reports final global model accuracy.
    3 repo stars
  15. ▌
    Omni Recon Downstream Eval · qhjqhj00
    Evaluates a general-purpose NeRF framework on downstream 3D tasks, including real-time novel view synthesis, parameter-efficient 3D scene understanding, and text-guided 3D editing. It probes the model's ability to generalize to unseen scenes and adapt to various geometric and appearance tasks with minimal fine-tuning. Use when the user wants to benchmark on DTU, ScanNet, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  16. ▌
    Omnidpo Hallucination Eval · qhjqhj00
    This evaluation protocol assesses the ability of omni-modal large language models to avoid hallucinating non-existent audio or visual content when presented with contradictory or incomplete multimodal inputs. It specifically probes whether models can correctly identify real elements while resisting false affirmations of missing ones across text, vision, and audio modalities. Use when the user wants to benchmark on AVHBench, CMM, or asks about evaluating this task. Reports F1 Score, Hallucination Resistance (HR).
    3 repo stars
  17. ▌
    Omnigaiatoolreasoning Eval · qhjqhj00
    Evaluates native omni-modal AI agents' ability to perform multi-hop cross-modal reasoning across video, audio, and image inputs. It probes their capacity to integrate external tools (web search, browser, code execution) for evidence gathering and to produce verifiable open-form answers under varying task difficulties. Use when the user wants to benchmark on OmniGAIA, or asks about evaluating this task. Reports Pass@1.
    3 repo stars
  18. ▌
    Ophthalmic Multimodal Eval · qhjqhj00
    Evaluates multimodal large language models on clinical ophthalmic image interpretation, specifically diagnosing retinal and macular diseases from fundus photographs and optical coherence tomography (OCT) scans. It probes the models' ability to recognize normal conditions and identify specific pathological states across diverse disease categories. Use when the user wants to benchmark on Ophthalmic Multimodal Benchmark, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  19. ▌
    Opinion Summarization Eval · qhjqhj00
    Evaluates a model's ability to synthesize multiple product reviews into a single, coherent opinion summary. It probes the model's capacity to capture diverse aspects, maintain factual accuracy, and adhere to domain-specific stylistic constraints without relying on surface-level lexical overlap. Use when the user wants to benchmark on Amazon, Oposum+, Flipkart, or asks about evaluating this task. Reports human evaluation.
    3 repo stars
  20. ▌
    Pareto Interpretation Eval · qhjqhj00
    This evaluation probes a model's ability to synthesize interpretable surrogate models (decision diagrams) that explicitly trade off prediction fidelity against structural simplicity. It measures how well a method can navigate the Pareto front between accuracy and explainability without collapsing them into a single weighted objective. Use when the user wants to benchmark on Airplane Perception Module (AP), Bank Loan Predictor (BL), Theorem Prover Solvability Predictor (TP), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    PDF Malware Poisoning Eval · qhjqhj00
    Evaluates the robustness of embedded feature selection methods (LASSO, Ridge, Elastic Net) against training data poisoning attacks in a PDF malware detection setting. It measures how injected malicious samples manipulate feature selection stability and degrade classification performance. Use when the user wants to benchmark on Contagio + Web Benign PDFs, or asks about evaluating this task. Reports classification error.
    3 repo stars
  22. ▌
    Phrase Level Polarity Eval · qhjqhj00
    Classifies the sentiment polarity of specific phrases within tweets, requiring models to handle contextual disambiguation, sarcasm, and informal language. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports macro-averaged F1.
    3 repo stars
  23. ▌
    Pico Tinyml Benchmark Eval · qhjqhj00
    Evaluates real-time performance and resource efficiency of TinyML models on embedded hardware by measuring inference latency, CPU/memory utilization, and prediction confidence across different platforms. Use when the user wants to benchmark on Gesture Classification, Keyword Spotting, MobileNet V2, or asks about evaluating this task. Reports Average Inference Latency (ms).
    3 repo stars
  24. ▌
    Pio Detailed Labeling Eval · qhjqhj00
    Evaluates a model's ability to predict fine-grained hierarchical labels for tokens within identified P, I, and O spans. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.
    3 repo stars
  25. ▌
    Pointgat Molecule C10 Eval · qhjqhj00
    Evaluates a hybrid graph attention and 3D point cloud neural network's ability to predict quantum chemical properties and molecular physicochemical traits. It probes the model's capacity to integrate 2D topological graph features with 3D spatial geometry for accurate regression and classification of molecular energies and properties. Use when the user wants to benchmark on MoleculeNet, C10, or asks about evaluating this task. Reports MAE, R².
    3 repo stars
  26. ▌
    Possible Stories Ifsm Eval · qhjqhj00
    Assesses whether large language models can generate story endings that align with free-form instructions provided alongside a narrative context. It measures both instruction-following accuracy and the model's ability to produce distinct endings for different instructions. Use when the user wants to benchmark on Possible Stories, or asks about evaluating this task. Reports IFSM.
    3 repo stars
  27. ▌
    Preference Discerning Eval · qhjqhj00
    This benchmark evaluates a model's ability to dynamically adapt to evolving user preferences by conditioning on natural language preferences inferred from interaction history. It probes recommendation accuracy, fine- and coarse-grained preference steering, sentiment following, and history consolidation across multiple e-commerce and gaming datasets. Use when the user wants to benchmark on Amazon Beauty, Amazon Sports and Outdoors, Amazon Toys and Games, Steam, or asks about evaluating this task. Reports Recall@10.
    3 repo stars
  28. ▌
    Preset Voice Matching Eval · qhjqhj00
    Evaluates a privacy-regulated speech-to-speech translation framework that replaces voice cloning with matching to pre-consented preset voices. Probes classifier robustness, cross-lingual speech naturalness, and inference efficiency across multilingual scenarios. Use when the user wants to benchmark on RAVDESS, CGDD, CAFE, EmoDB, CREMA-D, or asks about evaluating this task. Reports NISQA.
    3 repo stars
  29. ▌
    Prior Polarity Degree Eval · qhjqhj00
    Evaluates the ability of sentiment lexicons or models to assign accurate real-valued polarity scores to individual terms, measuring rank correlation with gold standards. Use when the user wants to benchmark on Term test set, or asks about evaluating this task. Reports Kendall's τ coefficient.
    3 repo stars
  30. ▌
    Protein Generative AI Eval · qhjqhj00
    This evaluation framework probes the physical validity, functional success, and generalization capability of generative protein models. It emphasizes leakage-aware dataset splits, structural and docking quality metrics, and experimental validation pipelines to ensure designs are biologically plausible and functionally active. Use when the user wants to benchmark on PLINDER, PoseBusters, PDFBench, PoseX, VenusX, FragBench, GeomMotif, or asks about evaluating this task. Reports RMSD.
    3 repo stars
  31. ▌
    Protein Understanding Eval · qhjqhj00
    Evaluates protein language models on sequence understanding across multiple downstream tasks (structure, function, interactions, developability) and 3D structure prediction from single amino acid sequences. Use when the user wants to benchmark on CAMEO, CASP15, OOD Protein Sequences (UniProt), or asks about evaluating this task. Reports TM-score.
    3 repo stars
  32. ▌
    Psp Protein Structure Eval · qhjqhj00
    Evaluates protein structure prediction models on their ability to infer accurate 3D atomic coordinates from amino acid sequences. It specifically probes topological backbone similarity and side-chain accuracy when trained on large-scale distilled protein datasets. Use when the user wants to benchmark on CASP14 test set, PSP dataset, or asks about evaluating this task. Reports TM-score.
    3 repo stars
  33. ▌
    Ptbxl Multi Label Ecg Eval · qhjqhj00
    Evaluates deep learning architectures for multi-label classification of 12-lead ECG recordings into 23 diagnostic categories. Probes the trade-off between local morphological feature extraction and sequential temporal modeling under severe class imbalance. Use when the user wants to benchmark on PTB-XL, or asks about evaluating this task. Reports Macro AUROC.
    3 repo stars
  34. ▌
    Rakugo Listening Test Eval · qhjqhj00
    Evaluates the perceptual quality of synthesized rakugo speech by comparing it to professional human performances across multiple dimensions, including naturalness, character distinguishability, content understandability, entertainment value, and overall skill level. The benchmark probes whether TTS systems can capture the nuanced performance modeling required for traditional Japanese verbal entertainment. Use when the user wants to benchmark on Misomame, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
    3 repo stars
  35. ▌
    Recsys Challenge 2015 Eval · qhjqhj00
    Evaluates session-based recommendation models by predicting the next item in a user's clickstream sequence. It probes the model's ability to capture temporal dynamics and handle data sparsity in e-commerce sessions. Use when the user wants to benchmark on RecSys Challenge 2015, or asks about evaluating this task. Reports Recall@20.
    3 repo stars
  36. ▌
    Relative Mean Square Error · qhjqhj00
    Evaluates the prediction accuracy of data-driven neural network closures for kinetic theory of active fluids against reference kinetic simulations and traditional analytical closures. It probes the model's ability to capture rotational symmetries, generalize across parameter regimes, and resolve topological defects in fluid dynamics. Use when the user has predictions and gold and needs to compute RMSE.
    3 repo stars
  37. ▌
    Repeatnet Session Rec Eval · qhjqhj00
    Evaluates session-based recommendation models on predicting the next item in a user session. It probes the model's ability to capture sequential patterns and handle repeat consumption behaviors across e-commerce and music domains. Use when the user wants to benchmark on YOOCHOOSE, DIGINETICA, LASTFM, or asks about evaluating this task. Reports Recall@k.
    3 repo stars
  38. ▌
    Rl Dialogue Benchmark Eval · qhjqhj00
    Evaluates the robustness and generalization of reinforcement learning-based dialogue management policies across varying simulated environments. It probes how well RL algorithms handle different domain sizes, user behavior profiles, and noisy speech input channels in task-oriented spoken dialogue systems. Use when the user wants to benchmark on PyDial simulated environments, or asks about evaluating this task. Reports average success rate.
    3 repo stars
  39. ▌
    Robust Mean Relative Error · qhjqhj00
    Evaluates a model's ability to forecast future values in a financial time-series dataset given historical observations. It measures prediction accuracy using a capped relative error to prevent extreme outliers from dominating the score. Use when the user has predictions and gold and needs to compute robust mean relative error.
    3 repo stars
  40. ▌
    Rst Discourse Parsing Eval · qhjqhj00
    Evaluates a model's ability to predict RST-style discourse tree structure and nuclearity relations between elementary discourse units (EDUs). It probes both intra-domain and inter-domain generalization of discourse parsing across different text genres (news, instructions, reviews). Use when the user wants to benchmark on RST-DT, Instr-DT, MEGA-DT, Yelp13-DT, or asks about evaluating this task. Reports Parseval (Par.) / RST-Parseval (R-Par.).
    3 repo stars
  41. ▌
    Sam Zero Shot Medical Eval · qhjqhj00
    Evaluates the zero-shot segmentation capability of the Segment Anything Model (SAM) across diverse medical imaging modalities and anatomical structures. It probes the model's ability to generalize without task-specific fine-tuning to structured medical targets like organs, lesions, and retinal layers. Use when the user wants to benchmark on Skin Lesion Analysis Toward Melanoma Detection (ISIC), DoFE (Drishiti-GS, RIM-ONE-r3, REFUGE amalgamated), AMOS, MICCAI 2017 Robotic Instrument Segmentation, Chest X-ray, Rat Colon, AROI, or asks about evaluating this task. Reports Dice similarity coefficient.
    3 repo stars
  42. ▌
    Sasrec Sequential Rec Eval · qhjqhj00
    Evaluates a model's ability to predict the next item in a user's interaction sequence based on historical behavior. It probes the model's capacity to capture long-range dependencies and adapt to varying data sparsity across different domains. Use when the user wants to benchmark on Amazon (Beauty), Amazon (Games), Steam, MovieLens-1M, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  43. ▌
    Screen Spot Grounding Eval · qhjqhj00
    Evaluates a model's ability to locate specific GUI elements from a screenshot given a text instruction. It measures both coarse localization accuracy and fine-grained bounding box overlap across desktop, mobile, and web platforms. Use when the user wants to benchmark on ScreenSpot, or asks about evaluating this task. Reports grounding accuracy.
    3 repo stars
  44. ▌
    Script Identification Eval · qhjqhj00
    Evaluates the accuracy of a script identification tool on multilingual web corpora by checking if the predicted writing system matches the admissible scripts for the corpus's assigned language. It also measures script representation coverage in multilingual LLM tokenizers. Use when the user wants to benchmark on Multilingual C4 (mC4), OSCAR 22.01, or asks about evaluating this task. Reports ACC.
    3 repo stars
  45. ▌
    Seismic Wavefield Ctf Eval · qhjqhj00
    Evaluates machine learning models for seismic wavefield forecasting, reconstruction, and generalization under realistic constraints like noise and limited data. It probes model robustness and dynamic learning by comparing performance across multiple tasks against naive baselines. Use when the user wants to benchmark on global wavefields, DAS, synthetic 3D crustal wavefields, or asks about evaluating this task. Reports multi-metric scoring.
    3 repo stars
  46. ▌
    Sheet Music Benchmark Eval · qhjqhj00
    Evaluates end-to-end optical music recognition (OMR) systems on their ability to transcribe scanned sheet music images into standardized **kern musical notation. It probes layout analysis, staff/page-level transcription accuracy, and robustness across diverse musical textures such as monophony, pianoform, and quartets. Use when the user wants to benchmark on Sheet Music Benchmark (SMB), or asks about evaluating this task. Reports OMR-NED.
    3 repo stars
  47. ▌
    Silicone Prep Anomaly Eval · qhjqhj00
    Evaluates multimodal vision-language models' ability to detect context-dependent visual anomalies in robotic scientific laboratory workflows using first-person imagery and stage-specific textual prompts. Use when the user wants to benchmark on Silicone Preparation Workflow, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  48. ▌
    Singer Identification Eval · qhjqhj00
    Evaluates the ability of audio embedding and classification models to correctly identify the vocalist performing a song track. It specifically probes robustness to synthetic/deepfake voices and generalization across different music datasets and genre contexts. Use when the user wants to benchmark on Train, Validation, Closed, FMA, MTG, Cloned, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  49. ▌
    Sketch Less Retrieval Eval · qhjqhj00
    Evaluates early retrieval performance using partial or low-quality sketches combined with text, measuring how quickly the correct target image appears in the ranked list as the sketch is drawn. It probes robustness to incomplete visual inputs and multimodal fusion. Use when the user wants to benchmark on FS2K-SDE1, FS2K-SDE2, User-SDE, or asks about evaluating this task. Reports m@A, m@B.
    3 repo stars
  50. ▌
    Socialcounterfactuals Eval · qhjqhj00
    Probes intersectional social bias in Large Vision-Language Models by measuring how model outputs vary when only perceived race, gender, or physical attributes change in counterfactual images. It specifically evaluates toxicity, stereotypical language, and competency ratings across different demographic groups. Use when the user wants to benchmark on SocialCounterfactuals, or asks about evaluating this task. Reports MaxToxicity.
    3 repo stars
  51. ▌
    Spatial Reasoning Vqa Eval · qhjqhj00
    Evaluates a vision-language model's ability to perform spatial reasoning tasks, including relative positioning, counting, size comparison, and cross-dataset generalization. It probes whether models learn transferable spatial concepts rather than memorizing dataset-specific patterns or visual artifacts. Use when the user wants to benchmark on GRAID-BDD, GRAID-NuImages, BLINK, A-OKVQA, NaturalBench, RealWorldQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  52. ▌
    Spatiotemporal Kmeans Eval · qhjqhj00
    Evaluates clustering algorithms on their ability to track moving and static clusters in collective animal behavior data across space and time. It probes robustness in low-data regimes and the capacity to produce stable, interpretable cluster trajectories without relying on ground-truth labels for hyperparameter tuning. Use when the user wants to benchmark on Cakmak et al. Spatiotemporal Benchmark, or asks about evaluating this task. Reports total AMI.
    3 repo stars
  53. ▌
    Speech Enhancement Ms Eval · qhjqhj00
    Evaluates semi-supervised speech enhancement algorithms by measuring how well they recover clean speech from noisy mixtures across varying SNRs and noise types. It probes both perceptual quality and speech intelligibility preservation. Use when the user wants to benchmark on IEEE Speech + Environmental/Industrial Noise Database, or asks about evaluating this task. Reports HASQI.
    3 repo stars
  54. ▌
    Spinnaker2 Benchmarks Eval · qhjqhj00
    Evaluates the energy efficiency and computational capability of the SpiNNaker2 processing element architecture across a suite of neuromorphic and deep learning workloads, including classical spiking neural networks, hybrid SNN/DNN frameworks, and standard DNN layers. Use when the user wants to benchmark on SpiNNaker2 Benchmark Suite, or asks about evaluating this task. Reports energy_efficiency.
    3 repo stars
  55. ▌
    Sqi Separation Margin Eval · qhjqhj00
    Evaluates the ability of signal quality indices (SQIs) to predict downstream task performance on medical time series. It measures how well an SQI correlates with and separates high-quality from low-quality signal segments for specific tasks like R-peak detection and atrial fibrillation classification. Use when the user wants to benchmark on Glasgow University database (GUDb), MIT-BIH Atrial Fibrillation Database (MIT-BIH AF), Deepbeat test subset, or asks about evaluating this task. Reports optimal separation margin ($\Delta^*$).
    3 repo stars
  56. ▌
    SQL Hadoop Comparison Eval · qhjqhj00
    This evaluation compares the interactive analytics performance of four SQL-on-Hadoop systems (Impala, Drill, Spark SQL, Phoenix) by measuring query response times and resource utilization. It characterizes how each system's optimizer and execution engine handle join orders, operator selection, and data scanning across different storage formats and scaling configurations. Use when the user wants to benchmark on Unspecified SQL workloads (text/parquet), or asks about evaluating this task. Reports query_rt.
    3 repo stars
  57. ▌
    Subgroup Benchmarking Eval · qhjqhj00
    Evaluates the precision of statistical estimators (Empirical Bayes, Synthetic Regression, Direct Training) for estimating model performance on data subgroups with limited observations. It probes how well these methods reduce mean squared error and produce reliable confidence intervals when benchmarking LLMs, vision models, and tabular classifiers on niche tasks. Use when the user wants to benchmark on LLM MC QA tasks, Computer Vision tasks (LAION CLIP benchmark), COCO Captions, Tabular Fairness Datasets (ACS, COMPAS, Student), or asks about evaluating this task. Reports MSE.
    3 repo stars
  58. ▌
    Summarization Metrics Eval · qhjqhj00
    Evaluates the correlation between automatic summarization metrics and LLM-as-a-Judge models against human judgments across five quality criteria (coherence, consistency, fluency, relevance, 5W1H) in Spanish and Basque. Use when the user wants to benchmark on BASSE, or asks about evaluating this task. Reports Spearman's $ ho$.
    3 repo stars
  59. ▌
    Svg Sophia Refinement Eval · qhjqhj00
    Evaluates a model's ability to refine and correct imperfect SVG code, measuring structural accuracy, visual fidelity, and code efficiency. Use when the user wants to benchmark on SVG-Sophia Code Refinement Benchmark, or asks about evaluating this task. Reports SR.
    3 repo stars
  60. ▌
    Swimba Standard Bench Eval · qhjqhj00
    Evaluates language understanding, reasoning, and knowledge recall capabilities of a Switch Mamba model on standard multiple-choice benchmarks. It also measures inference efficiency (throughput, latency, FLOPs) to assess the computational trade-offs of the parameter-space mixture-of-experts design. Use when the user wants to benchmark on BoolQ, OpenBookQA, RTE, MMLU, PIQA, WinoGrande, HellaSwag, ARC-Challenge, ARC-Easy, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  61. ▌
    Tabular QA Confidence Eval · qhjqhj00
    Evaluates the calibration and reliability of confidence scores produced by LLMs when answering questions over tabular data. It probes how well predicted confidence aligns with actual accuracy across different elicitation methods and dataset complexities. Use when the user wants to benchmark on WikiTableQuestions, TableBench, or asks about evaluating this task. Reports smooth ECE.
    3 repo stars
  62. ▌
    Taobao Ctr Prediction Eval · qhjqhj00
    This benchmark evaluates the ability of machine learning models to predict click-through rates (CTR) for advertisements on a large-scale e-commerce platform. It specifically probes how well models capture static user-ad interactions versus dynamic, temporal user behavior sequences to forecast future clicks. Use when the user wants to benchmark on Alibaba's Taobao Advertising Dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  63. ▌
    Task Graph Scheduling Eval · qhjqhj00
    Evaluates task graph scheduling algorithms by measuring their makespan on standard and adversarially modified network topologies. It probes how well algorithms minimize execution time across heterogeneous datasets and reveals performance reversals under minor structural changes. Use when the user wants to benchmark on Parallel Chains (and 15 other SAGA framework datasets), or asks about evaluating this task. Reports makespan.
    3 repo stars
  64. ▌
    Telemedicine Feedback Eval · qhjqhj00
    Predicts whether a patient will give positive feedback (thumbs-up) for a doctor's response in a Romanian telemedicine platform. It probes the model's ability to leverage clinical communication features, patient/doctor history, and metadata to forecast user satisfaction. Use when the user wants to benchmark on Romanian Telemedicine Platform Dataset, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  65. ▌
    Test Time Scaling Vlm Eval · qhjqhj00
    Evaluates the impact of test-time scaling (TTS) inference strategies on Vision-Language Models across multimodal reasoning and perception tasks. It measures how techniques like Chain-of-Thought, Best-of-N, Self-Consistency, and Self-Refinement improve or degrade performance on open-source versus closed-source models. Use when the user wants to benchmark on MathVista, MMMU, MMBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  66. ▌
    Time Series Benchmark Eval · qhjqhj00
    Evaluates zero-shot forecasting performance of time series foundation models across diverse datasets and horizons. Probes model capability to capture structural temporal patterns (trend, seasonality, stationarity, complexity) and generalizes to unseen data without leakage. Use when the user wants to benchmark on TIME Benchmark, or asks about evaluating this task. Reports MASE, CRPS.
    3 repo stars
  67. ▌
    Topic Trend Detection Eval · qhjqhj00
    Measures the change in sentiment trend towards a specific topic over time or across datasets, requiring temporal or comparative analysis. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports avgDiff.
    3 repo stars
  68. ▌
    Toxicity Perspectives Eval · qhjqhj00
    Evaluates how well automated toxicity classifiers align with diverse human perceptions of harmful content, specifically measuring how demographic background and personal harassment experiences influence toxicity judgments. Use when the user wants to benchmark on Toxicity Perspectives Dataset, or asks about evaluating this task. Reports interrater agreement (Cohen's kappa).
    3 repo stars
  69. ▌
    Trajectory Generation Eval · qhjqhj00
    Evaluates the statistical fidelity and practical utility of synthetic human trajectory generation models by measuring how well generated trajectories perform on downstream mobility tasks compared to real trajectories. It probes whether synthetic data can replace real data without performance degradation across recommendation, prediction, labeling, and simulation tasks. Use when the user wants to benchmark on Foursquare Tokyo (TKY), Foursquare Istanbul (IST), Foursquare New York City (NYC), or asks about evaluating this task. Reports MAPE.
    3 repo stars
  70. ▌
    Transport Forecasting Eval · qhjqhj00
    Evaluates spatio-temporal forecasting models on traffic speed, volume, and bike flow prediction tasks. It specifically probes whether simple baselines that account for weekly stationarity (historical average plus linear regression on residuals) can match or outperform complex deep learning architectures across diverse transport datasets. Use when the user wants to benchmark on PeMSD7(M), Urban1, NYC Citi Bike, PeMSD4, SZ-taxi, METR-LA, PEMS-BAY, NYC Bike in- and out-flows, Seattle traffic speeds, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  71. ▌
    Trec2022 Fair Ranking Eval · qhjqhj00
    Evaluates retrieval systems on balancing topical relevance with intersectional fairness in Wikipedia article rankings. It probes static single-query ranking for coordinators and dynamic multi-query ranking for editors under fairness constraints across demographic attributes. Use when the user wants to benchmark on TREC 2022 Fair Ranking Track, or asks about evaluating this task. Reports M1, EE-L.
    3 repo stars
  72. ▌
    Tte Clutter Filtering Eval · qhjqhj00
    This protocol evaluates a 3D convolutional auto-encoder for removing reverberation artifacts (clutter) from transthoracic echocardiographic (TTE) sequences. It measures how well the network preserves cardiac structures while suppressing simulated artifacts, using synthetic data with known ground truth for training and validation, and normal in-vivo sequences for testing. Use when the user wants to benchmark on Synthetic TTE sequences, or asks about evaluating this task. Reports reconstruction loss ($L_{rec}$).
    3 repo stars
  73. ▌
    Tweac Agent Selection Eval · qhjqhj00
    This evaluation probes a model's ability to correctly route natural language questions to the most appropriate domain-specific QA agent from a large, heterogeneous pool. It measures both sample efficiency (performance with few training examples per agent) and scalability (maintaining accuracy as the number of candidate agents grows to hundreds). Use when the user wants to benchmark on QA-Tasks, Many-Agents, or asks about evaluating this task. Reports Accuracy@1.
    3 repo stars
  74. ▌
    Uav Wildlife Tracking Eval · qhjqhj00
    This evaluation probes an autonomous UAV navigation model's ability to track wildlife by predicting flight commands that match expert pilot behavior. It measures how well the model maintains optimal camera framing and altitude for behavioral video collection. Use when the user wants to benchmark on KABR, or asks about evaluating this task. Reports % of actions matching original flight.
    3 repo stars
  75. ▌
    Un Corpus Translation Eval · qhjqhj00
    Evaluates zero-shot and supervised machine translation quality across multiple language pairs using a multilingual encoder-decoder architecture. It probes the model's ability to translate between unseen language pairs (e.g., Spanish-French) using only monolingual data and reinforcement learning, without parallel training data for the target pair. Use when the user wants to benchmark on United Nations Parallel Corpus (UN corpus), or asks about evaluating this task. Reports BLEU.
    3 repo stars
  76. ▌
    Universalimagequalityindex · qhjqhj00
    Compute the UniversalImageQualityIndex metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute UniversalImageQualityIndex, or asks how to score with UniversalImageQualityIndex.
    3 repo stars
  77. ▌
    Unseen Object 6d Pose Eval · qhjqhj00
    Evaluates a model's ability to estimate the 6D pose (rotation and translation) of novel, unseen 3D objects in real-world scenes without retraining, using only their mesh models and partial RGBD inputs. It specifically probes robustness to pose ambiguity, partial observability, and real-world noise. Use when the user wants to benchmark on GraspNet-1Billion, YCB-Video, or asks about evaluating this task. Reports IADD.
    3 repo stars
  78. ▌
    Vae Malware Detection Eval · qhjqhj00
    Evaluates the effectiveness of Variational Autoencoder (VAE)-derived latent space features for malware classification using traditional machine learning models. It probes robustness to data partitioning, random seed initialization, and computational efficiency without hyperparameter tuning. Use when the user wants to benchmark on EMBER, BODMAS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  79. ▌
    Vga Gui Comprehension Eval · qhjqhj00
    Evaluates a vision-language model's ability to understand graphical user interfaces (GUIs) and answer user questions based on visual content. It specifically probes the model's capacity to avoid hallucinations by grounding responses in actual GUI elements rather than relying solely on textual priors. Use when the user wants to benchmark on GUI Comprehension Bench, or asks about evaluating this task. Reports GPT evaluation score.
    3 repo stars
  80. ▌
    Virus Host Prediction Eval · qhjqhj00
    Evaluates computational tools and genomic features for predicting prokaryotic virus-host interactions. It probes the ability of models to correctly link viral sequences to their host taxa using either pairwise link prediction or taxonomic classification formulations. Use when the user wants to benchmark on RefSeq-VHDB, MetaHiC-VHDB, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  81. ▌
    Visual Text Grounding Eval · qhjqhj00
    Evaluates multimodal large language models' ability to perform precise spatial reasoning and visual text grounding in document images. It tests whether models can generate accurate bounding boxes that support their textual answers, both from scratch (OCR-free) and when provided with OCR text (OCR-based), while also measuring their instruction-following capability. Use when the user wants to benchmark on ChartQA, DocVQA, InfographicsVQA, TRINS, or asks about evaluating this task. Reports IoU.
    3 repo stars
  82. ▌
    Voice Morph Threshold Eval · qhjqhj00
    This evaluation probes auditory self-recognition boundaries by measuring how much AI voice morphing a participant can tolerate before they stop recognizing their own voice. It assesses perceptual thresholds, decision latency, and the influence of acoustic embedding distances and demographic factors on voice identity perception. Use when the user wants to benchmark on VoiceMorph Experimental Dataset, or asks about evaluating this task. Reports lowess_T.
    3 repo stars
  83. ▌
    Weakly Supervised Ner Eval · qhjqhj00
    This evaluation probes a model's ability to perform named entity recognition under weak supervision, where training labels are noisy and derived from multiple distant supervision sources. It measures how well the model can denoise these labels and generalize entity patterns across general, biomedical, and review domains. Use when the user wants to benchmark on CoNLL 2003, LaptopReview, NCBI-Disease, BC5CDR, or asks about evaluating this task. Reports entity-level F1.
    3 repo stars
  84. ▌
    Weather Robustness Od Eval · qhjqhj00
    Evaluates object detection robustness to adverse weather by measuring performance degradation when models trained on clear-weather datasets are tested on weather-corrupted images. It specifically quantifies dataset bias by comparing baseline in-distribution performance against out-of-distribution performance on the DAWN dataset. Use when the user wants to benchmark on DAWN, Pascal VOC 2012, Microsoft COCO 2017, or asks about evaluating this task. Reports mAP.
    3 repo stars
  85. ▌
    Wide Deep Recommender Eval · qhjqhj00
    Evaluates a hybrid recommender model's ability to balance memorization of frequent user-item interactions with generalization to unseen combinations for ranking candidate apps. The protocol measures predictive accuracy on a static holdout set and business impact via live A/B testing on user acquisition rates. Use when the user wants to benchmark on Google Play App Store (internal), or asks about evaluating this task. Reports Online Acquisition Gain.
    3 repo stars
  86. ▌
    Wmt22 Ted Translation Eval · qhjqhj00
    Evaluates machine translation models on English-to-multiple-target-language pairs using standard development sets. It measures translation quality via BLEU scores to compare bilingual versus multilingual decoder representations and capacity. Use when the user wants to benchmark on WMT22 General Machine Translation, Multitarget TED talks, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  87. ▌
    Xu1998hz Sescore German Mt · qhjqhj00
    Compute xu1998hz/sescore_german_mt via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of xu1998hz/sescore_german_mt.
    3 repo stars
  88. ▌
    Yuyijiong Quad Match Score · qhjqhj00
    Compute yuyijiong/quad_match_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of yuyijiong/quad_match_score.
    3 repo stars
  89. ▌
    Zero Shot Commonsense Eval · qhjqhj00
    Evaluates zero-shot commonsense reasoning capabilities of language models using multiple-choice questions. It specifically probes how prompt engineering and probability calibration strategies affect accuracy across different model sizes and architectures. Use when the user wants to benchmark on CommonsenseQA, COPA, OpenBookQA, PIQA, Social IQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  90. ▌
    3d Scene Understanding Eval · qhjqhj00
    Evaluates a 3D vision-language model's ability to perform visual grounding, dense captioning, and situated question answering on indoor RGB-D scenes. It probes the model's capacity for precise object referencing, spatial reasoning, and open-ended language generation conditioned on 3D scene context. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports Acc@0.5.
    3 repo stars
  91. ▌
    Ad Personalization Ips Eval · qhjqhj00
    Evaluates the predictive accuracy and decision-making value of ad targeting policies using different information sets (contextual, geographical, behavioral). It specifically tests whether geographical and behavioral data act as complements or substitutes in improving click-through rates, accounting for user exposure history. Use when the user wants to benchmark on Real-world Ad Impression Dataset, or asks about evaluating this task. Reports IPS policy value.
    3 repo stars
  92. ▌
    Adapter Fedllm Privacy Eval · qhjqhj00
    Evaluates the privacy vulnerability of adapter-based federated large language models against gradient inversion attacks. It measures how accurately an adversary can reconstruct private training text from shared adapter gradients under varying batch sizes, model architectures, and defensive mechanisms. Use when the user wants to benchmark on CoLA, SST, Rotten Tomatoes, or asks about evaluating this task. Reports ROUGE-1.
    3 repo stars
  93. ▌
    Adaptive Query Routing Eval · qhjqhj00
    Evaluates retrieval and answer generation methods across structured financial, legal, and medical documents. It probes how well different architectures handle varying query complexities, cross-references, and domain-specific structural requirements. Use when the user wants to benchmark on Controlled Multi-Domain Corpus, FinanceBench, or asks about evaluating this task. Reports Quality.
    3 repo stars
  94. ▌
    Adversarial Robustness Eval · qhjqhj00
    Evaluates the robustness of image classification models against adversarial perturbations and natural distribution shifts. It measures how well a model maintains prediction accuracy on clean data while recovering performance on out-of-distribution or adversarially attacked inputs. Use when the user wants to benchmark on MNIST, CIFAR10, ImageNet, or asks about evaluating this task. Reports Relative Robustness (RR).
    3 repo stars
  95. ▌
    AI Face Fairness Bench Eval · qhjqhj00
    Evaluates the fairness and utility of AI-generated face detectors across demographic attributes (skin tone, gender, age) and intersectional groups. It measures how well detectors distinguish real vs. AI-generated faces while ensuring equitable performance across demographic subgroups. Use when the user wants to benchmark on AI-Face, or asks about evaluating this task. Reports $F_{MEO}$.
    3 repo stars
  96. ▌
    Alpaca Eval Lc Winrate Eval · qhjqhj00
    Evaluates the alignment quality of language models by measuring their win rate against a baseline on the AlpacaEval benchmark. It specifically uses length-controlled (LC) win rates to mitigate the known bias toward longer model outputs in standard auto-annotator evaluations. Use when the user wants to benchmark on alpaca_eval, or asks about evaluating this task. Reports AlpacaEval length-controlled (LC) win rate.
    3 repo stars
  97. ▌
    Appearance Free Action Eval · qhjqhj00
    Evaluates zero-shot generalization of action recognition models to appearance-free videos generated by warping noise or random dots with optical flow, testing reliance on motion cues over static shape and texture. Use when the user wants to benchmark on UCF5, AFD5, AFF5, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  98. ▌
    Artifact Understanding Eval · qhjqhj00
    Evaluates vision-language models on their ability to detect, spatially localize, and explain visual artifacts in AI-generated images. It probes the model's capacity for fine-grained visual reasoning and artifact-aware grounding beyond standard natural image understanding. Use when the user wants to benchmark on ArtiBench, LOKI, or asks about evaluating this task. Reports accuracy, mIoU, ROUGE.
    3 repo stars
  99. ▌
    Asr Clinical Continual Eval · qhjqhj00
    This evaluation probes an ASR model's ability to continuously adapt to noisy, rural clinical telephony speech while retaining its baseline performance on standard general-domain speech. It specifically measures the trade-off between target-domain transcription accuracy and catastrophic forgetting of pre-trained linguistic knowledge. Use when the user wants to benchmark on Gram Vaani, Kathbath, or asks about evaluating this task. Reports WER.
    3 repo stars
  100. ▌
    Audio Compositionality Eval · qhjqhj00
    Probes whether audio encoders preserve algebraic consistency when identical sources are added to different base scenes (A-COAT), and whether representations can be accurately reconstructed from discrete attribute-level primitives like timbre, pitch, rate, and amplitude (A-TRE). Use when the user wants to benchmark on Synthetic Audio Scenes, or asks about evaluating this task. Reports A-COAT.
    3 repo stars