all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 23 of 76

  1. ▌
    Coedit Text Editing Eval · qhjqhj00
    Evaluates a model's ability to follow task-specific and composite text editing instructions across diverse tasks like grammar correction, simplification, coherence, style transfer, and paraphrasing. It also assesses generalization to unseen instructions and human-perceived writing efficiency. Use when the user wants to benchmark on JFLEG, TurkCorpus, ASSET, ITER (Coherence split), DISCOFUSE, ITER (Iterative text revision), GYAFC, WNC, MRPC, STS, QQP, or asks about evaluating this task. Reports evaluation metrics.
    3 repo stars
  2. ▌
    Cold Offensive Rate Eval · qhjqhj00
    This benchmark probes the safety and bias of Chinese generative language models by measuring how frequently they produce offensive content when prompted with various inputs, including offensive, non-offensive, and anti-bias contexts. Use when the user wants to benchmark on COLDataset, or asks about evaluating this task. Reports offensive rate.
    3 repo stars
  3. ▌
    Community Forensics Eval · qhjqhj00
    This evaluation protocol probes the ability of fake image detectors to generalize across a wide variety of generative models and architectures. It measures how well classifiers trained on diverse synthetic data can distinguish real from generated images in both in-distribution and out-of-distribution settings. Use when the user wants to benchmark on Wang et al. [129], Ojha et al. [90], Synthbuster [7], GenImage [137], Community Forensics (Ours), or asks about evaluating this task. Reports mAP.
    3 repo stars
  4. ▌
    Confready Checklist Eval · qhjqhj00
    This benchmark evaluates a model's ability to accurately answer conference submission checklist questions based on manuscript content. It specifically probes long-form document understanding, retrieval-augmented generation (RAG) effectiveness, and the model's capacity to reflect on ethical considerations, reproducibility, and societal impacts. Use when the user wants to benchmark on ConfReady Evaluation Set, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  5. ▌
    Contextiq Retrieval Eval · qhjqhj00
    Evaluates zero-shot text-to-video retrieval for contextual advertising. It probes a model's ability to rank relevant video content based on natural language queries using multimodal signals (vision, audio, captions, metadata). Use when the user wants to benchmark on Val-1, Val-2, or asks about evaluating this task. Reports P@K.
    3 repo stars
  6. ▌
    Cosmos Drive Dreams Eval · qhjqhj00
    Evaluates the effectiveness of a synthetic driving data generation pipeline by measuring performance gains in downstream autonomous driving perception tasks, including 3D lane detection, 3D object detection, and LiDAR-based detection, particularly under challenging conditions like extreme weather and nighttime. Use when the user wants to benchmark on Waymo Open Dataset, RDS-HQ, RDS-HQ-HL, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  7. ▌
    Crowdsensing Id Dfl Eval · qhjqhj00
    Evaluates the capability of decentralized federated learning (DFL) models to detect malware and classify benign states in IoT crowdsensing environments. It probes robustness under varying node counts, peer-to-peer network topologies, and data heterogeneity (IID vs. non-IID Dirichlet splits). Use when the user wants to benchmark on Crowdsensing Intrusion Detection Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  8. ▌
    Ctr Prediction Dhan Eval · qhjqhj00
    Evaluates a model's ability to predict the next item a user will click or review based on their historical interaction sequence. It probes hierarchical user interest modeling across multiple product dimensions and abstraction levels in recommendation systems. Use when the user wants to benchmark on Amazon Review (Six-Category, Kindle Shop, Electronics), or asks about evaluating this task. Reports AUC.
    3 repo stars
  9. ▌
    Curriculum Word2vec Eval · qhjqhj00
    Evaluates how the ordering of training data (curriculum) affects the quality of task-specific word embeddings. It probes whether optimized data sequencing improves downstream performance in sentiment analysis, NER, POS tagging, and parsing compared to random or heuristic orderings. Use when the user wants to benchmark on Wikipedia Paragraph Corpus, or asks about evaluating this task. Reports test results.
    3 repo stars
  10. ▌
    Cvc Value Alignment Eval · qhjqhj00
    Evaluates how well large language models align with culturally grounded Chinese value rules compared to Western benchmarks. It probes moral reasoning, preference alignment, and boundary separation across six sensitive themes like drugs, firearms, politics, and suicide. Use when the user wants to benchmark on CVC, or asks about evaluating this task. Reports preference.
    3 repo stars
  11. ▌
    Cyberseceval3 Human Eval · qhjqhj00
    This evaluation probes the impact of LLM assistance on human cybersecurity practitioners' ability to execute novel cyberattack challenges. It measures objective performance metrics (phase completion rates and time) and subjective perception (sentiment/mental effort) across inexperienced and highly skilled cohorts, comparing LLM-assisted versus unassisted conditions. Use when the user wants to benchmark on CYBERSECEVAL 3 Challenge Set, or asks about evaluating this task. Reports phase completion time.
    3 repo stars
  12. ▌
    Dharmaocr Benchmark Eval · qhjqhj00
    Evaluates structured OCR extraction fidelity and text degeneration rates on printed, handwritten, and legal documents. Measures how well models adhere to JSON schemas while minimizing pathological generation loops. Use when the user wants to benchmark on DharmaOCR-Benchmark, or asks about evaluating this task. Reports Score.
    3 repo stars
  13. ▌
    Dialectalarabicmmlu Eval · qhjqhj00
    Evaluates large language models' ability to understand and reason across multiple Arabic dialects and standard Arabic across diverse academic and professional domains. It measures dialectal generalization and sensitivity to linguistic context by comparing performance under default, dialect-conditioned, and dialect-identification prompts. Use when the user wants to benchmark on DialectalArabicMMLU, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  14. ▌
    Diffaware Ctxtaware Eval · qhjqhj00
    Evaluates whether LLMs recognize meaningful demographic group differences (Difference Awareness) and understand when differential treatment is contextually appropriate (Contextual Awareness), challenging the standard 'color-blind' fairness paradigm. Use when the user wants to benchmark on DiffAware and CtxtAware Benchmark Suite, or asks about evaluating this task. Reports win rate.
    3 repo stars
  15. ▌
    Dinids Cross Domain Eval · qhjqhj00
    This benchmark evaluates the robustness of network intrusion detection models against distribution shifts between different network environments. It specifically probes cross-domain generalization by training on one NetFlow dataset and testing on another, measuring how well domain-invariant feature extraction mitigates performance degradation when facing unseen attack distributions. Use when the user wants to benchmark on NFv2-UNSW-NB15, NFv2-CIC-2018, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  16. ▌
    Dmcontrol Metaworld Eval · qhjqhj00
    Evaluates sample efficiency, asymptotic performance, and generalization of reinforcement learning agents on high-dimensional continuous control tasks with varying observation modalities (state, pixels, multi-modal) and reward structures (dense, sparse, goal-conditioned). Use when the user wants to benchmark on DeepMind Control Suite (DMControl), Meta-World v2, or asks about evaluating this task. Reports Cumulative Episode Return.
    3 repo stars
  17. ▌
    Dotkaio Competition Math · qhjqhj00
    Compute dotkaio/competition_math via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of dotkaio/competition_math.
    3 repo stars
  18. ▌
    Dpo Ppo Multi Bench Eval · qhjqhj00
    This protocol evaluates language models across factual knowledge, mathematical reasoning, instruction following, code generation, truthfulness, and safety/refusal capabilities. It uses standardized benchmarks to measure how preference optimization methods and data quality impact model performance. Use when the user wants to benchmark on MMLU, GSM8k, Big Bench Hard, TruthfulQA, AlpacaEval, IFEval, HumanEval+, MBPP+, ToxiGen, XSTest, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  19. ▌
    Dvcs Cff Extraction Eval · qhjqhj00
    This benchmark evaluates machine learning models' ability to extract Compton Form Factors (CFFs) from deeply virtual Compton scattering (DVCS) cross-section data. It specifically probes how well models adhere to quantum chromodynamics (QCD) constraints, generalize across kinematic regions, and accurately quantify both aleatoric and epistemic uncertainties during the extraction process. Use when the user wants to benchmark on DVCS unpolarized proton target data, or asks about evaluating this task. Reports predictive uncertainty.
    3 repo stars
  20. ▌
    Dvfs Latency Energy Eval · qhjqhj00
    Evaluates the accuracy of a data-driven DVFS-aware latency model for DNN inference on GPUs against a traditional FLOPs-based benchmark. It probes the model's ability to predict real-world inference time and energy consumption under varying frequency settings, deadlines, and cooperative offloading scenarios. Use when the user wants to benchmark on CIFAR10, or asks about evaluating this task. Reports inference time (ms).
    3 repo stars
  21. ▌
    Dynamic Nerf Soccer Eval · qhjqhj00
    Evaluates the ability of dynamic NeRF models to perform photorealistic novel view synthesis in large-scale, dynamic sports environments. It probes spatiotemporal modeling capabilities, specifically how well models handle fast-moving small objects (like a soccer ball) and scale variations across different camera configurations. Use when the user wants to benchmark on Synthetic Soccer Scenes (Multi-Camera), or asks about evaluating this task. Reports PSNR.
    3 repo stars
  22. ▌
    Ecg Language Models Eval · qhjqhj00
    Evaluates the ability of encoder-free ECG-language models to process raw ECG signals alongside textual queries for medical question answering and instruction following. Probes whether models genuinely leverage physiological ECG data or rely on language priors and benchmark artifacts. Use when the user wants to benchmark on PTB-XL ECG-QA, PULSE ECG-Bench, ECG-Chat Instruct, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  23. ▌
    Edge Case Detection Eval · qhjqhj00
    Tests the model's ability to classify whether a respondent's message represents an edge case that falls outside the scope of existing coordination policies and requires user escalation. Use when the user wants to benchmark on Edge Case Detection Test Suite, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  24. ▌
    Effective Dimensionality · qhjqhj00
    Effective Dimensionality (ED) quantifies the number of independent signals or latent axes captured by a benchmark, measuring how much redundancy exists across its tasks. It probes whether a benchmark's claimed breadth actually reflects diverse evaluation dimensions or merely correlated task performance. Use when the user has predictions and gold and needs to compute Effective Dimensionality (ED).
    3 repo stars
  25. ▌
    Ege Math Assessment Eval · qhjqhj00
    This benchmark evaluates vision-language models' ability to assess handwritten mathematical solutions against a standardized educational rubric. It probes the models' capacity for error diagnosis, step-by-step reasoning alignment, and accurate grade assignment under varying levels of contextual guidance. Use when the user wants to benchmark on EGE-Math Solutions Assessment Benchmark, or asks about evaluating this task. Reports final_score.
    3 repo stars
  26. ▌
    Embodied Nav Safety Eval · qhjqhj00
    This protocol evaluates the safety and navigation performance of embodied agents against physical and model-based attacks. It measures task completion efficiency, path optimality, and goal satisfaction across diverse simulated and real-world environments. Use when the user wants to benchmark on Li et al. (2023), Kim et al. (2024), Khanna et al. (2024), Yin et al. (2024), Wang et al. (2024b), or asks about evaluating this task. Reports Success weighted by Path Length (SPL).
    3 repo stars
  27. ▌
    Embodiedgpt Control Eval · qhjqhj00
    This evaluation probes the model's ability to generate executable sub-goal plans from visual inputs and translate them into low-level control actions in simulated robotic environments. It specifically tests closed-loop planning and few-shot policy adaptation across standard embodied AI benchmarks. Use when the user wants to benchmark on Franka Kitchen, Meta-World, or asks about evaluating this task. Reports success rate.
    3 repo stars
  28. ▌
    Emotion Recognition Eval · qhjqhj00
    Evaluates vision models' ability to classify emotions in images across multiple affective categories. It also measures the alignment between emotions expressed in text prompts and those visually present in generated images. Use when the user wants to benchmark on EmoSet, or asks about evaluating this task. Reports macro-averaged F1-score.
    3 repo stars
  29. ▌
    Emotional Prompting Eval · qhjqhj00
    Evaluates how emotional stimuli (joy, encouragement, anger, insecurity) and their intensity affect LLM behavior across factual accuracy, sycophancy, and toxicity. It measures the performance delta when emotional prompt add-ons are applied to base prompts. Use when the user wants to benchmark on Anthropic’s SycophancyEval subset, Toxicity dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  30. ▌
    Emotiw2018 Group Er Eval · qhjqhj00
    Evaluates a model's ability to recognize and aggregate facial emotions across multiple individuals in a crowd scene. It probes the system's capacity to capture complex spatial dependencies among overlapping facial expressions to predict the dominant group-level emotion. Use when the user wants to benchmark on EmotiW2018, GECV, or asks about evaluating this task. Reports mean accuracy (mAC).
    3 repo stars
  31. ▌
    Empatheticdialogues Eval · qhjqhj00
    This benchmark evaluates a model's ability to generate or retrieve empathetic, relevant, and fluent responses to emotionally grounded personal stories. It probes how well dialogue systems can acknowledge and react to a speaker's feelings in open-domain conversations. Use when the user wants to benchmark on EmpatheticDialogues, or asks about evaluating this task. Reports Human Empathy/Relevance/Fluency.
    3 repo stars
  32. ▌
    Ept15 Weather Bench Eval · qhjqhj00
    Evaluates the accuracy of AI weather forecasting models against established numerical models and ground-truth observations. It probes the model's ability to predict atmospheric variables (e.g., wind speed, solar radiation) at hourly resolution over 20-day lead times. Use when the user wants to benchmark on ERA5, IFS HRES IC, Weather Stations, or asks about evaluating this task. Reports Skill Score (SS).
    3 repo stars
  33. ▌
    Error Detection Hmc Eval · qhjqhj00
    Evaluates a model's ability to detect classification errors and recover hierarchical multi-label constraints without prior knowledge. It probes the system's capacity to generate interpretable logical rules from failure patterns and improve downstream model consistency. Use when the user wants to benchmark on Military Vehicles, ImageNet50, OpenImage36, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  34. ▌
    Exoplanet Detection Eval · qhjqhj00
    Evaluates high-contrast imaging algorithms' ability to detect injected exoplanet companions in real and simulated ADI sequences. It measures detection sensitivity and specificity across different noise regimes and instrument datasets, comparing performance against standard PCA and deep learning baselines. Use when the user wants to benchmark on EIDC (Exoplanet Imaging Data Challenge) Phase 1, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  35. ▌
    Explained Variance Score · qhjqhj00
    Compute the explained_variance_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute explained_variance_score, or asks how to score with explained_variance_score.
    3 repo stars
  36. ▌
    Fact Checking Arena Eval · qhjqhj00
    Evaluates LLMs on multi-hop fact-checking by measuring claim extraction, evidence retrieval, and justification quality. Uses an arena-style pairwise comparison framework with LLM judges to rank models across multiple reasoning dimensions. Use when the user wants to benchmark on HOVER, FEVERIOUS, or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  37. ▌
    Fair Credit Scoring Eval · qhjqhj00
    This benchmark evaluates the trade-off between algorithmic fairness and financial profitability in credit scoring models. It probes how well various fairness-aware preprocessing, in-processing, and post-processing techniques maintain predictive accuracy and demographic parity while minimizing economic loss for lenders. Use when the user wants to benchmark on german, bene, taiwan, uk, pakdd, gmsc, homecredit, or asks about evaluating this task. Reports Profitability (Profit per EUR).
    3 repo stars
  38. ▌
    Fairness Algorithms Eval · qhjqhj00
    This evaluation probes the trade-off between predictive performance and group fairness across various machine learning pipelines. It systematically compares fairness-unaware baselines against preprocessing and in-training fairness interventions, measuring how well algorithms maintain accuracy while satisfying demographic parity and equalized odds constraints. Use when the user wants to benchmark on Titanic, German, Adult, S-D, S-P, I-D, or asks about evaluating this task. Reports Fair Efficiency (Theta_AUC_DI / Theta_AUC_EO).
    3 repo stars
  39. ▌
    Fairness Comparison Eval · qhjqhj00
    Evaluates the comparative performance of fairness-enhancing machine learning interventions across multiple datasets. It probes how different algorithmic strategies trade off predictive accuracy against a comprehensive set of fairness metrics under standardized preprocessing and fixed train-test splits. Use when the user wants to benchmark on Standardized benchmark datasets (unspecified in excerpt), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  40. ▌
    Fairness Sequential Eval · qhjqhj00
    Evaluates sequential decision policies under simulated historical and measurement bias to measure how accounting for unrealized outcomes affects fairness disparities and cumulative utility. It probes whether uncertainty-aware exploration mitigates selection rate differences and false positive rate parity violations without sacrificing profit. Use when the user wants to benchmark on Synthetic Sequential Simulation, or asks about evaluating this task. Reports selection rate difference.
    3 repo stars
  41. ▌
    Financial Retrieval Eval · qhjqhj00
    Evaluates the ability of text embedding models to retrieve relevant financial document passages given complex, long-form queries. It probes domain-specific retrieval capabilities, including sensitivity to company names, tickers, financial metrics, and date-specific information. Use when the user wants to benchmark on Financial Document Retrieval Dataset, or asks about evaluating this task. Reports Recall@1.
    3 repo stars
  42. ▌
    Forge Manufacturing Eval · qhjqhj00
    Evaluates multimodal large language models on fine-grained manufacturing tasks, including workpiece verification, surface defect inspection, and assembly verification. It probes the models' ability to combine visual grounding with domain-specific knowledge to identify anomalies or classify conditions. Use when the user wants to benchmark on FORGE, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  43. ▌
    Fractal Pretraining Eval · qhjqhj00
    Evaluates the downstream transfer capability of fractal-based pre-trained visual representations on fine-grained image classification and medical image segmentation tasks. It measures how effectively synthetic Iterated Function System (IFS) pre-training captures transferable features compared to training from scratch or using ImageNet/FractalDB pre-training. Use when the user wants to benchmark on CUB-2011, Stanford Cars, Stanford Dogs, FGVC Aircraft, CIFAR-100, GlaS, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  44. ▌
    Frogent Drug Design Eval · qhjqhj00
    Evaluates an end-to-end agentic framework for small-molecule drug design across eight benchmarks spanning the full discovery pipeline, from target identification and knowledge retrieval to virtual screening, interaction profiling, de novo design, and retrosynthetic planning. Use when the user wants to benchmark on Humanity’s Last Exam (HLE), UniProt, Open Targets Platform, ADMETLab 3.0, DAVIS, PLIP, CrossDocked, USPTO-50k, PaRoute, or asks about evaluating this task. Reports score.
    3 repo stars
  45. ▌
    Gedit Imgedit Bench Eval · qhjqhj00
    Evaluates the quality of AI-generated image edits by measuring semantic consistency, perceptual quality, and overall fidelity against source images and edit instructions. Use when the user wants to benchmark on GEdit-Bench, ImgEdit-Bench, or asks about evaluating this task. Reports Overall (GEdit-Bench).
    3 repo stars
  46. ▌
    Gnn Op Runtime Benchmark · qhjqhj00
    Evaluates the runtime and memory efficiency of low-level Graph Neural Network and sparse tensor operations on NVIDIA A100 GPUs. It probes how input sparsity, tensor dimensions, and reduce factors affect computational overhead when operations are pushed to near-full GPU memory capacity. Use when the user has predictions and gold and needs to compute median_runtime.
    3 repo stars
  47. ▌
    Goemotions Transfer Eval · qhjqhj00
    Evaluates cross-domain generalization of emotion classification models by measuring how well a model fine-tuned on GoEmotions transfers to external benchmarks with limited labeled data, compared to training from scratch. Use when the user wants to benchmark on ISEAR, EmoInt, Emotion-Stimulus, or asks about evaluating this task. Reports average F1-score.
    3 repo stars
  48. ▌
    Granular Change Accuracy · qhjqhj00
    Evaluates Dialogue State Tracking (DST) models by measuring performance based on per-turn belief state changes rather than raw slot or turn-level accuracy. It aims to provide partial credit for partially correct predictions and reduce bias from error timing and distribution across dialogue turns. Use when the user has predictions and gold and needs to compute Granular Change Accuracy.
    3 repo stars
  49. ▌
    Graph Gen Benchmark Eval · qhjqhj00
    This benchmark evaluates how effectively graph generative models can produce synthetic graphs that serve as reliable proxies for benchmarking Graph Neural Networks. It measures the fidelity of generated graphs by comparing GNN performance metrics trained on synthetic data against those trained on the original real-world graphs. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, AmazonC, AmazonP, MS CS, MS Physic, or asks about evaluating this task. Reports MSE.
    3 repo stars
  50. ▌
    Greek LLM Benchmark Eval · qhjqhj00
    Evaluates open-source (Llama-70b) and closed-source (GPT-4o mini) LLMs across seven distinct NLP tasks in Modern Greek. It probes capabilities in classification, sequence labeling, text generation, and machine translation to assess model performance in a lesser-resourced language setting. Use when the user wants to benchmark on SemEval-2020 Task 12 (OffensEval-2020 Greek), Greek Native Corpus (GNC), Global Voices Greek MT Corpus, Areios Pagos Legal Summarization Corpus, University Help Desk Intent Classification Dataset, Greek NER Annotated Dataset, Greek Treebank (POS), or asks about evaluating this task. Reports macro-F1, BERTScore F1.
    3 repo stars
  51. ▌
    Gsc Speech Commands Eval · qhjqhj00
    Evaluates keyword spotting models trained on real versus synthetic speech data, measuring how ASR-based filtering of hallucinated synthetic commands affects classification accuracy on the Google Speech Commands dataset. Use when the user wants to benchmark on Google Speech Commands (GSC), or asks about evaluating this task. Reports Accuracy (%).
    3 repo stars
  52. ▌
    Gui Grounding Agent Eval · qhjqhj00
    Evaluates a GUI agent's ability to precisely locate UI elements (grounding) and execute multi-step tasks in real-world desktop/web environments. It probes spatial reasoning, text/icon matching, and long-horizon planning robustness without lookahead. Use when the user wants to benchmark on ScreenSpot-Pro, ScreenSpot-V2, OSWorld-G, OSWorld, WindowsAgentArena, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  53. ▌
    Guide Research Idea Eval · qhjqhj00
    This evaluation probes a system's ability to act as a scientific advisor by predicting whether research hypotheses will be accepted at a top-tier AI conference. It measures alignment with expert peer-review decisions using ranking-based precision and recall metrics on a held-out set of conference submissions. Use when the user wants to benchmark on ICLR 2025 Submissions Test Set, or asks about evaluating this task. Reports Top-30% Precision.
    3 repo stars
  54. ▌
    Hage2000 Code Eval Stdio · qhjqhj00
    Compute hage2000/code_eval_stdio via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of hage2000/code_eval_stdio.
    3 repo stars
  55. ▌
    Halvest Contrastive Eval · qhjqhj00
    Evaluates language models' ability to capture authorial style and stylometric patterns in scholarly text, independent of topical content. It probes whether models can distinguish documents by the same author across different topics and languages using triplet classification and document retrieval tasks. Use when the user wants to benchmark on HALvest-Contrastive, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  56. ▌
    Har Wearable Sensor Eval · qhjqhj00
    Evaluates deep learning architectures (CNNs, LSTMs, DNNs) for frame-by-frame human activity recognition using wearable sensor time-series data. It probes the models' capacity to capture temporal dependencies and generalize across diverse domains (kitchen gestures, lifestyle/exercise, and medical gait analysis) while handling severe class imbalance. Use when the user wants to benchmark on Opportunity, PAMAP2, Daphnet Gait, or asks about evaluating this task. Reports mean f1-score.
    3 repo stars
  57. ▌
    Hate Speech Ordinal Eval · qhjqhj00
    Evaluates deep learning models' ability to predict continuous, interval-scaled hate speech scores from raw text comments. It benchmarks against existing APIs and transformer baselines using cross-validated error and correlation metrics. Use when the user wants to benchmark on Custom hate speech corpus (YouTube, Reddit, Twitter), or asks about evaluating this task. Reports RMSE.
    3 repo stars
  58. ▌
    Hd209458b Retrieval Eval · qhjqhj00
    Evaluates atmospheric retrieval models on exoplanet transmission spectra to constrain chemical abundances, temperature-pressure profiles, and cloud properties. It probes the model's ability to disentangle spectral features across multi-wavelength observations and quantify detection significance of trace gases. Use when the user wants to benchmark on JWST NIRCam transmission spectra, HST WFC3 transmission spectra, HST STIS transmission spectra, or asks about evaluating this task. Reports reduced chi-squared (χ²_red).
    3 repo stars
  59. ▌
    Helmet Long Context Eval · qhjqhj00
    Evaluates a model's ability to retain, process, and reason over extended contexts (8K to 128K tokens) across retrieval-augmented generation (RAG) and long-range question answering (LongQA) tasks. It probes robustness to noise, multi-hop reasoning, and memorization in long-context settings. Use when the user wants to benchmark on HELMET, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  60. ▌
    Hindi LLM Benchmark Eval · qhjqhj00
    Evaluates instruction-following, mathematical reasoning, code/function-calling, and retrieval-augmented generation capabilities of LLMs in Hindi. The benchmark specifically probes the models' ability to handle culturally and linguistically nuanced prompts that go beyond direct English translation. Use when the user wants to benchmark on IFEval-Hi, MT-Bench-Hi, GSM8K-Hi, ChatRAG-Hi, BFCL-Hi, or asks about evaluating this task. Reports score.
    3 repo stars
  61. ▌
    Hirid Icu Benchmark Eval · qhjqhj00
    Evaluates machine learning and deep learning models on predicting clinical outcomes from high-resolution ICU time-series data. It probes capabilities in handling class imbalance, long temporal dependencies, and varying data resolutions across stay-level and online monitoring tasks. Use when the user wants to benchmark on HiRID, or asks about evaluating this task. Reports AUPRC.
    3 repo stars
  62. ▌
    Hlss Histopathology Eval · qhjqhj00
    Evaluates the quality of self-supervised visual representations learned from histopathology images by measuring downstream classification performance at patch, slide, and patient levels. It probes the model's ability to capture hierarchical pathological structures and align them with clinical diagnostic categories. Use when the user wants to benchmark on OpenSRH, TCGA, or asks about evaluating this task. Reports kNN classification accuracy (ACC).
    3 repo stars
  63. ▌
    Image Transcreation Eval · qhjqhj00
    Evaluates multimodal models' ability to culturally adapt images (transcreation) while preserving semantics, layout, and naturalness. Probes cross-cultural visual alignment and context-aware editing capabilities for global audiences. Use when the user wants to benchmark on Image Transcreation Dataset, or asks about evaluating this task. Reports culture-concept.
    3 repo stars
  64. ▌
    Imagenet Top1 Error Eval · qhjqhj00
    Evaluates the top-1 classification accuracy of a model on the ImageNet dataset. It probes the model's ability to correctly classify images into one of 1000 categories under various training conditions, specifically testing the impact of large minibatch sizes and learning rate scaling strategies on optimization and generalization. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports top-1 error (%).
    3 repo stars
  65. ▌
    India Weather Bench Eval · qhjqhj00
    Evaluates data-driven regional weather forecasting models over India under varying boundary conditioning strategies. It probes the ability of architectures to accurately predict multi-variable meteorological fields at high resolution and assesses their robustness during extreme weather events like heatwaves. Use when the user wants to benchmark on IndiaWeatherBench, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  66. ▌
    Instruction Editing Eval · qhjqhj00
    Evaluates instruction-guided image editing models on their ability to modify an input image according to a text prompt. It measures semantic alignment with the prompt and visual fidelity to a ground-truth edit, plus human preference in pairwise comparisons. Use when the user wants to benchmark on MagicBrush (MagBr), ZONE, or asks about evaluating this task. Reports CLIP-T.
    3 repo stars
  67. ▌
    Ist Unbabel 2022 Qe Eval · qhjqhj00
    Evaluates machine translation quality estimation (QE) by predicting human quality scores at the sentence level and identifying error locations at the word level. It also assesses the model's ability to generate faithful explanations for predicted errors. Use when the user wants to benchmark on IST-Unbabel 2022 QE Shared Task, or asks about evaluating this task. Reports Spearman's rank correlation, Matthew's correlation coefficient (MCC), Recall@K (R@K).
    3 repo stars
  68. ▌
    Jointdnn Benchmarks Eval · qhjqhj00
    Evaluates a layer-granular DNN offloading framework by measuring inference latency and energy consumption across standard discriminative, generative, and autoencoder neural network architectures on mobile-cloud setups. Use when the user wants to benchmark on JointDNN Deep Architecture Benchmarks (AlexNet, OverFeat, VGG16, Deep Speech, ResNet, NiN, Chair, Pix2Pix), or asks about evaluating this task. Reports latency.
    3 repo stars
  69. ▌
    Kabr Drone Behavior Eval · qhjqhj00
    Evaluates an automated video classification pipeline for multi-species wildlife behavior monitoring by comparing machine learning predictions against expert manual annotations and traditional ground-based sampling methods. Use when the user wants to benchmark on KABR Drone Behavioral Dataset (custom), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  70. ▌
    Kilian Group Arxiv Score · qhjqhj00
    Compute kilian-group/arxiv_score via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of kilian-group/arxiv_score.
    3 repo stars
  71. ▌
    Kimi K1 5 Benchmark Eval · qhjqhj00
    Evaluates multimodal reasoning, coding, and instruction-following capabilities across text, code, and vision tasks using a standardized suite of academic benchmarks. Use when the user wants to benchmark on MMLU, IF-Eval, CLUEWSC, C-EVAL, HumanEval-Mul, LiveCodeBench, Codeforces, AIME 2024, MATH-500, MMMU, MATH-Vision, MathVista, or asks about evaluating this task. Reports exact-match accuracy (EM).
    3 repo stars
  72. ▌
    Kits21 Segmentation Eval · qhjqhj00
    This benchmark evaluates the ability of deep learning models to perform multi-organ and multi-lesion semantic segmentation on 3D medical imaging data. It specifically probes a model's capacity to accurately delineate kidneys, renal tumors, and renal cysts from corticomedullary-phase CT scans, testing both volumetric overlap and boundary precision. Use when the user wants to benchmark on KiTS21, or asks about evaluating this task. Reports dice.
    3 repo stars
  73. ▌
    Kitti Subt Vo Depth Eval · qhjqhj00
    Evaluates unsupervised monocular visual odometry and depth estimation methods on challenging driving and subterranean environments. Probes the model's ability to predict consistent 6-DoF ego-motion and recover accurate depth maps without ground-truth supervision. Use when the user wants to benchmark on KITTI, DARPA Subterranean Challenge, or asks about evaluating this task. Reports relative translation error ($t_{err}$), relative rotation error ($r_{err}$).
    3 repo stars
  74. ▌
    Klip Postprocessing Eval · qhjqhj00
    Evaluates the accuracy and computational efficiency of PCA-based PSF subtraction algorithms for high-contrast astronomical imaging. Specifically, it measures how well the algorithm mitigates speckle noise while recovering planetary signals compared to a reference implementation. Use when the user wants to benchmark on Beta Pictoris ($\beta$ Pic), HR8799, or asks about evaluating this task. Reports SNR.
    3 repo stars
  75. ▌
    Knowledge Based Vqa Eval · qhjqhj00
    Probes a multimodal LLM's ability to answer visual questions that require external knowledge retrieval. It specifically tests the model's capacity to dynamically decide when to retrieve information and assess the relevance of retrieved documents using self-reflective tokens, without degrading performance on standard visual-only queries. Use when the user wants to benchmark on Encyclopedic-VQA, InfoSeek, or asks about evaluating this task. Reports BERT matching score (BEM), VQA accuracy.
    3 repo stars
  76. ▌
    Llava Onevision 1 5 Eval · qhjqhj00
    This evaluation probes the multimodal reasoning, visual question answering, OCR, and chart understanding capabilities of large multimodal models. It tests the model's ability to process high-resolution images, extract fine-grained text, and perform complex reasoning across diverse visual domains. Use when the user wants to benchmark on MMStar, MMEBench, MME-RealWorld, SeedBench, CV-Bench, RealWorldQA, MathVista, WeMath, MathVision, MMMU, MMMU-Pro, ChartQA, CharXiv, DocVQA, OCRBench, AI2D, InfoVQA, PixmoCount, CountBench, VL-RewardBench, V*, or asks about evaluating this task. Reports accuracy / benchmark-specific score.
    3 repo stars
  77. ▌
    LLM Self Correction Eval · qhjqhj00
    Evaluates the intrinsic self-correction capability of LLMs across safety, reasoning, and vision-language tasks. It measures how iterative self-refinement reduces model uncertainty and improves calibration, toxicity mitigation, and bias reduction. Use when the user wants to benchmark on AdvBench, CommonGen-Hard, BBQ, MMVP, MS-COCO (Visual Grounding Subset), Real Toxicity Prompts, or asks about evaluating this task. Reports toxicity_score.
    3 repo stars
  78. ▌
    Lora Fewshot Aerial Eval · qhjqhj00
    Evaluates parameter-efficient fine-tuning (LoRA) for cross-domain few-shot object detection on aerial imagery. It probes the model's ability to generalize to new domains with limited labeled data while mitigating overfitting. Use when the user wants to benchmark on DOTA, DIOR, or asks about evaluating this task. Reports mAP@0.5.
    3 repo stars
  79. ▌
    Lunara Aesthetic Ii Eval · qhjqhj00
    Evaluates the quality, contextual variation isolation, aesthetic appeal, and identity preservation of the Lunara Aesthetic II image variation dataset compared to other web-scale datasets. Use when the user wants to benchmark on Lunara-II-Variations, Lunara-I, CC3M, LAION-2B-Aesthetic, WIT, or asks about evaluating this task. Reports LAION Aesthetics v2 score.
    3 repo stars
  80. ▌
    Magic Apt Detection Eval · qhjqhj00
    Evaluates the ability of a self-supervised graph representation learning model to detect Advanced Persistent Threats (APTs) in system audit logs. It probes multi-granularity anomaly detection (batched log-level and system entity-level) under a strict unsupervised setting where only benign data is available for training. Use when the user wants to benchmark on StreamSpot, Unicorn Wget, DARPA Engagement 3, or asks about evaluating this task. Reports Precision.
    3 repo stars
  81. ▌
    Malware Adversarial Eval · qhjqhj00
    Evaluates the robustness of deep neural network malware classifiers against adversarial attacks. It probes whether an attacker can successfully misclassify malicious Android applications by adding a limited number of valid features to their manifest files. Use when the user wants to benchmark on DREBIN, or asks about evaluating this task. Reports misclassification rate.
    3 repo stars
  82. ▌
    Mammo Fm Diagnostic Eval · qhjqhj00
    Evaluates a breast-specific foundational model's ability to generalize across in-distribution and out-of-distribution mammographic datasets for zero-shot diagnosis, linear probing, full fine-tuning, and pathology localization. It probes the model's robustness, data efficiency, and representation quality for clinical tasks like cancer detection and risk prediction. Use when the user wants to benchmark on EMBED, VinDr, RSNA, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  83. ▌
    Manifold Robustness Eval · qhjqhj00
    Evaluates the robustness of a dimensionality reduction pipeline (Isomap + Procrustes alignment + TDA clustering) against ambient noise, outliers, and hyperparameter variation. It tests whether the method can consistently recover a low-distortion 2D embedding of a contractible manifold or correctly detect topological failure on non-contractible data. Use when the user wants to benchmark on Swiss roll, Buckyball, or asks about evaluating this task. Reports Persistent homology features ($PH_1$, $PH_0$).
    3 repo stars
  84. ▌
    Maniskill2 Softbody Eval · qhjqhj00
    Evaluates a robot policy's ability to perform long-horizon manipulation tasks involving deformable soft bodies (e.g., clay, noodles, liquid, plasticine). It probes spatial reasoning, contact dynamics, and precise end-effector control under varying initial conditions. Use when the user wants to benchmark on ManiSkill2 Challenge (Soft-body Track), or asks about evaluating this task. Reports Success Metric.
    3 repo stars
  85. ▌
    Materials LLM Probe Eval · qhjqhj00
    Evaluates large language models' ability to retrieve materials science knowledge and predict continuous physical properties. It probes the fundamental asymmetry in LLM behavior between symbolic tasks (classification, link prediction) and numerical regression tasks, assessing how fine-tuning affects accuracy and output consistency across modalities. Use when the user wants to benchmark on MatKG, Crystal System Classification, Bandgap Prediction, Dielectric Constant Prediction, or asks about evaluating this task. Reports RMSE, Top-1 accuracy.
    3 repo stars
  86. ▌
    Mbib Political Bias Eval · qhjqhj00
    Evaluates large language models' ability to detect political bias in media text using in-context learning and chain-of-thought prompting. It probes the model's capacity to distinguish biased from unbiased content across diverse textual chunks without fine-tuning. Use when the user wants to benchmark on Media Bias Identification Benchmark (MBIB), or asks about evaluating this task. Reports Macro-F1.
    3 repo stars
  87. ▌
    Medical AI Security Eval · qhjqhj00
    Probes the vulnerability of medical AI models to jailbreaking and privacy extraction attacks across different clinical specialties. It assesses how models handle synthetic patient data requests for harmful or sensitive information, measuring compliance rates and protected health information (PHI) leakage severity. Use when the user wants to benchmark on Medical AI Security Attack Scenarios, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  88. ▌
    Medical LLM Merging Eval · qhjqhj00
    Evaluates the effectiveness of various model merging techniques for consolidating knowledge in medical large language models. It probes whether merged models can outperform their base and parent models across diverse medical and general reasoning benchmarks. Use when the user wants to benchmark on MedQA, PubMedQA, HellaSwag, MedMCQA, MMLU Professional Medicine, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  89. ▌
    Medical Mllm Safety Eval · qhjqhj00
    Evaluates the safety robustness and medical capability of multimodal large language models against general and medical-specific threats, including cross-modality jailbreak attacks. It measures the trade-off between restoring safety guardrails and preserving domain-specific accuracy. Use when the user wants to benchmark on HarmBench, CATQA, HEx-PHI, MedSafetyBench, CARES, MedSentry, 3D-Tiny-1K, VQA_RAD, MedQA, PubMedQA, SuperGPQA, CMExam, Medbullets, or asks about evaluating this task. Reports Safety Score (1-ASR).
    3 repo stars
  90. ▌
    Medquad Behavioural Eval · qhjqhj00
    Evaluates a medical AI system's clinical reasoning behavior, focusing on uncertainty handling, deferral, and safety rather than raw answer accuracy. It probes the model's ability to maintain clinician-aligned reasoning, avoid speculative completions, and preserve context across diverse medical queries. Use when the user wants to benchmark on MedQuAD benchmark, or asks about evaluating this task. Reports Benchmark Completion Rate.
    3 repo stars
  91. ▌
    Mimic Iii Benchmark Eval · qhjqhj00
    Evaluates clinical risk prediction and intervention forecasting using structured EHR time-series data. Probes a model's ability to handle missingness, temporal gaps, and class imbalance in ICU patient records. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  92. ▌
    Mimic Iv Tu Dataset Eval · qhjqhj00
    Evaluates self-supervised graph representation learning on electronic health records and general graph classification tasks. It probes the model's ability to learn temporal and structural patient representations without task-specific fine-tuning, and its robustness across clinical and non-clinical domains. Use when the user wants to benchmark on MIMIC-IV, TUDataset, or asks about evaluating this task. Reports binary classification accuracy.
    3 repo stars
  93. ▌
    Mlaad Cross Dataset Eval · qhjqhj00
    Evaluates the cross-dataset generalization capability of voice anti-spoofing models. It probes whether models trained on one synthetic audio dataset can accurately detect deepfake or spoofed speech when tested on entirely different datasets, including those with only spoof samples or different languages. Use when the user wants to benchmark on ASVspoof19, ASVspoof21-DF, ASVspoof21-LA, FakeOrReal, InTheWild, MLAAD v1, Voc.v, WaveFake, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  94. ▌
    Mobile Construction Eval · qhjqhj00
    Tests a robot's ability to simultaneously navigate and construct a target structure in 1D/2D/3D grid worlds under partial observability and environmental uncertainty. It evaluates the bi-directional coupling between localization and manipulation planning. Use when the user wants to benchmark on Mobile Construction Benchmark, or asks about evaluating this task. Reports IoU.
    3 repo stars
  95. ▌
    Mobile Dl Inference Eval · qhjqhj00
    Evaluates the effectiveness of parallel deep learning inference strategies across heterogeneous mobile processors (CPU, GPU, DSP) under varying workloads and dynamic system conditions. It probes how operator support, scheduling granularity, and competing processes impact inference latency, resource utilization, and system responsiveness. Use when the user wants to benchmark on Standard DL Models (YOLOv2, VGG-16, PoseNet, FST, RetinaFace, ResNet-18, ResNet-50), or asks about evaluating this task. Reports inference latency (ms).
    3 repo stars
  96. ▌
    Mobile R1 Benchmark Eval · qhjqhj00
    Evaluates a vision-language model's ability to navigate and complete multi-step tasks on mobile GUIs. It measures step-level action correctness, full trajectory success, and robustness to intermediate errors in a simulated Android environment. Use when the user wants to benchmark on Chinese Mobile Agent Benchmark, or asks about evaluating this task. Reports Accuracy (Acc.).
    3 repo stars
  97. ▌
    Mobility Timeseries Eval · qhjqhj00
    Evaluates the accuracy of time series forecasting models on urban mobility data across different prediction horizons. It probes how well traditional, deep learning, and foundation models capture short-term, medium-term, and long-term temporal dependencies in bike-sharing flows. Use when the user wants to benchmark on BikeNYC, BikeVIE, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  98. ▌
    Molecular Structure Eval · qhjqhj00
    Evaluates the accuracy of quantum chemistry methods in predicting equilibrium molecular geometries and vibrational properties against experimental benchmarks. It probes the ability of different basis sets and ansatzes to capture electron correlation and potential energy surface curvature. Use when the user wants to benchmark on Small Molecule Benchmark (H2, LiH, BeH2, H2O), or asks about evaluating this task. Reports vibrational_frequency_error_pct.
    3 repo stars
  99. ▌
    Molvision Benchmark Eval · qhjqhj00
    Evaluates vision-language models on molecular property prediction by combining skeletal structure images with textual prompts. It probes the model's ability to perform binary classification, numerical regression, and textual description generation across diverse chemical properties. Use when the user wants to benchmark on BACE-V, BBBP-V, HIV-V, ClinTox-V, Tox21-V, ESOL-V, LD50-V, QM9-V, PCQM4Mv2-V, ChEBI-V, or asks about evaluating this task. Reports True/False accuracy.
    3 repo stars
  100. ▌
    Moviegraphs Emotion Eval · qhjqhj00
    Predicts multi-label emotions and mental states for movie scenes and individual characters using multimodal inputs (video, dialog, character appearance). It probes long-form video understanding and the ability to integrate visual and linguistic cues for affect recognition. Use when the user wants to benchmark on MovieGraphs, or asks about evaluating this task. Reports mAP.
    3 repo stars