all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 45 of 76

  1. ▌
    Moleculariq Eval · qhjqhj00
    Evaluates large language models' ability to perform symbolic reasoning on molecular graphs, including counting atomic features, indexing substructures, and generating constrained molecular structures. It probes whether models understand chemical topology and composition rather than relying on memorized token patterns or canonical SMILES conventions. Use when the user wants to benchmark on MOLECULARIQ, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  2. ▌
    Monkey Jump Eval · qhjqhj00
    Evaluates the performance and parameter/memory/throughput efficiency of a gradient-free MoE-style PEFT routing mechanism across text, image, and video benchmarks compared to standard and MoE-PEFT baselines. Use when the user wants to benchmark on 47 Benchmarks (Text/Image/Video), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  3. ▌
    Mos Rmbench Eval · qhjqhj00
    Evaluates the ability of speech quality reward models to correctly rank pairs of audio samples based on their Mean Opinion Score (MOS). It probes fine-grained perceptual discrimination and cross-dataset generalization in preference-based audio modeling. Use when the user wants to benchmark on BVCC, NISQA, SingMOS, SOMOS, TMHINT-QI, VMC’23, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  4. ▌
    Mosaic Cdsr Eval · qhjqhj00
    Evaluates a model's ability to perform next-item recommendation in cross-domain sequential settings by decomposing user intent into orthogonal preference components. It probes how well the model leverages shared and domain-specific signals across multiple item categories to predict future interactions. Use when the user wants to benchmark on Amazon Reviews (Movie–Book), Amazon Reviews (Movie–Music), Douban (Movie–Book), or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  5. ▌
    Mtbi Speech Eval · qhjqhj00
    Evaluates the automatic speech recognition accuracy and zero-shot generalization capabilities of Speech Large Language Models. It probes robustness across mathematical reasoning, speaker role inference, and prompt adaptation tasks. Use when the user wants to benchmark on LibriSpeech, GSM8K, Generalization Test Set, or asks about evaluating this task. Reports WER.
    3 repo stars
  6. ▌
    Mteb Eng V2 Eval · qhjqhj00
    Evaluates the semantic discrimination and generalization capabilities of text embedding models across diverse NLP tasks including retrieval, classification, clustering, and semantic similarity. It uses a zero-shot English-only benchmark to measure performance without task-specific fine-tuning. Use when the user wants to benchmark on MTEB(eng, v2), or asks about evaluating this task. Reports average score across all tasks.
    3 repo stars
  7. ▌
    Mteb Subset Eval · qhjqhj00
    Evaluates text embedding models across diverse semantic tasks including retrieval, reranking, clustering, pair classification, classification, and semantic textual similarity to measure the quality of dense vector representations. Use when the user wants to benchmark on MTEB (15-task subset), or asks about evaluating this task. Reports MTEB average score.
    3 repo stars
  8. ▌
    Muharaf Htr Eval · qhjqhj00
    Evaluates handwritten text recognition (HTR) systems on historical Arabic manuscripts. It probes the model's ability to accurately transcribe cursive text at both the page and line levels, handling contextual character variations and layout structures. Use when the user wants to benchmark on Muharaf, or asks about evaluating this task. Reports CER.
    3 repo stars
  9. ▌
    Multi Bench Eval · qhjqhj00
    Evaluates the emotional intelligence (EI) capabilities of spoken dialogue models in multi-turn interactive settings. It probes basic emotion understanding, advanced emotion support, paralinguistic analysis, and style inference across both Chinese and English dialogues. Use when the user wants to benchmark on MULTI-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Multibanana Eval · qhjqhj00
    Evaluates text-to-image generation models on their ability to synthesize images from multiple reference images and text prompts. It probes adherence to complex instructions, consistency with reference attributes, and robustness to domain mismatches, scale discrepancies, rare concepts, and multilingual text. Use when the user wants to benchmark on MultiBanana, or asks about evaluating this task. Reports MultiBanana score.
    3 repo stars
  11. ▌
    Multiclasslogauc · qhjqhj00
    Compute the MulticlassLogAUC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassLogAUC, or asks how to score with MulticlassLogAUC.
    3 repo stars
  12. ▌
    Multiclassrecall · qhjqhj00
    Compute the MulticlassRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MulticlassRecall, or asks how to score with MulticlassRecall.
    3 repo stars
  13. ▌
    Multiconer2 Eval · qhjqhj00
    Evaluates a model's ability to perform fine-grained named entity recognition and entity linking across multiple languages. It probes whether external knowledge retrieval improves classification of ambiguous or low-frequency entities compared to context-only baselines. Use when the user wants to benchmark on MultiCoNER2, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  14. ▌
    Multidialog Eval · qhjqhj00
    This benchmark evaluates the semantic coherence, acoustic fidelity, and audio-visual synchronization of end-to-end spoken dialogue systems that generate face-to-face conversational audio and video. It probes a model's ability to maintain contextually appropriate dialogue while producing synchronized multimodal outputs without relying on intermediate text representations. Use when the user wants to benchmark on MultiDialog, or asks about evaluating this task. Reports PPL.
    3 repo stars
  15. ▌
    Multifinben Eval · qhjqhj00
    Evaluates large language models on financial reasoning, comprehension, and generation across multiple modalities (text, vision, audio), languages (English, Chinese, Japanese, Spanish, Greek), and task types (IE, QA, summarization, etc.), using a difficulty-aware selection framework to ensure balanced and discriminative assessment. Use when the user wants to benchmark on IESC, FinRED, FINER-ORD, Headlines, TATSA, BRL-Math, FinQA, TATQA, CECTSUM, TGEDTSUM, RMCCF, BigData22, MDSFT, RRE, AIE, LNE, FinanceIQ, chabsa, MultiFin, EFPA, FNS-2023, GRFinNUM, GRMultiFin, GRFinQA, GRFNS-2023, DOLFIN, PolyFiQA-Easy, PolyFiQA-Expert, EnglishOCR, JapaneseOCR, SpanishOCR, GreekOCR, TableBench, MDRM-test, FinAudioSum, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  16. ▌
    Multilabellogauc · qhjqhj00
    Compute the MultilabelLogAUC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelLogAUC, or asks how to score with MultilabelLogAUC.
    3 repo stars
  17. ▌
    Multilabelrecall · qhjqhj00
    Compute the MultilabelRecall metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultilabelRecall, or asks how to score with MultilabelRecall.
    3 repo stars
  18. ▌
    Multilexsum Eval · qhjqhj00
    Evaluates abstractive and extractive summarization models on real-world civil rights lawsuits, testing their ability to synthesize information from extremely long multi-document sources and generate summaries at three distinct length granularities (long, short, tiny). Use when the user wants to benchmark on Multi-LexSum, or asks about evaluating this task. Reports ROUGE-2 F1.
    3 repo stars
  19. ▌
    Multimed St Eval · qhjqhj00
    Evaluates the capability of speech translation models to accurately convert medical speech across five languages (English, Vietnamese, German, French, Mandarin Chinese) into text. It probes both end-to-end and cascaded architectures, as well as the impact of multilingual vs. bilingual training and code-switching handling in a specialized medical domain. Use when the user wants to benchmark on MultiMed-ST, or asks about evaluating this task. Reports BLEU, BERTScore.
    3 repo stars
  20. ▌
    Multisports Eval · qhjqhj00
    Evaluates multi-person spatio-temporal action detection in sports videos, probing the model's ability to localize fine-grained actions across multiple concurrent persons, handle occlusion, and model long-range temporal context. Use when the user wants to benchmark on MultiSports, or asks about evaluating this task. Reports frame-mAP@0.5.
    3 repo stars
  21. ▌
    Multitaskwrapper · qhjqhj00
    Compute the MultitaskWrapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute MultitaskWrapper, or asks how to score with MultitaskWrapper.
    3 repo stars
  22. ▌
    Multiwoz2 1 Eval · qhjqhj00
    Evaluates a model's ability to track dialogue state (domain, slot, value triplets) across conversation turns, specifically probing its robustness to user mind-changes or 'turnback' utterances that modify previously stated intentions. Use when the user wants to benchmark on MultiWOZ 2.1, or asks about evaluating this task. Reports joint goal accuracy.
    3 repo stars
  23. ▌
    Musciclaims Eval · qhjqhj00
    Multimodal scientific claim verification, requiring models to read complex figures and captions to determine if a scientific claim is supported, neutral, or contradicted. It also probes evidence localization, basic visual understanding, cross-modal aggregation, and epistemic sensitivity. Use when the user wants to benchmark on MuSciClaims, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  24. ▌
    Musdb18 Sdr Eval · qhjqhj00
    Evaluates the quality of separated audio sources (vocals, drums, bass, other) from mixed music tracks using source-to-distortion ratio, probing the model's ability to perform supervised music source separation. Use when the user wants to benchmark on MUSDB18, or asks about evaluating this task. Reports SDR.
    3 repo stars
  25. ▌
    Mvpaint T2t Eval · qhjqhj00
    This evaluation probes a model's ability to generate high-quality, multi-view consistent 3D textures on arbitrary meshes conditioned on text instructions. It measures visual fidelity, distributional similarity to ground truth, and cross-view consistency through both automated generative metrics and human preference studies. Use when the user wants to benchmark on Objaverse T2T benchmark, GSO T2T benchmark, or asks about evaluating this task. Reports FID.
    3 repo stars
  26. ▌
    Narrativeqa Eval · qhjqhj00
    Evaluates English long-document retrieval on complex, narrative-style questions, probing deep comprehension and information extraction from lengthy texts. Use when the user wants to benchmark on NarrativeQA, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  27. ▌
    Needlebench Eval · qhjqhj00
    Evaluates large language models' ability to retrieve specific information and perform complex multi-point reasoning within long-context documents. It probes both information-sparse retrieval and information-dense reasoning (Ancestral Trace Challenge) across 32K and 128K token contexts. Use when the user wants to benchmark on NeedleBench, or asks about evaluating this task. Reports Overall.
    3 repo stars
  28. ▌
    Neucir 2023 Eval · qhjqhj00
    Evaluates neural cross-language and multilingual information retrieval systems on news collections in Chinese, Persian, and Russian. It probes a model's ability to rank relevant documents when queries are in English and documents are in different languages, as well as its capacity to unify rankings across multiple languages. Use when the user wants to benchmark on NeuCLIR 2023, or asks about evaluating this task. Reports nDCG@20.
    3 repo stars
  29. ▌
    Neural Stpp Eval · qhjqhj00
    Evaluates the ability of spatio-temporal point process models to accurately capture complex, history-dependent spatial and temporal distributions of discrete events. It probes how well models can compute exact likelihoods for sequences of events in continuous space and time across diverse domains like seismology, epidemiology, and neuroscience. Use when the user wants to benchmark on PINWHEEL, EARTHQUAKES, COVID-19 CASES, BOLD5000, or asks about evaluating this task. Reports log-likelihood per event.
    3 repo stars
  30. ▌
    Nlu Service Eval · qhjqhj00
    Evaluates commercial and open-source NLU platforms on intent classification and named entity recognition across multiple dialogue domains, highlighting limitations in multi-intent support and contextual modeling. Use when the user wants to benchmark on NLU Evaluation Dataset, or asks about evaluating this task. Reports Intent classification accuracy, Entity recognition precision.
    3 repo stars
  31. ▌
    Omniscience Eval · qhjqhj00
    Evaluates the semantic alignment and factual fidelity of automatically generated scientific image captions. It probes whether dense, context-aware captions can replace visual inputs for downstream reasoning tasks and how well they capture complex scientific figures compared to raw human-written captions. Use when the user wants to benchmark on OmniScience, AI2D, MMMU, MM-MT-Bench, MSEarth, or asks about evaluating this task. Reports cross-modal relevance score.
    3 repo stars
  32. ▌
    Omnispatial Eval · qhjqhj00
    This benchmark evaluates vision-language models on advanced spatial reasoning capabilities, specifically probing dynamic reasoning, complex spatial logic, spatial interaction, and perspective-taking. It measures how well models understand and manipulate spatial relationships, temporal changes, and viewpoint shifts beyond basic object recognition. Use when the user wants to benchmark on OmniSpatial, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  33. ▌
    Omnispectra Eval · qhjqhj00
    Evaluates a foundation model's ability to process native-resolution astronomical spectra of variable lengths without resampling, and assesses its zero-shot, few-shot, and supervised performance on stellar property estimation and source classification tasks across diverse spectroscopic surveys. Use when the user wants to benchmark on OmniSpectra Multi-Survey Corpus, or asks about evaluating this task. Reports Mean-Squared Error.
    3 repo stars
  34. ▌
    Onechart Se Eval · qhjqhj00
    Evaluates a model's ability to extract structured information from chart images, including textual OCR accuracy and precise numerical value parsing. It tests the model's capacity to convert visual chart elements into a standardized Python-dict representation, handling both annotated and unannotated charts across multiple languages and rendering styles. Use when the user wants to benchmark on ChartQA-SE, PlotQA-SE, ChartX-SE, ChartY-en, ChartY-zh, or asks about evaluating this task. Reports SCRM AP (strict/slight/high).
    3 repo stars
  35. ▌
    Ood Mol Opt Eval · qhjqhj00
    Out-of-domain molecular property prediction and Bayesian optimization for molecular design. Tests transferability of learned representations to novel tasks. Use when the user wants to benchmark on Out-of-domain molecular design tasks, or asks about evaluating this task. Reports Top performing molecule property.
    3 repo stars
  36. ▌
    Openfwi Fwi Eval · qhjqhj00
    Evaluates deep learning models for seismic full-waveform inversion (FWI) by predicting subsurface velocity models from seismic wavefield data. It probes the model's ability to generalize across varying geological complexities and out-of-distribution scenarios using parameter-efficient fine-tuning. Use when the user wants to benchmark on OpenFWI, or asks about evaluating this task. Reports SSIM.
    3 repo stars
  37. ▌
    Openml Cc18 Eval · qhjqhj00
    Evaluates machine learning classifiers on a curated collection of standardized classification tasks. It probes the reproducibility and comparability of algorithm performance across diverse datasets under consistent, machine-readable evaluation protocols. Use when the user wants to benchmark on OpenML-CC18, or asks about evaluating this task. Reports accuracy_score.
    3 repo stars
  38. ▌
    Openner 1 0 Eval · qhjqhj00
    Evaluates named entity recognition (NER) capabilities across 52 languages and 36 distinct corpora. It probes cross-lingual generalization, robustness to varying entity type ontologies, and the ability of both encoder-based models and LLMs to handle multilingual text with diverse annotation guidelines. Use when the user wants to benchmark on OpenNER 1.0, or asks about evaluating this task. Reports micro-averaged mention-level F1.
    3 repo stars
  39. ▌
    Pas Dataset Eval · qhjqhj00
    Evaluates the transferability and pretraining quality of vision models trained on synthetic domain-specific datasets compared to manually curated and general-domain datasets. It probes the model's ability to generalize to fine-grained classification and object detection tasks within specific domains like birds and food. Use when the user wants to benchmark on CUB-200-2011, NABirds, iNatbirds, Food-101, FoodX-251, Food-2K, or asks about evaluating this task. Reports Top-1 k-NN accuracy.
    3 repo stars
  40. ▌
    Pdbbind Lba Eval · qhjqhj00
    Predicts the binding affinity between a protein pocket and a ligand from 3D structural data. It probes the model's ability to quantify molecular interaction strength and generalize across protein sequence identities. Use when the user wants to benchmark on PDBBind, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  41. ▌
    Perrecbench Eval · qhjqhj00
    Evaluates whether LLMs can capture true personalized user preferences by ranking items or users in groups, while explicitly controlling for confounding factors like user rating bias and item quality. It probes the model's ability to perform comparative reasoning rather than simple rating prediction. Use when the user wants to benchmark on PerRecBench, or asks about evaluating this task. Reports Kendall’s tau.
    3 repo stars
  42. ▌
    Persian Ner Eval · qhjqhj00
    Evaluates the quality and cross-lingual transferability of machine-translated Persian named entity recognition datasets by measuring model performance against original English benchmarks. It probes how well translation-based dataset generation preserves entity boundaries and labels across languages with different scripts and linguistic structures. Use when the user wants to benchmark on CoNLL 2003, OntoNotes 5.0, NCBI Disease, WNUT 2017, or asks about evaluating this task. Reports F1.
    3 repo stars
  43. ▌
    Persian RAG Eval · qhjqhj00
    Evaluates retrieval-augmented generation (RAG) pipelines for Persian text across general, scientific, and formal domains. It probes the ability of sentence embeddings to retrieve relevant context and large language models to generate accurate, faithful, and relevant answers based on that context. Use when the user wants to benchmark on PQuad, Scientific-Specialized, Organizational Report, or asks about evaluating this task. Reports Context Recall.
    3 repo stars
  44. ▌
    Phi3 Safety Eval · qhjqhj00
    Evaluates the safety and refusal capabilities of language models across multiple risk categories including harmful content generation, jailbreaking, stereotype bias, privacy leaks, and toxicity detection. It measures how well models balance harmlessness (refusing unsafe prompts) and helpfulness (complying with safe prompts) in both single- and multi-turn interactions. Use when the user wants to benchmark on XSTest, DecodingTrust, ToxiGen, XSafety, RTP-LX, Microsoft Internal Automated Measurement, or asks about evaluating this task. Reports IPRR.
    3 repo stars
  45. ▌
    Phishnchips Eval · qhjqhj00
    Evaluates the security and robustness of autonomous LLM email agents against phishing attacks by measuring how different system prompt configurations affect detection sensitivity and operational false positive rates. It specifically probes the model's ability to maintain high recall while minimizing usability costs, and tests adversarial brittleness under infrastructure phishing conditions where attacker-controlled domains match sender addresses. Use when the user wants to benchmark on Synthetic Email Phishing Corpus, or asks about evaluating this task. Reports Net Effectiveness (Recall-FPR).
    3 repo stars
  46. ▌
    Phreshphish Eval · qhjqhj00
    Evaluates phishing website detection models on temporally disjoint, real-world data with realistic base rates, while mitigating training-to-test leakage and varying difficulty levels. Use when the user wants to benchmark on PhreshPhish, or asks about evaluating this task. Reports Precision-Recall.
    3 repo stars
  47. ▌
    Pii Masking Eval · qhjqhj00
    Evaluates the ability of PII masking models to correctly identify and classify sensitive information in text. It probes performance across diverse contexts, multilingual inputs, noisy formats, and evolving entity types. Use when the user wants to benchmark on PII Masking Dataset, or asks about evaluating this task. Reports non-identification.
    3 repo stars
  48. ▌
    Pii Tagging Eval · qhjqhj00
    Evaluates a model's ability to identify and extract private or sensitive information spans from text across legal, healthcare, and finance domains. It probes the model's capacity for fine-grained named entity recognition under limited labeled data and domain-specific privacy definitions. Use when the user wants to benchmark on ECHR, MACCROBAT, PUPA (Finance Subset), or asks about evaluating this task. Reports F1.
    3 repo stars
  49. ▌
    Plaba Track Eval · qhjqhj00
    Evaluates NLP systems and large language models on adapting biomedical abstracts to plain language for lay consumers. It probes capabilities in text simplification, term replacement, factual faithfulness, and conciseness while measuring alignment with human expert judgments. Use when the user wants to benchmark on TREC PLABA, or asks about evaluating this task. Reports SARI.
    3 repo stars
  50. ▌
    Planetarium Eval · qhjqhj00
    Evaluates an LLM's ability to translate natural language planning task descriptions into valid, semantically equivalent Planning Domain Definition Language (PDDL) code. It specifically probes the model's capacity to accurately capture initial states, goal states, and object relationships while adhering to formal planning semantics. Use when the user wants to benchmark on Planetarium, or asks about evaluating this task. Reports equivalence.
    3 repo stars
  51. ▌
    Pmindia Nmt Eval · qhjqhj00
    This evaluation probes the quality of automatic machine translation between English and 13 Indian languages using a parallel corpus. It measures how well NMT systems can handle diverse linguistic structures, including abugida scripts and agglutinative morphology, across low-resource language pairs. Use when the user wants to benchmark on PMIndia, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  52. ▌
    Polychartqa Eval · qhjqhj00
    Evaluates multimodal language models' ability to answer questions that require reasoning across multiple charts or images. It probes visual decomposition, sub-chart localization, and handling of complex multi-visual contexts versus single-chart inputs. Use when the user wants to benchmark on PolyChartQA, MultiChartQA-RQ1, or asks about evaluating this task. Reports L-Accuracy.
    3 repo stars
  53. ▌
    Pope Nocaps Eval · qhjqhj00
    Tests object perception and hallucination on images without captions, evaluating whether LVLMs can ground object detection purely from visual input without textual priors. Use when the user wants to benchmark on POPE-NoCaps, or asks about evaluating this task. Reports Acc.
    3 repo stars
  54. ▌
    Power Divergence · qhjqhj00
    Compute the power_divergence metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute power_divergence, or asks how to score with power_divergence.
    3 repo stars
  55. ▌
    Privacylens Eval · qhjqhj00
    Assesses an LLM agent's ability to understand and follow privacy norms while performing real-world tasks. It measures both helpfulness and the rate at which sensitive information is incorrectly exposed. Use when the user wants to benchmark on PrivacyLens, or asks about evaluating this task. Reports privacy leakage rate.
    3 repo stars
  56. ▌
    Progressftx Eval · qhjqhj00
    Evaluates a progressive feature transmission protocol for split inference at the wireless edge, measuring how efficiently features are transmitted to meet target inference accuracy or uncertainty thresholds under varying channel conditions. Use when the user wants to benchmark on GM dataset, MNIST, or asks about evaluating this task. Reports average communication latency.
    3 repo stars
  57. ▌
    Propsegment Eval · qhjqhj00
    Evaluates a model's ability to decompose sentences into atomic semantic units (propositional segmentation) and determine entailment relationships between text spans. It probes fine-grained compositional semantic alignment and partial entailment recognition beyond sentence-level NLI. Use when the user wants to benchmark on PropSegmEnt, or asks about evaluating this task. Reports Precision/Recall/F1w (macro-averaged).
    3 repo stars
  58. ▌
    Pulseimpute Eval · qhjqhj00
    Evaluates models on imputing missing values in pulsative physiological signals (ECG and PPG) under realistic, data-driven missingness patterns. It further assesses clinical utility by measuring downstream performance on heartbeat detection and cardiac classification tasks. Use when the user wants to benchmark on ECG, PPG, or asks about evaluating this task. Reports MSE, F1 Score.
    3 repo stars
  59. ▌
    Pushupbench Eval · qhjqhj00
    Evaluates video-language models on long-form repetition counting and temporal reasoning. It probes whether models can accurately track state changes and count actions across extended video clips, revealing weaknesses in spatio-temporal tracking compared to supervised baselines. Use when the user wants to benchmark on PushupBench, or asks about evaluating this task. Reports Exact Match.
    3 repo stars
  60. ▌
    Pyvision Rl Eval · qhjqhj00
    Evaluates open-weight multimodal agentic models on visual search, multimodal mathematical reasoning, multi-turn tool use, and video spatial reasoning. It probes the model's ability to dynamically construct context, invoke tools, and perform long-horizon reasoning with high visual token efficiency. Use when the user wants to benchmark on V*, HRBench-4K, HRBench-8K, MathVerse, MathVision, WeMath, DynaMath, TIR-Bench, VSI-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  61. ▌
    Qianfan Ocr Eval · qhjqhj00
    This evaluation probes a unified vision-language model's ability to perform end-to-end document intelligence, including specialized OCR, general text recognition, document understanding, and key information extraction across diverse document types and multilingual scenarios. Use when the user wants to benchmark on Omni-Doc-Bench v1.5, OLMOCRBench, OCRBench, DocVQA, ChartQA, Nanonets KIE, or asks about evaluating this task. Reports normalized accuracy (0-100).
    3 repo stars
  62. ▌
    Qualispeech Eval · qhjqhj00
    Evaluates auditory large language models' ability to perceive and describe low-level speech quality aspects, including noise, distortion, speed, continuity, listening effort, naturalness, and overall quality. It probes both numerical score prediction and natural language reasoning/description generation. Use when the user wants to benchmark on QualiSpeech, or asks about evaluating this task. Reports PCC.
    3 repo stars
  63. ▌
    Quito Bench Eval · qhjqhj00
    Evaluates time series forecasting models on a regime-balanced benchmark stratified by trend, seasonality, and forecastability. It probes a model's ability to handle varying context lengths, forecast horizons, and multivariate dependencies while mitigating domain bias and information leakage. Use when the user wants to benchmark on QuitoBench, or asks about evaluating this task. Reports MAE.
    3 repo stars
  64. ▌
    Radioactive Eval · qhjqhj00
    Evaluates interactive 3D medical image segmentation by measuring how well models segment target structures under varying human-in-the-loop prompting strategies (points, boxes, scribbles) and iterative refinement protocols. It specifically probes the trade-off between interaction effort and segmentation accuracy across 2D and 3D architectures. Use when the user wants to benchmark on RadioActive, or asks about evaluating this task. Reports Dice.
    3 repo stars
  65. ▌
    RAG Medical Eval · qhjqhj00
    Evaluates the effectiveness and efficiency of Retrieval-Augmented Generation (RAG) systems across medical and general knowledge domains. It probes how different RAG pipeline components (chunking, indexing, query classification, augmentation, and prompting) impact answer accuracy and response latency on question-answering and information extraction tasks. Use when the user wants to benchmark on MMLU, PubMedQA, PromptNER, Query Classification Dataset, or asks about evaluating this task. Reports accuracy (acc).
    3 repo stars
  66. ▌
    Ragen Agent Eval · qhjqhj00
    Evaluates LLM agents' multi-turn decision-making and reasoning capabilities across symbolic planning, risk-sensitive reasoning, and realistic web interaction environments. It probes the agent's ability to complete interactive tasks under noisy or probabilistic feedback while maintaining exploration and training stability. Use when the user wants to benchmark on Bandit, Sokoban, Frozen Lake, WebShop, or asks about evaluating this task. Reports success rate.
    3 repo stars
  67. ▌
    Randumb Ocl Eval · qhjqhj00
    Evaluates continual learning methods in online, exemplar-free, and low-exemplar regimes by measuring how well a model retains knowledge of previously seen classes after processing a single pass of sequential data. It specifically tests whether fixed random representations can match or exceed learned representations in these constrained settings. Use when the user wants to benchmark on MNIST, CIFAR10, CIFAR100, TinyImageNet200, miniImageNet100, or asks about evaluating this task. Reports average_accuracy.
    3 repo stars
  68. ▌
    Recruitview Eval · qhjqhj00
    This benchmark evaluates multimodal models on predicting continuous personality traits and interview performance scores from video, audio, and text inputs. It probes the model's ability to perform fine-grained behavioral analysis and regression across psychometric targets. Use when the user wants to benchmark on RecruitView, or asks about evaluating this task. Reports Spearman's ρ.
    3 repo stars
  69. ▌
    Repobench C Eval · qhjqhj00
    Evaluates autoregressive language models on predicting the next line of code using provided in-file and cross-file contexts. Use when the user wants to benchmark on RepoBench-C, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  70. ▌
    Repobench P Eval · qhjqhj00
    Evaluates an end-to-end pipeline that first retrieves cross-file snippets and then predicts the next line of code using both the in-file context and retrieved snippets. Use when the user wants to benchmark on RepoBench-P, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  71. ▌
    Repobench R Eval · qhjqhj00
    Evaluates a model's ability to retrieve relevant cross-file code snippets given an in-file context for predicting the next line of code. Use when the user wants to benchmark on RepoBench-R, or asks about evaluating this task. Reports acc@1.
    3 repo stars
  72. ▌
    Repro Bench Eval · qhjqhj00
    Evaluates whether agentic AI systems can accurately assess the computational reproducibility of social science research by comparing original paper findings against results reproduced from provided raw data and code. It probes end-to-end agentic reasoning, including command execution, debugging, and result interpretation in a simulated research environment. Use when the user wants to benchmark on REPRO-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  73. ▌
    Researchgym Eval · qhjqhj00
    Evaluates the capability of LLM agents to conduct closed-loop scientific research by proposing hypotheses, executing experiments, and outperforming human baselines on repurposed real-world AI papers. It probes long-horizon planning, resource management, and autonomous experimentation under realistic tool constraints. Use when the user wants to benchmark on ResearchGym, or asks about evaluating this task. Reports improvement over baselines.
    3 repo stars
  74. ▌
    Respondeoqa Eval · qhjqhj00
    This benchmark evaluates large language models on bilingual Latin-English question answering across knowledge-based, skill-based (grammar, scansion, literary devices), multihop reasoning, and translation tasks. It probes models' ability to handle classical language morphology, poetic meter analysis, and cross-lingual generation under constrained and unconstrained settings. Use when the user wants to benchmark on RespondeoQA, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  75. ▌
    Retrievalfallout · qhjqhj00
    Compute the RetrievalFallOut metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalFallOut, or asks how to score with RetrievalFallOut.
    3 repo stars
  76. ▌
    Retrievalhitrate · qhjqhj00
    Compute the RetrievalHitRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalHitRate, or asks how to score with RetrievalHitRate.
    3 repo stars
  77. ▌
    Retrotrucks Eval · qhjqhj00
    Evaluates video anomaly detection models on dashcam footage, specifically testing their ability to detect complex traffic anomalies like collisions and skidding in dynamic, real-world driving scenes. It also benchmarks performance against standard pedestrian anomaly detection datasets to highlight challenges posed by moving cameras and contextual anomalies. Use when the user wants to benchmark on RetroTrucks, UCSD Ped1, UCSD Ped2, ShanghaiTech, or asks about evaluating this task. Reports AUC-ROC.
    3 repo stars
  78. ▌
    Reviewbench Eval · qhjqhj00
    Evaluates the quality and accuracy of automated peer reviews generated by LLMs. It probes both the semantic alignment of review text against paper-specific rubrics and the precision of predicted numerical ratings and acceptance decisions. Use when the user wants to benchmark on ReviewBench, or asks about evaluating this task. Reports Rubric Overall Score.
    3 repo stars
  79. ▌
    Rextthewild Eval · qhjqhj00
    Evaluates multimodal models' ability to understand real-world medical photographs by answering clinician-verified multiple-choice questions across seven clinical domains. It probes capabilities in geometric perception, anatomical localization, clinical characterization, and causal reasoning. Use when the user wants to benchmark on ReXInTheWild, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  80. ▌
    Rgbt Ground Eval · qhjqhj00
    Evaluates multi-modal visual grounding capabilities by requiring models to localize objects in images using both RGB and thermal infrared (TIR) modalities guided by text queries. It specifically probes robustness under complex real-world conditions such as low-light environments, small object sizes, and diverse weather/illumination variations. Use when the user wants to benchmark on RGBT-Ground, or asks about evaluating this task. Reports Acc@0.5.
    3 repo stars
  81. ▌
    Rgt Seismic Eval · qhjqhj00
    Evaluates a model's ability to perform continuous regression for Relative Geologic Time (RGT) estimation from 2D seismic images. It probes the model's capacity to learn stratigraphic continuity and structural consistency across diverse geological settings, testing generalization from synthetic labeled data to unlabeled real-world field data. Use when the user wants to benchmark on Field Seismic Dataset, Synthetic Seismic Dataset, or asks about evaluating this task. Reports regression.
    3 repo stars
  82. ▌
    Safe Speech Eval · qhjqhj00
    Evaluates the capability of classifiers to detect sexist, abusive, offensive, and hate speech in conversational text across multiple granularity levels and established benchmarks. It probes fine-grained toxicity detection, cross-dataset generalization, and performance against strong supervised and LLM baselines. Use when the user wants to benchmark on EDOS (SemEval 2023), OffensEval 2019, AbusEval, HatEval, or asks about evaluating this task. Reports F1.
    3 repo stars
  83. ▌
    Safeprotein Eval · qhjqhj00
    Evaluates the biosafety risks and jailbreak vulnerabilities of protein foundation models by measuring their ability to reconstruct harmful protein sequences and 3D structures from partially masked inputs. It probes whether models can bypass safety filters and generate biologically dangerous proteins when given sequence and structural prompts. Use when the user wants to benchmark on SafeProtein-Bench, or asks about evaluating this task. Reports jailbreak success rate.
    3 repo stars
  84. ▌
    Safetunebed Eval · qhjqhj00
    Evaluates the safety alignment preservation and task utility of LLMs after parameter-efficient fine-tuning under data-poisoning attacks. It measures how well defenses maintain core capabilities while resisting harmful behavior injection. Use when the user wants to benchmark on MMLU, MT-Bench, AdvBench, PolicyEval, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  85. ▌
    Sagalee Asr Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) performance on the Oromo language using real-world, crowd-sourced audio data. It measures how well different model architectures (Conformer trained from scratch, Whisper fine-tuned) transcribe spoken Oromo into text under varying acoustic conditions. Use when the user wants to benchmark on Sagalee, or asks about evaluating this task. Reports WER.
    3 repo stars
  86. ▌
    Sailcompass Eval · qhjqhj00
    Evaluates large language models on language proficiency, reading comprehension, reasoning, and cultural understanding across three Southeast Asian languages (Indonesian, Vietnamese, Thai). It covers eight diverse tasks including question answering, machine translation, text summarization, multiple-choice exams, commonsense reasoning, machine reading comprehension, natural language inference, and sentiment analysis. Use when the user wants to benchmark on XQuAD, TyDiQA, Flores-200, ThaiSum, IndoSum, XLSUM, M3Exam, XCOPA, BELEBELE, XNLI, IndoNLI, Wisesight, Indolem, VSMEC, or asks about evaluating this task. Reports Exact Match.
    3 repo stars
  87. ▌
    Salad Bench Eval · qhjqhj00
    Evaluates the safety, robustness, and helpfulness of Large Language Models across a hierarchical taxonomy of 6 domains, 16 tasks, and 66 categories. It measures performance on benign, adversarial (attack-enhanced), and defense-enhanced prompts, as well as multiple-choice safety questions, while also benchmarking the effectiveness of various attack and defense strategies. Use when the user wants to benchmark on SALAD-Bench, ToxicChat, Beavertails, SafeRLHF, Harmbench, Lifetox, AdvBench-50, or asks about evaluating this task. Reports Safety Rate, Attack Success Rate (ASR).
    3 repo stars
  88. ▌
    Sarena Icon Eval · qhjqhj00
    Evaluates a model's ability to generate scalable vector graphics (SVG) from text prompts and reference images, measuring visual fidelity, semantic alignment, structural success, and code efficiency. Use when the user wants to benchmark on SArena-Icon, or asks about evaluating this task. Reports SR.
    3 repo stars
  89. ▌
    Scatspotter Eval · qhjqhj00
    Evaluates object detection and instance segmentation capabilities on real-world images of dog feces. It specifically probes model robustness to camouflage, occlusion, varying lighting conditions, and small object detection in outdoor urban environments. Use when the user wants to benchmark on ScatSpotter, or asks about evaluating this task. Reports mAP.
    3 repo stars
  90. ▌
    Scene Bench Eval · qhjqhj00
    Evaluates the factual consistency and scene graph adherence of text-to-image generation models. It probes whether generated images accurately preserve specified objects and their spatial/relational configurations as defined by input scene graphs, rather than just measuring aesthetic quality or text-image alignment. Use when the user wants to benchmark on Visual Genome (VG) test set, MegaSG, or asks about evaluating this task. Reports SGScore.
    3 repo stars
  91. ▌
    Scene Smith Eval · qhjqhj00
    Evaluates text-to-3D indoor scene generation systems on their ability to produce dense, physically plausible, and prompt-faithful environments. It probes both visual realism and simulation-readiness, measuring collision-free layouts and stable physics properties required for robotics policy testing. Use when the user wants to benchmark on SceneSmith Prompt Corpus, or asks about evaluating this task. Reports Realism Win%.
    3 repo stars
  92. ▌
    Scenicrules Eval · qhjqhj00
    Evaluates autonomous driving agents on their ability to navigate stochastic traffic scenarios while satisfying a hierarchical set of multi-objective specifications. It probes how well agents balance conflicting goals like collision avoidance, road compliance, passenger comfort, and progress under varying priority constraints. Use when the user wants to benchmark on ScenicRules Benchmark, or asks about evaluating this task. Reports Violation Score (VS).
    3 repo stars
  93. ▌
    Scholawrite Eval · qhjqhj00
    Evaluates an LLM's ability to predict human scholarly writing intentions from a LaTeX draft and to iteratively edit the draft according to those intentions. It measures lexical diversity, topic consistency, and intention coverage across a 100-iteration self-writing process. Use when the user wants to benchmark on SCHOLAWRITE, or asks about evaluating this task. Reports intention coverage.
    3 repo stars
  94. ▌
    Scigenbench Eval · qhjqhj00
    Evaluates the logical correctness, structural fidelity, and information utility of AI-generated scientific images. It probes whether generated visuals accurately encode domain-specific facts and geometric relationships, and whether they are indispensable for solving visually grounded scientific quizzes. Use when the user wants to benchmark on SciGenBench, or asks about evaluating this task. Reports inverse_validation_rate.
    3 repo stars
  95. ▌
    Screen Spot Eval · qhjqhj00
    This evaluation probes a vision-language model's ability to localize specific UI elements within graphical user interfaces based on natural language instructions. It tests precise coordinate prediction and cross-resolution generalization across mobile, desktop, and web platforms. Use when the user wants to benchmark on ScreenSpot, ScreenSpot-v2, ScreenSpot-Pro, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Se Toxicity Eval · qhjqhj00
    Evaluates the ability of contemporary toxicity detection models to correctly identify toxic language in software engineering contexts, such as code reviews and developer chat logs. It probes whether general-purpose classifiers can handle domain-specific terminology and contextual nuances without significant performance degradation. Use when the user wants to benchmark on Jigsaw Sample, Code Review, Gitter Ethereum, or asks about evaluating this task. Reports F-Score.
    3 repo stars
  97. ▌
    Seas Safety Eval · qhjqhj00
    This evaluation probes the safety alignment and refusal capabilities of LLMs when exposed to harmful or adversarial prompts. It measures the frequency of unsafe model outputs to quantify vulnerability, while simultaneously tracking general instruction-following scores to ensure that safety hardening does not degrade overall utility. Use when the user wants to benchmark on SEAS-Test, BeaverTrail, HH-RLHF, XSTest, or asks about evaluating this task. Reports Attack Success Rate (ASR).
    3 repo stars
  98. ▌
    Secretbench Eval · qhjqhj00
    Evaluates the capability of automated secret detection tools to accurately identify hardcoded secrets (e.g., API keys, passwords, private keys) in source code repositories. It probes the tools' ability to balance high recall for true secrets against low false positive rates to mitigate alert fatigue. Use when the user wants to benchmark on SecretBench, or asks about evaluating this task. Reports Precision.
    3 repo stars
  99. ▌
    Sega Layout Eval · qhjqhj00
    Evaluates a model's ability to generate content-aware graphic layouts from background images and instructions. It probes spatial reasoning, adherence to design principles (alignment, overlap, occlusion), and aesthetic quality. Use when the user wants to benchmark on PKU, CGL, Crello, or asks about evaluating this task. Reports Ali.
    3 repo stars
  100. ▌
    Semantic Kg Eval · qhjqhj00
    Evaluates the ability of semantic similarity methods to correctly classify pairs of natural language statements as semantically similar (label 1) or dissimilar (label 0). It specifically probes how well models handle controlled semantic variations (node and edge perturbations) across general and domain-specific knowledge domains. Use when the user wants to benchmark on Semantic-KG Benchmark, or asks about evaluating this task. Reports F1-score.
    3 repo stars