all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 60 of 76

  1. ▌
    Diacr Ita Eval · qhjqhj00
    Binary classification of lexical semantic change for target words across two diachronic time periods. It probes whether models can reliably detect meaning shifts in Italian using corpus pairs from newspapers and books. Use when the user wants to benchmark on DIACR-Ita, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  2. ▌
    Dipco Wer Eval · qhjqhj00
    Evaluates distant speech recognition and noise robustness by measuring word error rate on close-talk and far-field recordings of natural dinner conversations. It tests the model's ability to handle uncontrolled acoustic conditions, background music, and spatially diverse microphone placements. Use when the user wants to benchmark on DiPCo, or asks about evaluating this task. Reports WER.
    3 repo stars
  3. ▌
    Discoeval Eval · qhjqhj00
    Evaluates whether sentence representations capture discourse-aware semantics by testing performance on tasks involving sentence ordering, discourse relations, and coherence across multiple domains. Use when the user wants to benchmark on DiscoEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  4. ▌
    Discophon Eval · qhjqhj00
    This benchmark evaluates the ability of discrete speech models to unsupervisedly discover phoneme inventories and capture phonemic contrasts. It measures how well predicted units align with gold phonemes through phonetic similarity, recognition error, and temporal segmentation accuracy across multiple languages. Use when the user wants to benchmark on discoPhon, or asks about evaluating this task. Reports PNMI.
    3 repo stars
  5. ▌
    Docgenome Eval · qhjqhj00
    This benchmark evaluates multi-modal large language models on their ability to parse and understand complex scientific documents. It probes capabilities across document classification, visual grounding of text elements, open-ended single- and multi-page question answering, layout detection, and modality-to-LaTeX transformation. Use when the user wants to benchmark on DocGenome, or asks about evaluating this task. Reports GPT-acc.
    3 repo stars
  6. ▌
    Domainnet Eval · qhjqhj00
    Evaluates multi-source domain adaptation methods on image classification tasks across multiple domains with varying visual styles and categories. Use when the user wants to benchmark on DomainNet, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  7. ▌
    Domainsum Eval · qhjqhj00
    Evaluates abstractive text summarization models under varying degrees of domain shift (genre, style, topic) to measure performance degradation and distributional divergence across hierarchical granularity levels. Use when the user wants to benchmark on DomainSum, or asks about evaluating this task. Reports ROUGE.
    3 repo stars
  8. ▌
    Dpg Bench Eval · qhjqhj00
    Evaluates dense prompt following on multi-requirement prompts by decomposing them into dependency-structured VQA checks spanning entity presence, attributes, relations, and counts. Use when the user wants to benchmark on DPG-Bench, or asks about evaluating this task. Reports Overall.
    3 repo stars
  9. ▌
    Dr Spider Eval · qhjqhj00
    Evaluates the robustness of text-to-SQL models against semantic-preserving and semantic-changing perturbations across database schemas, natural language questions, and SQL queries. It measures how well models maintain execution accuracy when inputs are altered, revealing architectural vulnerabilities in entity linking, decoder design, and value prediction. Use when the user wants to benchmark on Dr.Spider, or asks about evaluating this task. Reports execution accuracy (EX).
    3 repo stars
  10. ▌
    Dre Bench Eval · qhjqhj00
    Evaluates large language models' fluid intelligence and abstract rule generalization across four hierarchical cognitive levels (Attribute, Spatial, Sequential, Conceptual). It probes the model's ability to dynamically adapt to varying task complexity and apply learned rules to novel grid-based reasoning problems. Use when the user wants to benchmark on DRE-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  11. ▌
    Droidspan Eval · qhjqhj00
    Evaluates the sustainability and robustness of a dynamic behavioral profiling approach (DroidSpan) for Android malware detection over time and against code obfuscation, compared to a static baseline (MamaDroid). Use when the user wants to benchmark on all-data, oldBen+oldMal, MalObf, or asks about evaluating this task. Reports F1-measure.
    3 repo stars
  12. ▌
    Droughted Eval · qhjqhj00
    Evaluates time-series forecasting models on predicting U.S. drought severity across 1 to 6 week horizons using meteorological and static features. It probes both regression accuracy and multi-class classification performance for drought monitoring levels. Use when the user wants to benchmark on DroughtED, or asks about evaluating this task. Reports MAE.
    3 repo stars
  13. ▌
    Dualbench Eval · qhjqhj00
    Evaluates a model's ability to generate synchronized background audio and intelligible speech from video input, measuring audio quality, distribution matching, and audio-video temporal alignment. Use when the user wants to benchmark on DualBench, VGGSound, or asks about evaluating this task. Reports FAD↓.
    3 repo stars
  14. ▌
    Dualgauge Eval · qhjqhj00
    Probes the joint functional correctness and security of LLM-generated code, while also evaluating an automated framework's ability to execute code in sandboxes and semantically judge test outcomes against human ground truth. Use when the user wants to benchmark on DualGauge-Bench, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  15. ▌
    Dualtoken Eval · qhjqhj00
    Evaluates a unified vision tokenizer's capacity to decouple and jointly optimize low-level perceptual reconstruction and high-level semantic understanding. It probes zero-shot classification, cross-modal retrieval, image reconstruction fidelity, and downstream multimodal reasoning capabilities. Use when the user wants to benchmark on ImageNet-1K, Flickr8K, VQAv2, POPE, MME, SEED-IMG, MMBench, MM-Vet, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  16. ▌
    Dyad Arch Eval · qhjqhj00
    Evaluates a block-sparse linear layer approximation (DYAD) against dense baselines across standard NLP and vision benchmarks, measuring accuracy preservation and computational efficiency. Use when the user wants to benchmark on BLIMP, OPENLLM, GLUE+, MNIST, or asks about evaluating this task. Reports BLIMP accuracy.
    3 repo stars
  17. ▌
    E2e Gmner Eval · qhjqhj00
    Evaluates a model's ability to perform end-to-end grounded multimodal named entity recognition, jointly identifying entity spans in text, predicting their semantic types, and grounding them to corresponding bounding boxes in an associated image. It probes the model's capacity for multimodal alignment, structured generation, and robustness to annotation noise via chain-of-thought reasoning. Use when the user wants to benchmark on Twitter-GMNER, Twitter-FMNERG, or asks about evaluating this task. Reports GMNER.
    3 repo stars
  18. ▌
    E3d Bench Eval · qhjqhj00
    Evaluates the effectiveness, robustness, and inference efficiency of end-to-end 3D Geometric Foundation Models across sparse-view depth estimation, video depth estimation, and multi-view relative pose estimation. It probes models' ability to generalize across diverse domains including indoor, outdoor, aerial, and highly dynamic scenes under both normalized and metric-scale settings. Use when the user wants to benchmark on DTU, ETH3D, KITTI, Tanks and Temples, ScanNet, Bonn, TUM Dynamics, Sintel, PointOdyssey, Syndrome, CO3Dv2, RealEstate10K, ScanNet-eval, KITTI Odometry, ADT, ACID, ULTRRA, or asks about evaluating this task. Reports AbsRel.
    3 repo stars
  19. ▌
    Ears Wham Eval · qhjqhj00
    Evaluates speech enhancement models by measuring how well they recover clean speech from noisy mixtures across a wide range of signal-to-noise ratios and speaker demographics. The benchmark covers both controlled training/validation splits and a blind test set with unseen speakers and noise. Use when the user wants to benchmark on EARS-WHAM, or asks about evaluating this task. Reports SI-SDR.
    3 repo stars
  20. ▌
    Echochain Eval · qhjqhj00
    Evaluates how voice assistants handle mid-generation interruptions by testing their ability to revise in-progress responses while maintaining context and switching objectives. It probes state-update reasoning, contextual inertia, interruption amnesia, and objective displacement under full-duplex interaction conditions. Use when the user wants to benchmark on EchoChain, or asks about evaluating this task. Reports pass_fail.
    3 repo stars
  21. ▌
    Ediref Erc Efr · qhjqhj00
    Evaluates emotion recognition and emotion-flip reasoning in multi-party conversations, specifically identifying trigger utterances that cause emotional shifts in both code-mixed (Hindi-English) and monolingual English dialogues. Use when the user wants to benchmark on E-MaSaC, MELD-FR, or asks about evaluating this task. Reports weighted F1, F1 score for trigger utterances.
    3 repo stars
  22. ▌
    Eee Bench Eval · qhjqhj00
    Probes multimodal reasoning and visual diagram interpretation in electrical and electronics engineering. It tests whether models can integrate complex circuit and system diagrams with textual problem descriptions to apply domain-specific knowledge and perform accurate calculations or logical deductions. Use when the user wants to benchmark on EEE-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  23. ▌
    Egohumans Eval · qhjqhj00
    Evaluates the ability of models to perform robust multi-human tracking and identity association in unconstrained egocentric 3D environments. It probes how well algorithms handle severe occlusions, dynamic activities, and camera-agnostic spatial reasoning when fusing egocentric and secondary views. Use when the user wants to benchmark on EgoHumans, or asks about evaluating this task. Reports IDF1.
    3 repo stars
  24. ▌
    Egonormia Eval · qhjqhj00
    Evaluates vision-language models' ability to understand and reason about physical-social norms in egocentric video scenarios. It probes whether models can correctly select normative actions, justify them, and identify plausible alternatives in conflict-prone situations. Use when the user wants to benchmark on EgoNormia, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  25. ▌
    Egoschema Eval · qhjqhj00
    Probes long-term visual memory and temporal reasoning in video-language models by requiring them to answer multiple-choice questions about very long-form videos. It measures the model's ability to retain and retrieve information across extended durations without relying on short clip analysis. Use when the user wants to benchmark on EgoSchema, or asks about evaluating this task. Reports QA Accuracy.
    3 repo stars
  26. ▌
    Egoxtreme Eval · qhjqhj00
    Evaluates the robustness of 6D object pose estimation models under extreme real-world visual conditions, including severe motion blur, dynamic lighting, and smoke. It also benchmarks temporal tracking strategies in highly dynamic egocentric scenarios to assess motion-aware inference capabilities. Use when the user wants to benchmark on EgoXtreme, or asks about evaluating this task. Reports ADD(-S) recall.
    3 repo stars
  27. ▌
    Histo Vl Eval · qhjqhj00
    Evaluates vision-language foundation models on diverse histopathology clinical tasks (detection, subtyping, grading, mutation prediction) to assess their robustness to textual/visual perturbations, magnification changes, stain normalization, and model calibration. Use when the user wants to benchmark on CRC-100K, BreakHist, DataBiox, GasHisSDB, Breast IDC, LC25000-lung, or asks about evaluating this task. Reports balanced accuracy.
    3 repo stars
  28. ▌
    Hotpotqa Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform multi-hop question answering by reasoning across multiple documents. It specifically probes explainability through supporting fact prediction and tests robustness against distractor paragraphs and large-scale retrieval contexts. Use when the user wants to benchmark on HotpotQA, or asks about evaluating this task. Reports F1.
    3 repo stars
  29. ▌
    Hulu Med Eval · qhjqhj00
    Evaluates a unified vision-language model's capability to perform holistic medical understanding across text-only queries, 2D/3D medical images, and surgical videos. It probes visual question answering, radiology report generation, clinical reasoning, and multilingual medical dialogue. Use when the user wants to benchmark on MIMIC-CXR, CheXpert, IU X-ray, MedMNIST-2D, M3D, 3D-RAD, AMOS-MM, MedFrameQA, Cholec80-VQA, EndoVis18-VQA, PSI-AVA-VQA, SurgeryVideoQA, MMedBench, RareBench, HealthBench, MMLU-Pro-Med, MedXQA, Medbullets, SGPQA, MedMCQA, MedQA, PubMedQA, MedXpertQA, MMMU-Med, OmniMedVQA, PMC-VQA, VQA-RAD, SLAKE, PathVQA, or asks about evaluating this task. Reports RaTEScore.
    3 repo stars
  30. ▌
    Humanref Eval · qhjqhj00
    Evaluates a model's ability to detect all instances of a person matching a natural language description in an image, including handling multiple instances and correctly rejecting cases where the described person is absent. Use when the user wants to benchmark on HumanRef, or asks about evaluating this task. Reports DensityF1 Score.
    3 repo stars
  31. ▌
    Hummusqa Eval · qhjqhj00
    Evaluates large audio-language models on music understanding by testing their ability to answer multiple-choice questions about audio excerpts. It probes perceptual grounding, structural/harmonic/cultural reasoning, and robustness against text-only shortcuts or answer-position bias. Use when the user wants to benchmark on HumMusQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  32. ▌
    Hybridna Eval · qhjqhj00
    Evaluates the capability of DNA foundation models to perform short-range and long-range genomic understanding tasks, as well as their ability to generate biologically plausible cis-regulatory elements. It probes sequence classification, variant effect prediction, and generative design across multiple species and cell types. Use when the user wants to benchmark on GUE, BEND, LRB, CRE (regLM), or asks about evaluating this task. Reports MCC.
    3 repo stars
  33. ▌
    Hybridqa Eval · qhjqhj00
    Multi-hop question answering that requires integrating information from both tabular and textual sources. It probes a model's ability to perform cross-modal reasoning and extract precise answers from heterogeneous data. Use when the user wants to benchmark on HybridQA, or asks about evaluating this task. Reports exact match (EM).
    3 repo stars
  34. ▌
    Hyperglm Eval · qhjqhj00
    Evaluates a multimodal model's ability to generate and anticipate scene graphs from video frames, capturing spatial object relationships and causal temporal transitions. It also tests video question answering, captioning, and relation reasoning capabilities by leveraging hypergraph structures to model multi-way interactions. Use when the user wants to benchmark on VSGR, PVSG, Action Genome, or asks about evaluating this task. Reports Recall (R) / mean Recall (mR).
    3 repo stars
  35. ▌
    Ics Flow Eval · qhjqhj00
    Evaluates machine learning models' ability to detect and classify cyberattacks in Industrial Control Systems (ICS) network traffic. It probes the capability to distinguish between normal operations and specific attack types like DDoS, IP-Scan, MitM, Port-Scan, and Replay using flow-level features. Use when the user wants to benchmark on ICS-Flow, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  36. ▌
    Inbreast Eval · qhjqhj00
    Evaluates a deep learning model's ability to detect and classify malignant lesions in mammograms. It measures classification accuracy at the breast level and detection/localization sensitivity against false positive rates. Use when the user wants to benchmark on INbreast, or asks about evaluating this task. Reports AUC.
    3 repo stars
  37. ▌
    Ineqmath Eval · qhjqhj00
    This benchmark evaluates large language models' ability to perform informal mathematical reasoning on Olympiad-level inequality problems. It probes step-wise deductive chain integrity by decomposing proofs into bound estimation and relation prediction subtasks, requiring models to generate logically sound derivations rather than just final answers. Use when the user wants to benchmark on IneqMath, or asks about evaluating this task. Reports LLM-as-judge accuracy.
    3 repo stars
  38. ▌
    Infoseek Eval · qhjqhj00
    Evaluates LLMs on complex, multi-step reasoning and agentic search tasks, including single-hop and multi-hop question answering as well as deep research benchmarks requiring web search and synthesis. Use when the user wants to benchmark on NQ, TQA, PopQA, HQA, 2Wiki, MSQ, Bamb, BrowseComp-Plus, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  39. ▌
    Innoeval Eval · qhjqhj00
    Evaluates an AI system's ability to assess scientific research ideas across classification, selection, ranking, and comparison tasks. It probes knowledge-grounded reasoning and multi-perspective evaluation aligned with human expert judgments and conference acceptance standards. Use when the user wants to benchmark on D_point, D_group, D_pair, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  40. ▌
    Interact Eval · qhjqhj00
    Evaluates the ability of generative models to synthesize physically plausible, contact-consistent 3D human-object interaction sequences conditioned on text, actions, or object shapes. It probes motion realism, contact accuracy, and alignment between linguistic/action prompts and generated kinematics. Use when the user wants to benchmark on InterAct, or asks about evaluating this task. Reports FID.
    3 repo stars
  41. ▌
    Ioaa LLM Eval · qhjqhj00
    Evaluates LLMs' ability to solve advanced astronomy and astrophysics problems, focusing on geometric/spatial reasoning, physical calculations, and multimodal data analysis. It benchmarks performance against human Olympiad participants using official scoring rubrics. Use when the user wants to benchmark on IOAA (International Olympiad on Astronomy and Astrophysics), or asks about evaluating this task. Reports score.
    3 repo stars
  42. ▌
    Iot Nids Eval · qhjqhj00
    Evaluates the ability of a graph neural network to detect network intrusions in IoT environments by classifying traffic flows as benign or malicious, and identifying specific attack types. It probes the model's capacity to leverage topological graph structures and edge features for robust intrusion detection across imbalanced, real-world network traffic datasets. Use when the user wants to benchmark on BoT-IoT, NF-BoT-IoT, ToN-IoT, NF-ToN-IoT, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  43. ▌
    Iquad V1 Eval · qhjqhj00
    Evaluates an agent's ability to navigate, perceive, and interact with dynamic 3D environments to answer visual questions. It probes spatial reasoning, long-term memory of object locations, and the ability to plan exploration based on question semantics. Use when the user wants to benchmark on iquad v1, or asks about evaluating this task. Reports Top-1 question answering accuracy.
    3 repo stars
  44. ▌
    Irpapers Eval · qhjqhj00
    Evaluates the ability of multimodal and text-only models to retrieve relevant scientific paper pages and answer questions based on those pages. It probes retrieval depth, modality complementarity, and the impact of context quantity on RAG performance. Use when the user wants to benchmark on IRPAPERS, or asks about evaluating this task. Reports Recall@1.
    3 repo stars
  45. ▌
    Ivy Fake Eval · qhjqhj00
    This benchmark evaluates multimodal AI-generated content (AIGC) detection and explainable reasoning capabilities. It probes a model's ability to classify images and videos as real or fake, and to generate natural-language explanations that localize and justify synthetic artifacts. Use when the user wants to benchmark on Ivy-Fake, GenImage, Chameleon, GenVideo, or asks about evaluating this task. Reports Accuracy (Acc).
    3 repo stars
  46. ▌
    Jaccard Score · qhjqhj00
    Compute the jaccard_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute jaccard_score, or asks how to score with jaccard_score.
    3 repo stars
  47. ▌
    Jfin Teb Eval · qhjqhj00
    Evaluates Japanese financial text embedding models across classification, retrieval, and clustering tasks. It probes the models' ability to capture domain-specific semantics, handle regulatory and economic terminology, and generalize to zero-shot financial scenarios. Use when the user wants to benchmark on JFinTEB, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  48. ▌
    Jrdb Par Eval · qhjqhj00
    Evaluates a model's ability to jointly recognize individual actions, social group activities, and global crowd-level activities in crowded panoramic scenes. It probes multi-granular activity recognition and hierarchical graph-based scene understanding. Use when the user wants to benchmark on JRDB-PAR, or asks about evaluating this task. Reports Overall F1 ($\mathcal{F}_a$).
    3 repo stars
  49. ▌
    Kgqa4mat Eval · qhjqhj00
    Evaluates a model's ability to translate natural language questions into correct graph database queries (Cypher or SPARQL) for knowledge graph question answering. It probes the model's capacity for logical reasoning, schema understanding, and formal language generation in both a domain-specific materials science setting and a general-domain multilingual benchmark. Use when the user wants to benchmark on KGQA4MAT, QALD-9, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  50. ▌
    Kitti Fc Eval · qhjqhj00
    Evaluates the robustness of optical flow estimation models when subjected to various digital, illumination, weather, noise, and blur corruptions. It measures both absolute performance degradation and relative robustness compared to clean data across in-domain and out-of-domain training settings. Use when the user wants to benchmark on KITTI-FC, or asks about evaluating this task. Reports EPE.
    3 repo stars
  51. ▌
    Klue Dst Eval · qhjqhj00
    Evaluates a model's ability to track and update dialogue state across turns in Korean conversations, testing multi-turn reasoning and slot filling. Use when the user wants to benchmark on KLUE-DST, or asks about evaluating this task. Reports Joint Accuracy.
    3 repo stars
  52. ▌
    Klue Ner Eval · qhjqhj00
    Evaluates a model's ability to identify and classify named entities (e.g., person, location, organization) within Korean text, testing token-level understanding. Use when the user wants to benchmark on KLUE-NER, or asks about evaluating this task. Reports F1.
    3 repo stars
  53. ▌
    Laobench Eval · qhjqhj00
    Evaluates large language models on their proficiency in Lao, a low-resource Southeast Asian language. It probes factual knowledge, K12 curriculum alignment, culturally grounded reasoning, bilingual translation fidelity, and open-ended generation quality through multiple-choice, translation, and pairwise arena tasks. Use when the user wants to benchmark on LaoBench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  54. ▌
    Lar Echr Eval · qhjqhj00
    Evaluates an LLM's ability to perform legal argument reasoning by predicting the correct continuation of a court's argument chain. Given case facts and preceding arguments, the model must select the most plausible next argument from multiple options, testing its understanding of legal logic and precedent application. Use when the user wants to benchmark on LAR-ECHR, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  55. ▌
    Lextreme Eval · qhjqhj00
    Evaluates multilingual legal NLP models across text classification and named entity recognition tasks. It probes the ability of models to handle domain-specific legal jargon, long-form documents, and cross-lingual generalization across 24 languages. Use when the user wants to benchmark on LEXTREME, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  56. ▌
    Lfree Da Eval · qhjqhj00
    Evaluates a malware detection model's ability to adapt to natural concept drift over time using a rolling monthly update setup on real-world Windows malware binaries. Use when the user wants to benchmark on MB-24+, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  57. ▌
    Libricss Eval · qhjqhj00
    Evaluates continuous speech separation and speaker diarization in reverberant, multi-microphone environments with varying speaker overlap ratios. It probes the system's ability to separate overlapping speech, estimate speaker locations (DOA), cluster them across time blocks, and produce accurate diarization and speech recognition outputs. Use when the user wants to benchmark on LibriCSS, or asks about evaluating this task. Reports DER.
    3 repo stars
  58. ▌
    Limitgen Eval · qhjqhj00
    Evaluates whether LLMs can accurately identify and articulate critical limitations in scientific research papers across methodological, experimental, analytical, and literature-related dimensions. The benchmark probes the model's ability to ground critiques in domain-specific best practices and produce actionable, substantive feedback rather than superficial presentation critiques. Use when the user wants to benchmark on LimitGen, or asks about evaluating this task. Reports Limitation Quality.
    3 repo stars
  59. ▌
    Liputan6 Eval · qhjqhj00
    Evaluates the ability of models to generate concise, accurate summaries of Indonesian news articles. It probes both extractive and abstractive summarization capabilities, highlighting performance gaps on more abstract summaries and the limitations of n-gram overlap metrics. Use when the user wants to benchmark on Liputan6, or asks about evaluating this task. Reports ROUGE F-1 (R1, R2, RL).
    3 repo stars
  60. ▌
    Liveclin Eval · qhjqhj00
    LiveClin probes real-world clinical reasoning and longitudinal case management by evaluating models on dynamic, multimodal patient scenarios. It tests the ability to maintain context across sequential diagnostic, treatment, and follow-up questions while resisting data contamination from static training corpora. Use when the user wants to benchmark on LiveClin, or asks about evaluating this task. Reports Case Accuracy.
    3 repo stars
  61. ▌
    Livefact Eval · qhjqhj00
    Evaluates LLMs on fake news detection under dynamic, time-evolving evidence streams. It probes both binary classification capability (Real/Fake) and reasoning capability (handling ambiguity when evidence is incomplete), while explicitly measuring benchmark data contamination and epistemic humility. Use when the user wants to benchmark on LiveFact November 2025 dataset, or asks about evaluating this task. Reports average_score.
    3 repo stars
  62. ▌
    Llamarec Eval · qhjqhj00
    This evaluation probes a model's ability to rank candidate items based on a user's sequential interaction history and item metadata. It measures how effectively the system retrieves and re-ranks relevant products or movies against a large candidate pool using standard recommendation metrics. Use when the user wants to benchmark on ML-100k, Beauty, Games, or asks about evaluating this task. Reports NDCG@k.
    3 repo stars
  63. ▌
    Llase G1 Eval · qhjqhj00
    Evaluates a LLaMA-based generative speech enhancement model's ability to perform multiple audio restoration tasks (noise suppression, packet loss concealment, target speaker extraction, acoustic echo cancellation, and speech separation) in a task-agnostic manner. It probes the model's capacity to preserve acoustic fidelity and semantic content while generalizing across different acoustic conditions and device types. Use when the user wants to benchmark on DNS Challenge blind test set (Interspeech 2020), ICASSP 2022 PLC-challenge blind test set, ICASSP 2023 DNS blind test set, Libri2mix test set, WSJ0_2mix test set, or asks about evaluating this task. Reports DNSMOS OVRL.
    3 repo stars
  64. ▌
    Llava Le Eval · qhjqhj00
    Evaluates multimodal instruction-following and complex geological reasoning on lunar surface imagery. It probes the model's ability to interpret crater morphology, degradation states, and inferred geological processes beyond simple visual description. Use when the user wants to benchmark on LUCID (held-out eval set), or asks about evaluating this task. Reports average_overall_score.
    3 repo stars
  65. ▌
    LLM Bias Eval · qhjqhj00
    This evaluation probes the propensity of large language models to generate stereotypical or biased predictions across multiple demographic and social categories. It measures how often models align with human-annotated stereotypes versus anti-stereotypes or neutral alternatives when completing masked sentences or answering repurposed benchmark questions. Use when the user wants to benchmark on StereoSet, WinoBias, UnQover, CrowS-Pairs, Real Toxicity Prompts (RTP), Equity Evaluation Corpus (EEC), or asks about evaluating this task. Reports bias_intensity.
    3 repo stars
  66. ▌
    Llmjudge Eval · qhjqhj00
    This benchmark evaluates the agreement and ranking consistency of LLM-generated relevance judgments against human assessments. It probes whether automated scoring methods can reliably replicate human relevance labels and maintain correct document ordering for information retrieval tasks. Use when the user wants to benchmark on LLMJudge test set, or asks about evaluating this task. Reports Cohen's \kappa.
    3 repo stars
  67. ▌
    Lnmbench Eval · qhjqhj00
    Evaluates the robustness of label noise learning (LNL) methods on medical image classification tasks. It probes model performance under varying noise types (symmetric, instance-dependent, real-world), noise ratios, and class imbalance distributions across multiple imaging modalities. Use when the user wants to benchmark on PathMNIST, DermaMNIST, BloodMNIST, OrganCMNIST, DRTiD, Kaggle DR+, CheXpert, or asks about evaluating this task. Reports average classification accuracy.
    3 repo stars
  68. ▌
    Locomo10 Eval · qhjqhj00
    This benchmark evaluates long-term conversational memory systems by testing their ability to retrieve relevant dialogue turns and answer questions over extended, multi-session histories. It probes semantic reasoning, temporal tracking, and adversarial robustness across five distinct question categories. Use when the user wants to benchmark on LoCoMo10, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  69. ▌
    Logiccat Eval · qhjqhj00
    This benchmark evaluates an LLM's ability to perform complex multi-step logical reasoning and chain-of-thought decomposition to generate correct SQL queries from natural language questions. It probes capabilities in mathematical deduction, physical knowledge integration, and cross-domain analytical querying. Use when the user wants to benchmark on LogicCat, or asks about evaluating this task. Reports Execution Accuracy (EX).
    3 repo stars
  70. ▌
    Long Ner Eval · qhjqhj00
    Evaluates named entity recognition capabilities, specifically probing a model's ability to handle class imbalance, out-of-vocabulary terms, and long or complex entity names across biomedical and general domain texts. Use when the user wants to benchmark on NCBI-disease, BC5CDR-disease, BC5CDR-chemical, BC4CHEMD, BC2GM, JNLPBA, LINNAEUS, Species-800, CoNLL-2003, WNUT-2017, or asks about evaluating this task. Reports F1.
    3 repo stars
  71. ▌
    Longlamp Eval · qhjqhj00
    Evaluates a model's ability to generate personalized long-form text by integrating retrieved user profiles into a retrieval-augmented generation framework. It probes how well models can adapt their writing style and content to specific user attributes across different domains like emails, abstracts, reviews, and topic-based writing. Use when the user wants to benchmark on LongLaMP, or asks about evaluating this task. Reports METEOR.
    3 repo stars
  72. ▌
    Loopserv Eval · qhjqhj00
    Evaluates an adaptive dual-phase LLM inference acceleration system for multi-turn dialogues, probing its ability to maintain generation accuracy and reduce computational overhead across varying query positions in long-context conversations. It specifically tests whether the system can generalize beyond positional heuristics used by static KV cache compression methods. Use when the user wants to benchmark on MFQA-en, 2WikiMQA, Musique, HotpotQA, NrtvQA, Qasper, MultiNews, GovReport, QMSum, TREC, SAMSUM, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  73. ▌
    M Portal Eval · qhjqhj00
    Evaluates multi-step spatial and physical reasoning in multimodal language models by requiring them to generate or validate chain-of-thought plans for solving Portal 2-inspired puzzle maps. The benchmark probes the model's ability to integrate visual map layouts with textual instructions to produce physically sound, multi-step traversal strategies. Use when the user wants to benchmark on M-Portal, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  74. ▌
    Macbench Eval · qhjqhj00
    Evaluates vision-language models' ability to perform multimodal scientific reasoning in chemistry and materials research. It probes capabilities across data extraction, experimental understanding, and data interpretation, specifically testing spatial reasoning, cross-modal synthesis, and multi-step inference. Use when the user wants to benchmark on MaCBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  75. ▌
    Maptrace Eval · qhjqhj00
    Evaluates fine-grained spatial reasoning and pixel-accurate route tracing on commercial map images. Models must generate precise path coordinates or masks corresponding to text-based navigation queries. Use when the user wants to benchmark on MapTrace, Map-Bench, or asks about evaluating this task. Reports NDTW.
    3 repo stars
  76. ▌
    Matbench Eval · qhjqhj00
    Evaluates machine learning models on predicting materials properties (e.g., elastic moduli, band gaps, formation energies) from crystal structures or compositions. It probes generalization across diverse data sizes, input types, and property domains using a standardized, pre-cleaned suite of 13 supervised tasks. Use when the user wants to benchmark on Matbench test suite v0.1, or asks about evaluating this task. Reports error estimation (RMSE/Accuracy).
    3 repo stars
  77. ▌
    Mathreal Eval · qhjqhj00
    Evaluates the ability of multimodal large language models to perform K-12 mathematical reasoning on real-world, mobile-captured images. It probes robustness to visual degradation (blur, rotation, handwritten annotations) and perspective variations, measuring how well models extract text and figures to solve math problems under imperfect conditions. Use when the user wants to benchmark on MathReal, or asks about evaluating this task. Reports Loose Accuracy (Acc).
    3 repo stars
  78. ▌
    Matsciml Eval · qhjqhj00
    Evaluates graph neural networks and equivariant point cloud networks on solid-state materials modeling tasks, including energy/force prediction, bandgap/fermi level regression, and crystal symmetry classification. Probes single-task, multi-task, and multi-dataset generalization capabilities. Use when the user wants to benchmark on OpenCatalyst (OC-20), Materials Project (MP), LiPS, OQMD, NOMAD, CMD, or asks about evaluating this task. Reports MSE.
    3 repo stars
  79. ▌
    Mavos Dd Eval · qhjqhj00
    Evaluates deepfake detection models on distinguishing real from fake audio-video content under closed-set and open-set conditions. It specifically probes cross-model and cross-lingual generalization by testing on unseen generation methods and languages. Use when the user wants to benchmark on MAVOS-DD, or asks about evaluating this task. Reports mAP.
    3 repo stars
  80. ▌
    Mc Bench Eval · qhjqhj00
    This benchmark evaluates multi-context visual grounding, requiring models to localize target objects across multiple images using open-ended, context-rich text prompts. It probes cross-image reasoning, fine-grained instance localization, and the ability to correctly group and reject irrelevant instances. Use when the user wants to benchmark on MC-Bench, or asks about evaluating this task. Reports AP50.
    3 repo stars
  81. ▌
    Mcts Vcb Eval · qhjqhj00
    Evaluates multimodal large language models on fine-grained video captioning by measuring how well generated descriptions capture verified, detailed key points from videos. It probes the model's ability to reason about and describe specific visual details like color, quantity, and position. Use when the user wants to benchmark on MCTS-VCB, or asks about evaluating this task. Reports F1.
    3 repo stars
  82. ▌
    Mdlbench Eval · qhjqhj00
    Evaluates the inference performance and overhead of six major mobile deep learning libraries across diverse model architectures, tasks, and mobile hardware configurations. It measures how software-level optimizations and library choices impact on-device latency compared to hardware capabilities and algorithmic optimizations like quantization. Use when the user wants to benchmark on MDLBench, or asks about evaluating this task. Reports inference time.
    3 repo stars
  83. ▌
    Mdpbench Eval · qhjqhj00
    MDPBench evaluates the capability of document parsing models to accurately extract text, formulas, tables, and layout structures from multilingual document images under real-world conditions. It specifically probes robustness to photographic degradation, non-Latin scripts, right-to-left reading orders, and cross-lingual generalization without prior language or image-type knowledge. Use when the user wants to benchmark on MDPBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  84. ▌
    Medbench Eval · qhjqhj00
    This benchmark evaluates Chinese large language models on clinical knowledge, diagnostic reasoning, and conversational ability. It probes performance across three standardized medical licensing exams and real-world clinical case scenarios, highlighting gaps in multi-hop reasoning, diagnostic precision, and response fluency. Use when the user wants to benchmark on MedBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  85. ▌
    Medg Krp Eval · qhjqhj00
    Evaluates large language models' ability to generate causal knowledge graphs from single medical concepts. It probes biomedical reasoning, causal understanding, and factual consistency by comparing generated graphs against human expert judgments and a ground-truth biomedical ontology (BIOS). Use when the user wants to benchmark on MedG-KRP, or asks about evaluating this task. Reports Precision.
    3 repo stars
  86. ▌
    Medi Aug Eval · qhjqhj00
    Evaluates the impact of six mix-based data augmentation methods on medical image classification performance. It probes how well models generalize to medical domains (brain MRI and eye fundus) when trained with different augmentation strategies and backbones. Use when the user wants to benchmark on Brain Tumor Classification Dataset, Eye Diseases Classification Dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  87. ▌
    Mediasum Eval · qhjqhj00
    Evaluates abstractive dialogue summarization models on long-form, multi-speaker media interviews. It measures how well models can condense multi-topic conversations into concise summaries, and assesses transfer learning capabilities to other dialogue domains. Use when the user wants to benchmark on MediaSum, or asks about evaluating this task. Reports ROUGE-1, ROUGE-2, ROUGE-L F1.
    3 repo stars
  88. ▌
    Medq Deg Eval · qhjqhj00
    Evaluates multimodal large language models' robustness and metacognitive reliability when processing medical images with various quality degradations (e.g., blur, noise, motion, artifacts) across different clinical capability dimensions. Use when the user wants to benchmark on MedQ-Deg, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  89. ▌
    Medrcube Eval · qhjqhj00
    Probes multimodal large language models' fine-grained capabilities in medical imaging across anatomical regions, imaging modalities, and cognitive hierarchies. It assesses reasoning reliability, shortcut behavior, and foundational perceptual skills to reveal how models handle clinical VQA beyond aggregate performance. Use when the user wants to benchmark on MedRCube, or asks about evaluating this task. Reports MedRCube Score.
    3 repo stars
  90. ▌
    Megawika Eval · qhjqhj00
    Evaluates the quality and evidential support of automatically generated Wikipedia question-answer pairs and their cited source documents. It probes whether sources actually contain the information claimed in passages, and measures the strength of support for QA pairs. Use when the user wants to benchmark on MegaWika, or asks about evaluating this task. Reports answerability.
    3 repo stars
  91. ▌
    Messirve Eval · qhjqhj00
    Evaluates information retrieval models on a large-scale, dialectally diverse Spanish dataset. It probes the ability of lexical and dense retrieval models to rank relevant Wikipedia documents for real-world Spanish search queries without fine-tuning. Use when the user wants to benchmark on MessIRve, or asks about evaluating this task. Reports nDCG@10.
    3 repo stars
  92. ▌
    Mets Cov Eval · qhjqhj00
    Evaluates named entity recognition capabilities on biomedical and social media text. It probes the model's ability to identify and classify domain-specific entities (e.g., diseases, drugs, vaccines) in both informal tweets and formal scientific abstracts under fully-supervised and few-shot learning conditions. Use when the user wants to benchmark on METS-CoV, BioRED, or asks about evaluating this task. Reports Micro F1.
    3 repo stars
  93. ▌
    Mfmdqwen Eval · qhjqhj00
    Evaluates large language models on multilingual financial misinformation detection across nine tasks in English, Chinese, Greek, and Bengali. It probes the model's ability to identify false or misleading financial claims, handle numerical sensitivity and reversed causality, and generalize across diverse linguistic settings. Use when the user wants to benchmark on MFMDBench, or asks about evaluating this task. Reports Macro-F1.
    3 repo stars
  94. ▌
    Microisp Eval · qhjqhj00
    Evaluates a deep learning-based image signal processing (ISP) model's ability to reconstruct high-resolution RGB images from RAW mobile sensor data. It measures reconstruction fidelity, visual quality, and inference efficiency across various mobile hardware platforms and resolutions. Use when the user wants to benchmark on Fujifilm UltraISP, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  95. ▌
    Mimicsql Eval · qhjqhj00
    Evaluates a model's ability to generate correct SQL queries from natural language clinical questions in the healthcare domain. It specifically probes retrieval-based reasoning, robustness to noisy or ambiguous medical terminology, and generalization under data scarcity conditions. Use when the user wants to benchmark on MIMICSQL, or asks about evaluating this task. Reports Execution Accuracy.
    3 repo stars
  96. ▌
    Mimii Dg Eval · qhjqhj00
    Evaluates domain generalization capabilities for anomalous sound detection by measuring how well models trained on a source domain of industrial machine sounds generalize to target domains with shifted operational parameters or background noise. Use when the user wants to benchmark on MIMII DG, or asks about evaluating this task. Reports AUC.
    3 repo stars
  97. ▌
    Mimotion Eval · qhjqhj00
    Evaluates models on predicting future 3D multi-person motion sequences given a short history of interacting subjects. It probes spatial-temporal modeling, interaction awareness, and long-horizon trajectory forecasting under varying scene complexities and prediction horizons. Use when the user wants to benchmark on MI-Motion, or asks about evaluating this task. Reports GJPE, AJPE, RFDE.
    3 repo stars
  98. ▌
    Mind2web Eval · qhjqhj00
    Evaluates a model's ability to act as a generalist web agent by completing multi-step tasks across diverse, unseen websites and domains. It probes out-of-distribution generalization, web element grounding, and sequential action planning in real-world browser environments. Use when the user wants to benchmark on Mind2Web, or asks about evaluating this task. Reports Step Success Rate.
    3 repo stars
  99. ▌
    Minimwob Eval · qhjqhj00
    Tests a visual GUI agent's ability to complete web automation tasks by interacting with simplified web environments based on screenshots. It measures task completion success rates across various interactive web widgets. Use when the user wants to benchmark on MiniWob, or asks about evaluating this task. Reports success rate.
    3 repo stars
  100. ▌
    Miroeval Eval · qhjqhj00
    Evaluates multimodal deep research agents on both the quality of their final synthesized reports and the underlying investigative process. It measures adaptive synthesis quality, factual grounding against heterogeneous sources, and process-centric attributes like search breadth, analytical depth, and alignment between intermediate findings and the final report. Use when the user wants to benchmark on MiroEval, or asks about evaluating this task. Reports Adaptive Synthesis Quality (S_quality).
    3 repo stars