all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 57 of 76

  1. ▌
    Nim4 Asr Eval · qhjqhj00
    Evaluates automatic speech recognition performance across diverse acoustic and linguistic domains, including English, Mandarin, dialects, code-switching, and in-car conversational scenarios. It measures transcription accuracy and hallucination rates to assess model robustness, latency, and customization capabilities. Use when the user wants to benchmark on LibriSpeech, VoxPopuli, MLS-English, AISHELL-1, AISHELL-2, AISHELL-2021-Eval, WeNetSpeech, SpeechIO, WeNetSpeech-Chuan, WeNetSpeech-Yue, KeSpeech, CS-Dialogue, ASCEND, M4Singer, Internal POI Benchmarks, Internal Media Benchmarks, Internal Device Control, Internal Conversational, or asks about evaluating this task. Reports WER, CER.
    3 repo stars
  2. ▌
    Noisyner Eval · qhjqhj00
    Evaluates the robustness of noise models and base models under realistic noisy label conditions, measuring how estimation accuracy and base model performance vary with different noise distributions and amounts of clean data. Use when the user wants to benchmark on NoisyNER, or asks about evaluating this task. Reports micro-average F1 score.
    3 repo stars
  3. ▌
    Nuinsseg Eval · qhjqhj00
    Evaluates the ability of models to perform instance-level segmentation of cell nuclei in H&E-stained histological images. It specifically probes robustness to tissue variability, high cell density, and ambiguous boundaries where manual annotation is difficult. Use when the user wants to benchmark on NuInsSeg, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  4. ▌
    Nuplan R Eval · qhjqhj00
    Evaluates autonomous driving planners in a closed-loop setting with realistic, reactive multi-agent traffic. It probes a planner's ability to handle complex, interactive driving scenarios by measuring overall success, robustness against catastrophic failures, and consistency across safety and comfort dimensions. Use when the user wants to benchmark on nuPlan-R, or asks about evaluating this task. Reports CLS.
    3 repo stars
  5. ▌
    Odysseys Eval · qhjqhj00
    Probes an agent's ability to perform realistic, long-horizon web navigation tasks that require sustained cross-site reasoning, context maintenance across multiple tabs, and efficient action execution. It evaluates whether models can complete complex, multi-step user journeys derived from real browsing behavior within strict step budgets. Use when the user wants to benchmark on Odysseys, or asks about evaluating this task. Reports Perfect Rubrics (%).
    3 repo stars
  6. ▌
    Oem Gfss Eval · qhjqhj00
    Evaluates a model's ability to perform generalized few-shot semantic segmentation on remote sensing imagery. It tests whether a model can accurately segment both previously seen (base) and new (novel) land cover classes simultaneously using only a few support examples (5-shot), probing generalization and resistance to class forgetting in low-data regimes. Use when the user wants to benchmark on OEM-GFSS, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  7. ▌
    Olymmath Eval · qhjqhj00
    Evaluates advanced mathematical reasoning capabilities on Olympiad-level problems. It probes a model's ability to perform rigorous, step-by-step logical deduction and numerical verification across algebra, geometry, number theory, and combinatorics. The benchmark also assesses cross-lingual reasoning performance between English and Chinese. Use when the user wants to benchmark on OlymMATH, or asks about evaluating this task. Reports Pass@1.
    3 repo stars
  8. ▌
    Omibench Eval · qhjqhj00
    Evaluates large vision-language models on Olympiad-level multi-image reasoning tasks across biology, chemistry, mathematics, and physics. It probes the model's ability to integrate complementary visual and textual evidence across multiple images to generate stepwise rationales and select or produce correct final answers. Use when the user wants to benchmark on OMIBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  9. ▌
    Omni Dpo Eval · qhjqhj00
    Evaluates the instruction-following and mathematical reasoning capabilities of LLMs fine-tuned with a dual-perspective preference optimization method. It measures conversational quality, adherence to instructions, and problem-solving accuracy across diverse open-ended and quantitative benchmarks. Use when the user wants to benchmark on AlpacaEval 2.0, Arena-Hard v0.1, IFEval, SedarEval, GSM8K, MATH 500, AIME 2024, AMC 2023, or asks about evaluating this task. Reports LC(%).
    3 repo stars
  10. ▌
    Omnifall Eval · qhjqhj00
    Evaluates human fall detection and action recognition capabilities across controlled (staged) and uncontrolled (wild) video domains. It probes a model's ability to classify a 10-class activity taxonomy, detect binary fall/fallen states, and segment action timelines, while measuring generalization gaps between in-distribution and out-of-distribution settings. Use when the user wants to benchmark on CMDFall, UP-Fall, Le2i, GMDCSA24, EDF, OCCU, CaucaFall, MCFD, OOPS-Fall, or asks about evaluating this task. Reports Balanced Accuracy, Macro F1.
    3 repo stars
  11. ▌
    Omniflow Eval · qhjqhj00
    Evaluates optical flow estimation models on synthetic omnidirectional human motion data. It probes the model's ability to handle fisheye distortions, domain-randomized environments, and varying amounts of fine-tuning data. Use when the user wants to benchmark on OmniFlow, or asks about evaluating this task. Reports optical flow error.
    3 repo stars
  12. ▌
    Omnigen2 Eval · qhjqhj00
    Evaluates a unified multimodal model's capabilities across visual understanding, text-to-image generation, instruction-based image editing, and in-context generation. It probes compositional prompt following, long-prompt adherence, edit accuracy versus preservation, and subject consistency across single, multiple, and scene contexts. Use when the user wants to benchmark on MMBench, MMMU, MM-Vet, GenEval, DPG-Bench, Emu-Edit, GEdit-Bench-EN, ImgEdit-Bench, OmniContext, or asks about evaluating this task. Reports GenEval Overall, DPG-Bench Overall.
    3 repo stars
  13. ▌
    Open Sdi Eval · qhjqhj00
    Evaluates a model's ability to localize manipulated regions in images generated by diffusion models, measuring both pixel-level segmentation precision and image-level binary detection capability across multiple unseen generators. Use when the user wants to benchmark on OpenSDI, or asks about evaluating this task. Reports pixel-level F1.
    3 repo stars
  14. ▌
    Openfake Eval · qhjqhj00
    Binary classification capability for detecting AI-generated images versus real photographs. It probes a model's ability to generalize across diverse generative models (diffusion, transformer-based) and real-world social media distributions. Use when the user wants to benchmark on OpenFake, or asks about evaluating this task. Reports F1 Score.
    3 repo stars
  15. ▌
    Openscan Eval · qhjqhj00
    Evaluates 3D vision models' ability to perform open-vocabulary scene understanding by querying for fine-grained object attributes (e.g., material, affordance, synonym) rather than standard object classes. It measures both 3D instance segmentation and 3D semantic segmentation capabilities on attribute-based queries across eight linguistic aspects. Use when the user wants to benchmark on OpenScan, or asks about evaluating this task. Reports AP, mIoU.
    3 repo stars
  16. ▌
    Optbench Eval · qhjqhj00
    Evaluates formal theorem proving capabilities specifically within the undergraduate optimization domain. It probes a model's ability to generate syntactically correct and semantically progressive Lean 4 proof steps or full scripts under strict verifier constraints, while measuring robustness against catastrophic forgetting on general math benchmarks. Use when the user wants to benchmark on OptBench, MiniF2F-test, ProofNet-test, or asks about evaluating this task. Reports Pass@32.
    3 repo stars
  17. ▌
    Orgforge Eval · qhjqhj00
    This benchmark evaluates RAG and retrieval-augmented agents on synthetic corporate corpora by testing their ability to retrieve artifacts, reason over causal and temporal chains, and detect knowledge gaps. It probes multi-hop reasoning, temporal knowledge-state tracking, and absence-of-evidence detection across structured enterprise artifacts like Slack, JIRA, and emails. Use when the user wants to benchmark on OrgForge Synthetic Corporate Corpus, or asks about evaluating this task. Reports OrgForgeScorer.
    3 repo stars
  18. ▌
    Orsi Sod Eval · qhjqhj00
    This benchmark evaluates optical remote sensing salient object detection models by measuring their ability to accurately segment prominent objects from complex, cluttered backgrounds. It probes structural consistency, boundary precision, and error magnitude across varying object scales and scene complexities. Use when the user wants to benchmark on ORSSD, EORSSD, ORSI-4199, or asks about evaluating this task. Reports maximum F-measure ($F_{\beta}^{max}$).
    3 repo stars
  19. ▌
    Oscbench Eval · qhjqhj00
    This benchmark probes a text-to-video model's ability to accurately render and maintain object state transformations (e.g., peeling, slicing) over time. It additionally measures semantic adherence to prompts, scene consistency, and overall perceptual quality to diagnose temporal coherence and physical realism in generated videos. Use when the user wants to benchmark on OSCBench, or asks about evaluating this task. Reports state-change accuracy.
    3 repo stars
  20. ▌
    Panmatch Eval · qhjqhj00
    Evaluates a unified vision model's ability to perform correspondence matching across stereo disparity estimation, optical flow, and feature matching in a zero-shot setting. It probes cross-domain generalization and robustness to challenging conditions like occlusion, lighting changes, and non-Lambertian surfaces. Use when the user wants to benchmark on Middlebury, ETH3D, KITTI, Infinigen, Spring, Sintel, Booster, or asks about evaluating this task. Reports PCA x.
    3 repo stars
  21. ▌
    Papillon Eval · qhjqhj00
    Evaluates a privacy-preserving LLM delegation pipeline that sanitizes user queries before sending them to a remote API model. It measures the trade-off between maintaining response quality and minimizing personally identifiable information (PII) leakage in the sanitized prompts. Use when the user wants to benchmark on PUPA-TNB, or asks about evaluating this task. Reports QUAL.
    3 repo stars
  22. ▌
    Pariksha Eval · qhjqhj00
    Evaluates multilingual and multi-cultural LLM performance across 10 Indic languages using culturally nuanced prompts. It measures model quality via pairwise comparisons (Elo ratings) and direct assessment scores, while also analyzing human-LLM evaluator agreement and various biases (position, verbosity, self-bias). Use when the user wants to benchmark on PARIKSHA, or asks about evaluating this task. Reports Elo rating, Direct Assessment score.
    3 repo stars
  23. ▌
    Parsinlu Eval · qhjqhj00
    Evaluates Persian language understanding across six distinct NLU tasks, including reading comprehension, textual entailment, sentiment analysis, and machine translation. It measures how well pre-trained monolingual and multilingual models perform on native-speaker annotated Persian data compared to human baselines. Use when the user wants to benchmark on ParsiNLU, or asks about evaluating this task. Reports F1, Accuracy.
    3 repo stars
  24. ▌
    Patenteb Eval · qhjqhj00
    Evaluates patent text embedding models across 15 diverse tasks including symmetric/asymmetric retrieval, classification, paraphrase detection, and clustering. It specifically probes domain-specific challenges like cross-domain retrieval, fragment-to-document matching, and temporal citation dynamics. Use when the user wants to benchmark on PatenTEB, or asks about evaluating this task. Reports NDCG@10, Macro-F1, Pearson r, V-measure.
    3 repo stars
  25. ▌
    Perla 3d Eval · qhjqhj00
    Evaluates a 3D vision-language model's ability to answer questions about indoor scenes and generate dense captions for 3D instances. It probes fine-grained spatial reasoning, object attribute recognition, and scene understanding by comparing generated text against ground-truth annotations using standard natural language generation metrics. Use when the user wants to benchmark on ScanNet, or asks about evaluating this task. Reports CiDEr.
    3 repo stars
  26. ▌
    Phase No Eval · qhjqhj00
    Evaluates the ability of neural phase pickers to detect P- and S-wave arrivals in continuous seismic waveforms across multi-station networks. It probes detection accuracy, timing precision, and generalization to out-of-distribution earthquake sequences under varying signal-to-noise conditions. Use when the user wants to benchmark on NCEDC 2020 Test Set, 2019 Ridgecrest Sequence, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  27. ▌
    Pheno Ca Eval · qhjqhj00
    Evaluates deep neural networks on phenotypic drug discovery tasks using high-content screening images. It probes the model's ability to deconvolve mechanisms of action, molecular targets, and compound identities from cellular phenotypes, as well as zero-shot compound retrieval for CRISPR perturbations. Use when the user wants to benchmark on Pheno-CA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  28. ▌
    Phibench Eval · qhjqhj00
    Evaluates diverse reasoning and coding capabilities using an internal benchmark designed to minimize data contamination and LLM-judge bias. It probes a model's ability to debug, extend, and explain code, as well as identify errors in mathematical proofs and generate related problems. Use when the user wants to benchmark on PhiBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  29. ▌
    Physgame Eval · qhjqhj00
    This benchmark evaluates a model's ability to reason about physical laws and detect physical commonsense violations in gameplay videos. It probes spatial, temporal, and meta-information-based physical reasoning through curated multi-choice questions. Use when the user wants to benchmark on PhysGame, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  30. ▌
    Piibench Eval · qhjqhj00
    This benchmark probes the ability of NER and PII detection systems to accurately identify and classify personally identifiable information spans across highly heterogeneous, cross-domain text sources. It specifically evaluates cross-domain generalization and robustness to diverse, fine-grained PII entity types that are rarely seen together in standard training corpora. Use when the user wants to benchmark on PIIBench, or asks about evaluating this task. Reports span-level F1.
    3 repo stars
  31. ▌
    Pinball Score · qhjqhj00
    Evaluates the sharpness and calibration of probabilistic net-load forecasts. It measures how closely predicted quantiles align with actual observations and how narrow the prediction intervals are while maintaining statistical reliability. Use when the user has predictions and gold and needs to compute Pinball Score.
    3 repo stars
  32. ▌
    Pixelrec Eval · qhjqhj00
    Evaluates the ability of recommender systems to rank items using raw pixel images instead of traditional ID embeddings. It probes cold-start item recommendation, cross-domain transfer learning, and end-to-end vision-based recommendation performance. Use when the user wants to benchmark on PixelRec, or asks about evaluating this task. Reports Recall@N.
    3 repo stars
  33. ▌
    Plantseg Eval · qhjqhj00
    Evaluates pixel-level segmentation capabilities for identifying and localizing plant diseases in real-world, uncontrolled agricultural imagery across 115 disease classes. The benchmark tests a model's ability to handle fine-grained lesion boundaries, overlapping disease symptoms, and high visual diversity typical of field-captured crops. Use when the user wants to benchmark on PlantSeg, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  34. ▌
    Pmlbmini Eval · qhjqhj00
    This benchmark evaluates the discriminative performance of machine learning models on tabular classification tasks under data-scarce conditions (sample sizes ≤500). It specifically probes whether complex AutoML and deep learning approaches can consistently outperform simple baselines like logistic regression when training data is limited. Use when the user wants to benchmark on PMLBmini, or asks about evaluating this task. Reports AUC.
    3 repo stars
  35. ▌
    Pnm Flow Eval · qhjqhj00
    Evaluates a topology-based pore network model's ability to predict flow-permeable surface area and hydraulic conductance in granular materials from micro-CT images. Use when the user wants to benchmark on Sphere Packing & High-Explosive Micro-CT Samples, or asks about evaluating this task. Reports conductance_ratio.
    3 repo stars
  36. ▌
    Polymath Eval · qhjqhj00
    Evaluates multi-modal mathematical and cognitive reasoning capabilities on visual puzzles. It probes spatial interpretation, relational understanding, pattern recognition, and long-horizon logical reasoning using diagram-based multiple-choice questions. Use when the user wants to benchmark on POLYMATH, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  37. ▌
    Pralekha Eval · qhjqhj00
    Evaluates cross-lingual document alignment (CLDA) techniques by measuring chunk/sentence-level alignment accuracy (intrinsic) and the resulting document-level machine translation quality (extrinsic) across English and 11 Indic languages. Use when the user wants to benchmark on Pralekha, or asks about evaluating this task. Reports F1 Score, DocCOMET.
    3 repo stars
  38. ▌
    Prefeval Eval · qhjqhj00
    Measures an agent's capability to retain and adhere to user preferences during long, multi-turn conversations. It tests whether the model can maintain consistency without external reminders or with explicit preference cues. Use when the user wants to benchmark on PrefEval, or asks about evaluating this task. Reports preference retention accuracy.
    3 repo stars
  39. ▌
    Prime Dp Eval · qhjqhj00
    Evaluates a pre-trained seismic model's multi-task capability on single-station waveforms, specifically phase picking (Pg, Sg, Pn, Sn), P-wave polarization classification, and seismic event type classification. The protocol tests generalization across temporal splits and transfer learning on local data to mitigate dataset imbalance. Use when the user wants to benchmark on CSNCD, or asks about evaluating this task. Reports recall.
    3 repo stars
  40. ▌
    Primesrl Eval · qhjqhj00
    Evaluates the quality of Semantic Role Labeling (SRL) systems by measuring precision and recall for predicate senses and argument annotations. It specifically probes a model's step-dependent error propagation by penalizing argument scores when the associated predicate sense is incorrect, while also handling discontinuous and reference arguments. Use when the user has predictions and gold and needs to compute PriMeSRL-Eval.
    3 repo stars
  41. ▌
    Proofnet Eval · qhjqhj00
    Evaluates a model's ability to translate between natural language mathematics and Lean 3 formal statements (autoformalization and informalization). It measures syntactic validity, semantic correctness, and lexical similarity to assess reasoning over undergraduate-level theory. Use when the user wants to benchmark on ProofNet, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  42. ▌
    Pubmedqa Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform biomedical research question answering by reasoning over structured scientific abstracts. It requires models to infer yes/no/maybe answers to questions derived from paper titles using only the non-conclusion sections of the abstract, without access to the final conclusion. Use when the user wants to benchmark on PubMedQA (PQA-L), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  43. ▌
    QA Fb15k Eval · qhjqhj00
    Evaluates cognition-based hallucination by testing whether LVLMs can leverage world knowledge stored in the LLM to answer entity and relation questions grounded in images. Use when the user wants to benchmark on QA-FB15K, or asks about evaluating this task. Reports Acc.
    3 repo stars
  44. ▌
    Qcaleval Eval · qhjqhj00
    Probes vision-language models' ability to interpret quantum calibration plots across diverse visual formats (1D traces, 2D maps, histograms) and perform structured scientific reasoning. It tests capabilities ranging from visual grounding and outcome classification to parameter extraction and operational calibration diagnosis, both in zero-shot and in-context learning settings. Use when the user wants to benchmark on QCalEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  45. ▌
    Qg Bench Eval · qhjqhj00
    Evaluates the ability of generative language models to produce paragraph-level questions conditioned on a target answer and a context sentence. It probes domain adaptability and multilingual generalization across diverse extractive QA datasets. Use when the user wants to benchmark on SQuAD v1.1, SQuADShifts, SubjQA, Multilingual QA (JAQuAD, GerQuAD, SberQuAD, KorQuAD, FQuAD, Spanish SQuAD, Italian SQuAD), or asks about evaluating this task. Reports automatic evaluation metrics.
    3 repo stars
  46. ▌
    Quantemp Eval · qhjqhj00
    Evaluates a model's ability to fact-check real-world numerical claims containing statistical and temporal expressions. It probes evidence retrieval, claim decomposition, and natural language inference to predict veracity (True, False, or Conflicting). Use when the user wants to benchmark on NumTemp, or asks about evaluating this task. Reports Macro-F1.
    3 repo stars
  47. ▌
    R U Maad Eval · qhjqhj00
    Evaluates a model's ability to perform unsupervised anomaly detection on multi-agent traffic trajectories in urban environments. It probes frame-wise recognition of rare and abnormal driving behaviors, including both individual maneuvers and context-dependent interactions between agents and static map features. Use when the user wants to benchmark on R-U-MAAD, or asks about evaluating this task. Reports frame-wise anomaly detection accuracy.
    3 repo stars
  48. ▌
    Re Verse Eval · qhjqhj00
    Evaluates vision-language models' ability to comprehend long-form sequential manga narratives, focusing on story synthesis, character grounding, and temporal reasoning across non-linear, multi-panel sequences. Use when the user wants to benchmark on Re:Verse, or asks about evaluating this task. Reports BERTScore.
    3 repo stars
  49. ▌
    Real Iad Eval · qhjqhj00
    Evaluates industrial anomaly detection models under standard unsupervised and fully unsupervised (noisy training) settings. It probes image-level, pixel-level, and multi-view sample-level defect detection capabilities. Use when the user wants to benchmark on Real-IAD, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  50. ▌
    Redbench Eval · qhjqhj00
    Evaluates LLM robustness against adversarial prompts (Attack Success Rate) and their tendency to over-defend on benign prompts (Rejection Rate). It probes safety alignment, refusal behavior, and cross-domain vulnerability across 22 risk categories and 19 domains. Use when the user wants to benchmark on RedBench, or asks about evaluating this task. Reports Rejection Rate (RR), Attack Success Rate (ASR).
    3 repo stars
  51. ▌
    Redis QA Eval · qhjqhj00
    Evaluates large language models' ability to answer questions about rare diseases, including diagnosis, symptoms, causes, and related properties. It probes the models' medical knowledge retrieval and reasoning capabilities in a specialized, low-resource domain. Use when the user wants to benchmark on ReDis-QA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  52. ▌
    Refcocom Eval · qhjqhj00
    Evaluates a model's ability to perform referring expression segmentation at both object and part levels. It probes fine-grained cross-modal alignment and pixel-level semantic understanding by requiring precise mask prediction for diverse textual references. Use when the user wants to benchmark on RefCOCOm, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  53. ▌
    Relbench Eval · qhjqhj00
    Evaluates the ability of graph neural networks to learn from relational databases by predicting entity attributes (classification and regression) and discovering relationships between entities (link prediction) using primary-foreign key graph structures. Use when the user wants to benchmark on RELBENCH, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  54. ▌
    Replaydf Eval · qhjqhj00
    Evaluates the robustness of audio deepfake detection models against replay attacks where deepfake audio is played back and re-recorded through real-world hardware, introducing acoustic distortions and room impulse responses. Use when the user wants to benchmark on ReplayDF, or asks about evaluating this task. Reports EER (%).
    3 repo stars
  55. ▌
    Reviewmt Eval · qhjqhj00
    Evaluates LLMs on simulating the academic peer review process across multi-turn dialogues. It probes the model's ability to generate relevant paper summaries, write comprehensive reviews, and make accurate acceptance or rejection decisions based on long-context interactions. Use when the user wants to benchmark on ReviewMT, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  56. ▌
    Rexbench Eval · qhjqhj00
    This benchmark evaluates the ability of LLM-based coding agents to autonomously implement research extensions by modifying existing AI/ML codebases based on domain-expert instructions. It probes complex, multi-step software engineering capabilities, including codebase navigation, hypothesis-driven implementation, and producing executable patches. Use when the user wants to benchmark on REXBench, or asks about evaluating this task. Reports final success rate.
    3 repo stars
  57. ▌
    Rm Bench Eval · qhjqhj00
    Evaluates reward models' ability to correctly identify preferred responses based on substantive content rather than superficial stylistic cues. It probes sensitivity to subtle correctness differences, resistance to verbosity/style bias, and performance across diverse domains like math, code, and safety. Use when the user wants to benchmark on RM-Bench, or asks about evaluating this task. Reports Average Accuracy.
    3 repo stars
  58. ▌
    Robocoin Eval · qhjqhj00
    Evaluates the effectiveness of a large-scale bimanual manipulation dataset and a hierarchical annotation framework on Vision-Language-Action models across different robotic platforms. It probes the models' ability to generalize across task complexities, leverage multi-resolution annotations, and benefit from trajectory quality filtering. Use when the user wants to benchmark on RoboCOIN, or asks about evaluating this task. Reports success_rate.
    3 repo stars
  59. ▌
    Roc Auc Score · qhjqhj00
    Compute the roc_auc_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute roc_auc_score, or asks how to score with roc_auc_score.
    3 repo stars
  60. ▌
    Rood Mri Eval · qhjqhj00
    Evaluates the robustness of deep learning segmentation models to out-of-distribution MRI data and synthetic corruptions (noise, contrast, resolution, spatial shifts, motion artifacts) across multiple severity levels. It measures performance degradation on anatomical and lesion segmentation tasks compared to clean data. Use when the user wants to benchmark on ROOD-MRI Benchmark (Hippocampus, Ventricle, WMH), or asks about evaluating this task. Reports DSC.
    3 repo stars
  61. ▌
    Routenlp Eval · qhjqhj00
    Evaluates a closed-loop LLM routing system's ability to dynamically select between a four-tier model portfolio based on task difficulty, balancing inference cost, response quality, and latency. The benchmark probes how well a router can escalate queries to more capable models only when necessary, while using distillation and conformal cascading to maintain performance at lower cost tiers. Use when the user wants to benchmark on EDGAR (NER), EDGAR (Summarization), BANKING77* (Intent Classification), BANKING77* (Response Generation), CUAD* (Clause Extraction), CUAD* (Risk Assessment), or asks about evaluating this task. Reports Quality Ratio.
    3 repo stars
  62. ▌
    Ruozhiba Eval · qhjqhj00
    Evaluates large language models' ability to solve complex, logic-heavy Chinese natural language puzzles and perform multi-step reasoning on diverse benchmark tasks. It probes the model's capacity for progressive reasoning, self-verification, and adaptability to structured prompt frameworks without manual tuning. Use when the user wants to benchmark on Ruozhiba, BIG-Bench-Hard, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  63. ▌
    Rvcbench Eval · qhjqhj00
    Evaluates the robustness of modern voice cloning models under realistic deployment conditions, including input variations (accents, text shifts, long context), cross-lingual synthesis, post-processing degradation, and adversarial perturbations. It probes the trade-offs between generation quality, content fidelity, speaker similarity, and deepfake detectability across diverse acoustic and linguistic stressors. Use when the user wants to benchmark on LibriTTS, VCTK, LibriSpeech, RVCBench, or asks about evaluating this task. Reports SIM, WER, MCD.
    3 repo stars
  64. ▌
    Saebench Eval · qhjqhj00
    Evaluates sparse autoencoder (SAE) architectures across multiple dimensions including reconstruction fidelity, feature disentanglement, concept detection, and practical interpretability tasks. It systematically compares how different SAE designs, dictionary sizes, and sparsity levels impact these capabilities. Use when the user wants to benchmark on Gemma-2-2B, Pythia-160M, or asks about evaluating this task. Reports Loss Recovered.
    3 repo stars
  65. ▌
    Safe Pro Eval · qhjqhj00
    This benchmark probes the safety judgment and alignment capabilities of professional-level AI agents. It evaluates whether agents can resist executing harmful or risky actions when given complex, domain-specific instructions in fields like finance, law, and healthcare. Use when the user wants to benchmark on SafePro, or asks about evaluating this task. Reports unsafe rate.
    3 repo stars
  66. ▌
    Sard Ocr Eval · qhjqhj00
    Evaluates the robustness and accuracy of OCR models on synthetic, book-style Arabic documents with high typographic diversity across 10 fonts. It measures character-level precision, word-level accuracy, and overall sequence fluency to benchmark vision-language and traditional OCR systems. Use when the user wants to benchmark on SARD, or asks about evaluating this task. Reports CER.
    3 repo stars
  67. ▌
    Sbr Srgi Eval · qhjqhj00
    Evaluates session-based recommendation models by predicting the next item in a user's session sequence. It probes the model's ability to capture sequential item transitions and leverage global item-transition patterns across sessions to improve ranking accuracy. Use when the user wants to benchmark on Diginetica, Tmall, Nowplaying, or asks about evaluating this task. Reports P@20.
    3 repo stars
  68. ▌
    Scannerf Eval · qhjqhj00
    Evaluates the rendering quality and novel-view synthesis capability of Neural Radiance Field (NeRF) methods on real-world inward-facing object scans. It probes how well models generalize to unseen camera poses when trained with varying image densities and localized acquisition patterns. Use when the user wants to benchmark on ScanNeRF, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  69. ▌
    Scibench Eval · qhjqhj00
    This benchmark evaluates large language models' ability to solve college-level scientific problems across mathematics, chemistry, and physics. It probes multi-step quantitative reasoning, unit conversion, physical derivations, and the efficacy of prompting strategies and external computational tools. Use when the user wants to benchmark on SciBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  70. ▌
    Scicm Scieval · qhjqhj00
    Evaluates cross-modality scientific information extraction by jointly predicting named entities, result entities, and relations from both full-text paragraphs and scientific tables. It probes a model's ability to handle long documents, align entities across modalities, and generalize across different scientific domains. Use when the user wants to benchmark on ScICM, or asks about evaluating this task. Reports F1.
    3 repo stars
  71. ▌
    Scp 116k Eval · qhjqhj00
    Evaluates the scientific reasoning and problem-solving capabilities of LLMs on graduate-level higher education science problems. It measures how well models can parse and solve complex scientific questions involving formulas and equations. Use when the user wants to benchmark on SCP-116K, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  72. ▌
    Screenpr Eval · qhjqhj00
    Evaluates a model's ability to read and describe the content and layout of a GUI screenshot at a specific pointed location. It probes layout-aware screen reading, spatial reasoning, and the capacity to generate focused descriptions for mobile agent navigation. Use when the user wants to benchmark on ScreenPR, or asks about evaluating this task. Reports Content Acc.
    3 repo stars
  73. ▌
    Sea Helm Eval · qhjqhj00
    Evaluates Thai language models across eight competencies including instruction following, multi-turn dialogue stability, natural language understanding, generation, reasoning, safety, and code-switching resistance. Use when the user wants to benchmark on SEA-HELM, or asks about evaluating this task. Reports SEA-HELM Average Score.
    3 repo stars
  74. ▌
    Seabench Eval · qhjqhj00
    Evaluates LLMs' ability to handle open-ended, daily interaction scenarios in Southeast Asian languages. It probes contextual adaptation, instruction following, and safety in real-world multilingual usage. Use when the user wants to benchmark on SeaBench, or asks about evaluating this task. Reports LLM-as-a-Judge Score.
    3 repo stars
  75. ▌
    Secbench Eval · qhjqhj00
    Evaluates large language models' cybersecurity knowledge retention and logical reasoning capabilities across multiple subdomains, languages, and difficulty levels using multiple-choice and short-answer questions. Use when the user wants to benchmark on SecBench, or asks about evaluating this task. Reports correctness percentage.
    3 repo stars
  76. ▌
    Seed Tts Eval · qhjqhj00
    Evaluates zero-shot voice conversion systems on linguistic preservation, speaker identity retention, and audio naturalness across English, Chinese, and cross-lingual settings. It also measures computational efficiency and latency for both streaming and offline inference modes. Use when the user wants to benchmark on Seed-TTS-Eval, or asks about evaluating this task. Reports WER (%).
    3 repo stars
  77. ▌
    Seisclip Eval · qhjqhj00
    Evaluates a seismology foundation model's ability to classify seismic event types, localize epicenters and depths, and determine focal mechanisms using multi-modal seismic data. It probes cross-dataset generalization and compares fine-tuned, frozen, and scratch-trained variants against spectrum-based baselines. Use when the user wants to benchmark on PNW dataset, SCSN dataset, or asks about evaluating this task. Reports AUC.
    3 repo stars
  78. ▌
    Senteval Eval · qhjqhj00
    Evaluates the transferability and quality of universal sentence embeddings across a standardized suite of downstream tasks. It probes capabilities in sentiment classification, natural language inference, semantic textual similarity, and cross-modal image-caption retrieval using fixed hyperparameters and consistent preprocessing. Use when the user wants to benchmark on MR, CR, SUBJ, MPQA, TREC, SST-2, SST-5, SNLI, SICK-E, SICK-R, STS14, MRPC, COCO, or asks about evaluating this task. Reports accuracy, pearson.
    3 repo stars
  79. ▌
    Sh Bench Eval · qhjqhj00
    Evaluates audio LLMs' ability to comprehend multi-speaker conversations while selectively focusing on a target speaker and ignoring bystanders for privacy. It measures both general audio understanding and selective hearing capability under different instruction modes. Use when the user wants to benchmark on SH-Bench, or asks about evaluating this task. Reports Selective Efficacy (SE).
    3 repo stars
  80. ▌
    Silicone Eval · qhjqhj00
    Evaluates a model's ability to perform sequence labelling on spoken dialogues, specifically predicting dialog acts (DA) and emotion/sentiment (E/S) labels per utterance within multi-utterance conversations. Use when the user wants to benchmark on SILICONE, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  81. ▌
    Simmc2 0 Eval · qhjqhj00
    Evaluates multimodal task-oriented dialogue capabilities, specifically focusing on dialogue state tracking, disambiguation, coreference resolution, and response generation using visual scene representations. Use when the user wants to benchmark on SIMMC 2.0, or asks about evaluating this task. Reports Intent-F1.
    3 repo stars
  82. ▌
    Skillret Eval · qhjqhj00
    Evaluates the ability of embedding models and rerankers to accurately retrieve relevant software skills from a large, noisy library based on long, scenario-rich user queries. It probes ranking quality, recall, and completeness in a two-stage retrieve-then-rerank pipeline, highlighting the need for domain-specific fine-tuning over general semantic matching. Use when the user wants to benchmark on SkillRet, or asks about evaluating this task. Reports NDCG@k.
    3 repo stars
  83. ▌
    Slt Pose Eval · qhjqhj00
    This benchmark evaluates how different pose estimation models impact the quality of sign language translation. It probes the robustness of pose estimators to occlusion, temporal instability, and missing hand keypoints, measuring their downstream effect on translation metrics. Use when the user wants to benchmark on RWTH-PHOENIX-Weather 2014, Signsuisse, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  84. ▌
    Soar Rna Eval · qhjqhj00
    Evaluates large language models on zero-shot and chain-of-thought cell type annotation tasks using single-cell RNA-seq gene expression profiles. It probes the models' ability to translate structured genomic data into textual descriptions and accurately predict cell type labels without fine-tuning. Use when the user wants to benchmark on SOAR-RNA, or asks about evaluating this task. Reports Average BLEU.
    3 repo stars
  85. ▌
    Soberdse Eval · qhjqhj00
    Evaluates a learning-based algorithm selection framework for High-Level Synthesis Design Space Exploration (DSE). It measures how accurately the model recommends the best-performing DSE algorithm for a given benchmark, and assesses the resulting optimization performance (ADRS) and runtime compared to heuristic and reinforcement learning baselines. Use when the user wants to benchmark on MachSuite & Polyhedral Benchmarks, or asks about evaluating this task. Reports recommendation_accuracy.
    3 repo stars
  86. ▌
    Somd2025 Eval · qhjqhj00
    Probes the capability of joint entity and relation extraction for identifying software mentions and their attributes (URLs, versions, licenses) in scholarly articles. It specifically tests in-distribution performance and out-of-distribution generalization across two competition phases. Use when the user wants to benchmark on SOMD 2025, or asks about evaluating this task. Reports F1-score.
    3 repo stars
  87. ▌
    Songbsab Eval · qhjqhj00
    Evaluates the effectiveness of an adversarial perturbation method (SongBsAb) designed to prevent illegal singing voice conversion. It probes the method's ability to disrupt singer identity and lyrical fidelity in converted audio while maintaining high audio quality and imperceptibility. Use when the user wants to benchmark on OpenSinger, NUS-48E, or asks about evaluating this task. Reports Lyric Word Error Rate (WER).
    3 repo stars
  88. ▌
    Sparc Cg Eval · qhjqhj00
    This benchmark probes a model's ability to perform compositional generalization in context-dependent Text-to-SQL. It evaluates whether models can correctly combine previously seen SQL query structures with novel modification patterns (e.g., new WHERE or ORDER BY clauses) in multi-turn dialogues. Use when the user wants to benchmark on SPARC-CG, or asks about evaluating this task. Reports question match (QM).
    3 repo stars
  89. ▌
    Sportmot Eval · qhjqhj00
    Evaluates multi-object tracking performance in sports scenes, specifically probing a model's ability to maintain track identities under fast, variable-speed motion and highly similar player appearances. Use when the user wants to benchmark on SportsMOT, or asks about evaluating this task. Reports HOTA.
    3 repo stars
  90. ▌
    Squad2 0 Eval · qhjqhj00
    Probes a model's ability to perform extractive reading comprehension while correctly identifying when a question cannot be answered from the provided context. It forces models to distinguish between answerable and unanswerable questions, testing knowledge gap detection and resistance to semantically relevant distractors. Use when the user wants to benchmark on SQuAD 2.0, or asks about evaluating this task. Reports F1.
    3 repo stars
  91. ▌
    Squality Eval · qhjqhj00
    Evaluates long-document, question-focused summarization quality through structured human ratings and automatic metric correlation. It probes a model's ability to generate accurate, comprehensive, and high-quality summaries that align with human preferences rather than relying on surface-level n-gram overlap. Use when the user wants to benchmark on SQuALITY, or asks about evaluating this task. Reports Human Rating (1-100).
    3 repo stars
  92. ▌
    Ssrbench Eval · qhjqhj00
    Evaluates vision-language models on spatial understanding and general question answering using image-text pairs. It probes capabilities such as object existence, attribute recognition, action identification, counting, positional reasoning, and object identification, specifically testing how well models leverage depth information and spatial reasoning. Use when the user wants to benchmark on SSRBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  93. ▌
    Stage Es Eval · qhjqhj00
    Assesses an LLM's capacity to abstract scene-level events into concise, free-form descriptions without schema constraints. It evaluates whether the generated events form a coherent, non-redundant structure and remain factually grounded in the screenplay text. Use when the user wants to benchmark on STAGE-ES, or asks about evaluating this task. Reports Event-Structure Consistency.
    3 repo stars
  94. ▌
    Stage Kg Eval · qhjqhj00
    Evaluates an LLM's ability to extract and structure narrative knowledge from full-length movie screenplays into canonical graphs. It probes entity and relation recognition while filtering out event-centric noise to focus on stable world-building elements. Use when the user wants to benchmark on STAGE-KG, or asks about evaluating this task. Reports Entity F1.
    3 repo stars
  95. ▌
    Stage QA Eval · qhjqhj00
    Tests retrieval-augmented question answering over long screenplay contexts. It probes the model's ability to locate relevant information across chunked text and synthesize accurate answers under different retrieval architectures. Use when the user wants to benchmark on STAGE-QA, or asks about evaluating this task. Reports Question Correctness.
    3 repo stars
  96. ▌
    Star SQL Eval · qhjqhj00
    Evaluates an LLM's ability to generate correct SQL queries from natural language questions over complex, multi-table database schemas. It probes the model's reasoning capabilities and schema generalization by requiring step-by-step rationales and testing on unseen databases. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports execution accuracy (EX).
    3 repo stars
  97. ▌
    Starflow Eval · qhjqhj00
    Evaluates vision-language models' ability to parse free-form workflow sketch images and generate structured JSON workflow definitions. It probes structural fidelity, trigger and component recognition, and hierarchical consistency in diagram-to-code translation. Use when the user wants to benchmark on StarFlow Dataset, or asks about evaluating this task. Reports FlowSim.
    3 repo stars
  98. ▌
    Step Gui Eval · qhjqhj00
    Evaluates a vision-language model's ability to perceive, locate, and interact with graphical user interfaces across desktop and mobile environments. It also measures general multimodal reasoning and OCR capabilities to ensure the model retains broad foundational skills after GUI-specific training. Use when the user wants to benchmark on ScreenSpot-Pro, ScreenSpot-v2, OSWorld-G, MMBench-GUI-L2, VisualWebBench, OSWorld-Verified, AndroidWorld, AndroidDaily, or asks about evaluating this task. Reports accuracy, pass@3.
    3 repo stars
  99. ▌
    Summeval Eval · qhjqhj00
    This benchmark evaluates how well automatic summarization metrics align with human judgments across multiple quality dimensions. It probes whether standard n-gram, embedding-based, and reference-less metrics reliably predict human-perceived coherence, consistency, fluency, and relevance of generated summaries. Use when the user wants to benchmark on SummEval, or asks about evaluating this task. Reports Kendall’s tau.
    3 repo stars
  100. ▌
    Svdquant Eval · qhjqhj00
    Evaluates the visual fidelity and text-image alignment of quantized diffusion models by generating images from text prompts and comparing them against reference outputs. It probes whether low-bit quantization preserves distributional similarity, perceptual quality, and human-preferred aesthetics compared to full-precision baselines. Use when the user wants to benchmark on MJHQ-30K, sDCI, or asks about evaluating this task. Reports FID.
    3 repo stars