qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Nim4 Asr Eval · qhjqhj00Evaluates automatic speech recognition performance across diverse acoustic and linguistic domains, including English, Mandarin, dialects, code-switching, and in-car conversational scenarios. It measures transcription accuracy and hallucination rates to assess model robustness, latency, and customization capabilities. Use when the user wants to benchmark on LibriSpeech, VoxPopuli, MLS-English, AISHELL-1, AISHELL-2, AISHELL-2021-Eval, WeNetSpeech, SpeechIO, WeNetSpeech-Chuan, WeNetSpeech-Yue, KeSpeech, CS-Dialogue, ASCEND, M4Singer, Internal POI Benchmarks, Internal Media Benchmarks, Internal Device Control, Internal Conversational, or asks about evaluating this task. Reports WER, CER.
- ▌ Noisyner Eval · qhjqhj00Evaluates the robustness of noise models and base models under realistic noisy label conditions, measuring how estimation accuracy and base model performance vary with different noise distributions and amounts of clean data. Use when the user wants to benchmark on NoisyNER, or asks about evaluating this task. Reports micro-average F1 score.
- ▌ Nuinsseg Eval · qhjqhj00Evaluates the ability of models to perform instance-level segmentation of cell nuclei in H&E-stained histological images. It specifically probes robustness to tissue variability, high cell density, and ambiguous boundaries where manual annotation is difficult. Use when the user wants to benchmark on NuInsSeg, or asks about evaluating this task. Reports Dice score.
- ▌ Nuplan R Eval · qhjqhj00Evaluates autonomous driving planners in a closed-loop setting with realistic, reactive multi-agent traffic. It probes a planner's ability to handle complex, interactive driving scenarios by measuring overall success, robustness against catastrophic failures, and consistency across safety and comfort dimensions. Use when the user wants to benchmark on nuPlan-R, or asks about evaluating this task. Reports CLS.
- ▌ Odysseys Eval · qhjqhj00Probes an agent's ability to perform realistic, long-horizon web navigation tasks that require sustained cross-site reasoning, context maintenance across multiple tabs, and efficient action execution. It evaluates whether models can complete complex, multi-step user journeys derived from real browsing behavior within strict step budgets. Use when the user wants to benchmark on Odysseys, or asks about evaluating this task. Reports Perfect Rubrics (%).
- ▌ Oem Gfss Eval · qhjqhj00Evaluates a model's ability to perform generalized few-shot semantic segmentation on remote sensing imagery. It tests whether a model can accurately segment both previously seen (base) and new (novel) land cover classes simultaneously using only a few support examples (5-shot), probing generalization and resistance to class forgetting in low-data regimes. Use when the user wants to benchmark on OEM-GFSS, or asks about evaluating this task. Reports mIoU.
- ▌ Olymmath Eval · qhjqhj00Evaluates advanced mathematical reasoning capabilities on Olympiad-level problems. It probes a model's ability to perform rigorous, step-by-step logical deduction and numerical verification across algebra, geometry, number theory, and combinatorics. The benchmark also assesses cross-lingual reasoning performance between English and Chinese. Use when the user wants to benchmark on OlymMATH, or asks about evaluating this task. Reports Pass@1.
- ▌ Omibench Eval · qhjqhj00Evaluates large vision-language models on Olympiad-level multi-image reasoning tasks across biology, chemistry, mathematics, and physics. It probes the model's ability to integrate complementary visual and textual evidence across multiple images to generate stepwise rationales and select or produce correct final answers. Use when the user wants to benchmark on OMIBench, or asks about evaluating this task. Reports accuracy.
- ▌ Omni Dpo Eval · qhjqhj00Evaluates the instruction-following and mathematical reasoning capabilities of LLMs fine-tuned with a dual-perspective preference optimization method. It measures conversational quality, adherence to instructions, and problem-solving accuracy across diverse open-ended and quantitative benchmarks. Use when the user wants to benchmark on AlpacaEval 2.0, Arena-Hard v0.1, IFEval, SedarEval, GSM8K, MATH 500, AIME 2024, AMC 2023, or asks about evaluating this task. Reports LC(%).
- ▌ Omnifall Eval · qhjqhj00Evaluates human fall detection and action recognition capabilities across controlled (staged) and uncontrolled (wild) video domains. It probes a model's ability to classify a 10-class activity taxonomy, detect binary fall/fallen states, and segment action timelines, while measuring generalization gaps between in-distribution and out-of-distribution settings. Use when the user wants to benchmark on CMDFall, UP-Fall, Le2i, GMDCSA24, EDF, OCCU, CaucaFall, MCFD, OOPS-Fall, or asks about evaluating this task. Reports Balanced Accuracy, Macro F1.
- ▌ Omniflow Eval · qhjqhj00Evaluates optical flow estimation models on synthetic omnidirectional human motion data. It probes the model's ability to handle fisheye distortions, domain-randomized environments, and varying amounts of fine-tuning data. Use when the user wants to benchmark on OmniFlow, or asks about evaluating this task. Reports optical flow error.
- ▌ Omnigen2 Eval · qhjqhj00Evaluates a unified multimodal model's capabilities across visual understanding, text-to-image generation, instruction-based image editing, and in-context generation. It probes compositional prompt following, long-prompt adherence, edit accuracy versus preservation, and subject consistency across single, multiple, and scene contexts. Use when the user wants to benchmark on MMBench, MMMU, MM-Vet, GenEval, DPG-Bench, Emu-Edit, GEdit-Bench-EN, ImgEdit-Bench, OmniContext, or asks about evaluating this task. Reports GenEval Overall, DPG-Bench Overall.
- ▌ Open Sdi Eval · qhjqhj00Evaluates a model's ability to localize manipulated regions in images generated by diffusion models, measuring both pixel-level segmentation precision and image-level binary detection capability across multiple unseen generators. Use when the user wants to benchmark on OpenSDI, or asks about evaluating this task. Reports pixel-level F1.
- ▌ Openfake Eval · qhjqhj00Binary classification capability for detecting AI-generated images versus real photographs. It probes a model's ability to generalize across diverse generative models (diffusion, transformer-based) and real-world social media distributions. Use when the user wants to benchmark on OpenFake, or asks about evaluating this task. Reports F1 Score.
- ▌ Openscan Eval · qhjqhj00Evaluates 3D vision models' ability to perform open-vocabulary scene understanding by querying for fine-grained object attributes (e.g., material, affordance, synonym) rather than standard object classes. It measures both 3D instance segmentation and 3D semantic segmentation capabilities on attribute-based queries across eight linguistic aspects. Use when the user wants to benchmark on OpenScan, or asks about evaluating this task. Reports AP, mIoU.
- ▌ Optbench Eval · qhjqhj00Evaluates formal theorem proving capabilities specifically within the undergraduate optimization domain. It probes a model's ability to generate syntactically correct and semantically progressive Lean 4 proof steps or full scripts under strict verifier constraints, while measuring robustness against catastrophic forgetting on general math benchmarks. Use when the user wants to benchmark on OptBench, MiniF2F-test, ProofNet-test, or asks about evaluating this task. Reports Pass@32.
- ▌ Orgforge Eval · qhjqhj00This benchmark evaluates RAG and retrieval-augmented agents on synthetic corporate corpora by testing their ability to retrieve artifacts, reason over causal and temporal chains, and detect knowledge gaps. It probes multi-hop reasoning, temporal knowledge-state tracking, and absence-of-evidence detection across structured enterprise artifacts like Slack, JIRA, and emails. Use when the user wants to benchmark on OrgForge Synthetic Corporate Corpus, or asks about evaluating this task. Reports OrgForgeScorer.
- ▌ Orsi Sod Eval · qhjqhj00This benchmark evaluates optical remote sensing salient object detection models by measuring their ability to accurately segment prominent objects from complex, cluttered backgrounds. It probes structural consistency, boundary precision, and error magnitude across varying object scales and scene complexities. Use when the user wants to benchmark on ORSSD, EORSSD, ORSI-4199, or asks about evaluating this task. Reports maximum F-measure ($F_{\beta}^{max}$).
- ▌ Oscbench Eval · qhjqhj00This benchmark probes a text-to-video model's ability to accurately render and maintain object state transformations (e.g., peeling, slicing) over time. It additionally measures semantic adherence to prompts, scene consistency, and overall perceptual quality to diagnose temporal coherence and physical realism in generated videos. Use when the user wants to benchmark on OSCBench, or asks about evaluating this task. Reports state-change accuracy.
- ▌ Panmatch Eval · qhjqhj00Evaluates a unified vision model's ability to perform correspondence matching across stereo disparity estimation, optical flow, and feature matching in a zero-shot setting. It probes cross-domain generalization and robustness to challenging conditions like occlusion, lighting changes, and non-Lambertian surfaces. Use when the user wants to benchmark on Middlebury, ETH3D, KITTI, Infinigen, Spring, Sintel, Booster, or asks about evaluating this task. Reports PCA x.
- ▌ Papillon Eval · qhjqhj00Evaluates a privacy-preserving LLM delegation pipeline that sanitizes user queries before sending them to a remote API model. It measures the trade-off between maintaining response quality and minimizing personally identifiable information (PII) leakage in the sanitized prompts. Use when the user wants to benchmark on PUPA-TNB, or asks about evaluating this task. Reports QUAL.
- ▌ Pariksha Eval · qhjqhj00Evaluates multilingual and multi-cultural LLM performance across 10 Indic languages using culturally nuanced prompts. It measures model quality via pairwise comparisons (Elo ratings) and direct assessment scores, while also analyzing human-LLM evaluator agreement and various biases (position, verbosity, self-bias). Use when the user wants to benchmark on PARIKSHA, or asks about evaluating this task. Reports Elo rating, Direct Assessment score.
- ▌ Parsinlu Eval · qhjqhj00Evaluates Persian language understanding across six distinct NLU tasks, including reading comprehension, textual entailment, sentiment analysis, and machine translation. It measures how well pre-trained monolingual and multilingual models perform on native-speaker annotated Persian data compared to human baselines. Use when the user wants to benchmark on ParsiNLU, or asks about evaluating this task. Reports F1, Accuracy.
- ▌ Patenteb Eval · qhjqhj00Evaluates patent text embedding models across 15 diverse tasks including symmetric/asymmetric retrieval, classification, paraphrase detection, and clustering. It specifically probes domain-specific challenges like cross-domain retrieval, fragment-to-document matching, and temporal citation dynamics. Use when the user wants to benchmark on PatenTEB, or asks about evaluating this task. Reports NDCG@10, Macro-F1, Pearson r, V-measure.
- ▌ Perla 3d Eval · qhjqhj00Evaluates a 3D vision-language model's ability to answer questions about indoor scenes and generate dense captions for 3D instances. It probes fine-grained spatial reasoning, object attribute recognition, and scene understanding by comparing generated text against ground-truth annotations using standard natural language generation metrics. Use when the user wants to benchmark on ScanNet, or asks about evaluating this task. Reports CiDEr.
- ▌ Phase No Eval · qhjqhj00Evaluates the ability of neural phase pickers to detect P- and S-wave arrivals in continuous seismic waveforms across multi-station networks. It probes detection accuracy, timing precision, and generalization to out-of-distribution earthquake sequences under varying signal-to-noise conditions. Use when the user wants to benchmark on NCEDC 2020 Test Set, 2019 Ridgecrest Sequence, or asks about evaluating this task. Reports F1 score.
- ▌ Pheno Ca Eval · qhjqhj00Evaluates deep neural networks on phenotypic drug discovery tasks using high-content screening images. It probes the model's ability to deconvolve mechanisms of action, molecular targets, and compound identities from cellular phenotypes, as well as zero-shot compound retrieval for CRISPR perturbations. Use when the user wants to benchmark on Pheno-CA, or asks about evaluating this task. Reports accuracy.
- ▌ Phibench Eval · qhjqhj00Evaluates diverse reasoning and coding capabilities using an internal benchmark designed to minimize data contamination and LLM-judge bias. It probes a model's ability to debug, extend, and explain code, as well as identify errors in mathematical proofs and generate related problems. Use when the user wants to benchmark on PhiBench, or asks about evaluating this task. Reports accuracy.
- ▌ Physgame Eval · qhjqhj00This benchmark evaluates a model's ability to reason about physical laws and detect physical commonsense violations in gameplay videos. It probes spatial, temporal, and meta-information-based physical reasoning through curated multi-choice questions. Use when the user wants to benchmark on PhysGame, or asks about evaluating this task. Reports accuracy.
- ▌ Piibench Eval · qhjqhj00This benchmark probes the ability of NER and PII detection systems to accurately identify and classify personally identifiable information spans across highly heterogeneous, cross-domain text sources. It specifically evaluates cross-domain generalization and robustness to diverse, fine-grained PII entity types that are rarely seen together in standard training corpora. Use when the user wants to benchmark on PIIBench, or asks about evaluating this task. Reports span-level F1.
- ▌ Pinball Score · qhjqhj00Evaluates the sharpness and calibration of probabilistic net-load forecasts. It measures how closely predicted quantiles align with actual observations and how narrow the prediction intervals are while maintaining statistical reliability. Use when the user has predictions and gold and needs to compute Pinball Score.
- ▌ Pixelrec Eval · qhjqhj00Evaluates the ability of recommender systems to rank items using raw pixel images instead of traditional ID embeddings. It probes cold-start item recommendation, cross-domain transfer learning, and end-to-end vision-based recommendation performance. Use when the user wants to benchmark on PixelRec, or asks about evaluating this task. Reports Recall@N.
- ▌ Plantseg Eval · qhjqhj00Evaluates pixel-level segmentation capabilities for identifying and localizing plant diseases in real-world, uncontrolled agricultural imagery across 115 disease classes. The benchmark tests a model's ability to handle fine-grained lesion boundaries, overlapping disease symptoms, and high visual diversity typical of field-captured crops. Use when the user wants to benchmark on PlantSeg, or asks about evaluating this task. Reports mIoU.
- ▌ Pmlbmini Eval · qhjqhj00This benchmark evaluates the discriminative performance of machine learning models on tabular classification tasks under data-scarce conditions (sample sizes ≤500). It specifically probes whether complex AutoML and deep learning approaches can consistently outperform simple baselines like logistic regression when training data is limited. Use when the user wants to benchmark on PMLBmini, or asks about evaluating this task. Reports AUC.
- ▌ Pnm Flow Eval · qhjqhj00Evaluates a topology-based pore network model's ability to predict flow-permeable surface area and hydraulic conductance in granular materials from micro-CT images. Use when the user wants to benchmark on Sphere Packing & High-Explosive Micro-CT Samples, or asks about evaluating this task. Reports conductance_ratio.
- ▌ Polymath Eval · qhjqhj00Evaluates multi-modal mathematical and cognitive reasoning capabilities on visual puzzles. It probes spatial interpretation, relational understanding, pattern recognition, and long-horizon logical reasoning using diagram-based multiple-choice questions. Use when the user wants to benchmark on POLYMATH, or asks about evaluating this task. Reports accuracy.
- ▌ Pralekha Eval · qhjqhj00Evaluates cross-lingual document alignment (CLDA) techniques by measuring chunk/sentence-level alignment accuracy (intrinsic) and the resulting document-level machine translation quality (extrinsic) across English and 11 Indic languages. Use when the user wants to benchmark on Pralekha, or asks about evaluating this task. Reports F1 Score, DocCOMET.
- ▌ Prefeval Eval · qhjqhj00Measures an agent's capability to retain and adhere to user preferences during long, multi-turn conversations. It tests whether the model can maintain consistency without external reminders or with explicit preference cues. Use when the user wants to benchmark on PrefEval, or asks about evaluating this task. Reports preference retention accuracy.
- ▌ Prime Dp Eval · qhjqhj00Evaluates a pre-trained seismic model's multi-task capability on single-station waveforms, specifically phase picking (Pg, Sg, Pn, Sn), P-wave polarization classification, and seismic event type classification. The protocol tests generalization across temporal splits and transfer learning on local data to mitigate dataset imbalance. Use when the user wants to benchmark on CSNCD, or asks about evaluating this task. Reports recall.
- ▌ Primesrl Eval · qhjqhj00Evaluates the quality of Semantic Role Labeling (SRL) systems by measuring precision and recall for predicate senses and argument annotations. It specifically probes a model's step-dependent error propagation by penalizing argument scores when the associated predicate sense is incorrect, while also handling discontinuous and reference arguments. Use when the user has predictions and gold and needs to compute PriMeSRL-Eval.
- ▌ Proofnet Eval · qhjqhj00Evaluates a model's ability to translate between natural language mathematics and Lean 3 formal statements (autoformalization and informalization). It measures syntactic validity, semantic correctness, and lexical similarity to assess reasoning over undergraduate-level theory. Use when the user wants to benchmark on ProofNet, or asks about evaluating this task. Reports Accuracy.
- ▌ Pubmedqa Eval · qhjqhj00This benchmark evaluates a model's ability to perform biomedical research question answering by reasoning over structured scientific abstracts. It requires models to infer yes/no/maybe answers to questions derived from paper titles using only the non-conclusion sections of the abstract, without access to the final conclusion. Use when the user wants to benchmark on PubMedQA (PQA-L), or asks about evaluating this task. Reports accuracy.
- ▌ QA Fb15k Eval · qhjqhj00Evaluates cognition-based hallucination by testing whether LVLMs can leverage world knowledge stored in the LLM to answer entity and relation questions grounded in images. Use when the user wants to benchmark on QA-FB15K, or asks about evaluating this task. Reports Acc.
- ▌ Qcaleval Eval · qhjqhj00Probes vision-language models' ability to interpret quantum calibration plots across diverse visual formats (1D traces, 2D maps, histograms) and perform structured scientific reasoning. It tests capabilities ranging from visual grounding and outcome classification to parameter extraction and operational calibration diagnosis, both in zero-shot and in-context learning settings. Use when the user wants to benchmark on QCalEval, or asks about evaluating this task. Reports accuracy.
- ▌ Qg Bench Eval · qhjqhj00Evaluates the ability of generative language models to produce paragraph-level questions conditioned on a target answer and a context sentence. It probes domain adaptability and multilingual generalization across diverse extractive QA datasets. Use when the user wants to benchmark on SQuAD v1.1, SQuADShifts, SubjQA, Multilingual QA (JAQuAD, GerQuAD, SberQuAD, KorQuAD, FQuAD, Spanish SQuAD, Italian SQuAD), or asks about evaluating this task. Reports automatic evaluation metrics.
- ▌ Quantemp Eval · qhjqhj00Evaluates a model's ability to fact-check real-world numerical claims containing statistical and temporal expressions. It probes evidence retrieval, claim decomposition, and natural language inference to predict veracity (True, False, or Conflicting). Use when the user wants to benchmark on NumTemp, or asks about evaluating this task. Reports Macro-F1.
- ▌ R U Maad Eval · qhjqhj00Evaluates a model's ability to perform unsupervised anomaly detection on multi-agent traffic trajectories in urban environments. It probes frame-wise recognition of rare and abnormal driving behaviors, including both individual maneuvers and context-dependent interactions between agents and static map features. Use when the user wants to benchmark on R-U-MAAD, or asks about evaluating this task. Reports frame-wise anomaly detection accuracy.
- ▌ Re Verse Eval · qhjqhj00Evaluates vision-language models' ability to comprehend long-form sequential manga narratives, focusing on story synthesis, character grounding, and temporal reasoning across non-linear, multi-panel sequences. Use when the user wants to benchmark on Re:Verse, or asks about evaluating this task. Reports BERTScore.
- ▌ Real Iad Eval · qhjqhj00Evaluates industrial anomaly detection models under standard unsupervised and fully unsupervised (noisy training) settings. It probes image-level, pixel-level, and multi-view sample-level defect detection capabilities. Use when the user wants to benchmark on Real-IAD, or asks about evaluating this task. Reports AUROC.
- ▌ Redbench Eval · qhjqhj00Evaluates LLM robustness against adversarial prompts (Attack Success Rate) and their tendency to over-defend on benign prompts (Rejection Rate). It probes safety alignment, refusal behavior, and cross-domain vulnerability across 22 risk categories and 19 domains. Use when the user wants to benchmark on RedBench, or asks about evaluating this task. Reports Rejection Rate (RR), Attack Success Rate (ASR).
- ▌ Redis QA Eval · qhjqhj00Evaluates large language models' ability to answer questions about rare diseases, including diagnosis, symptoms, causes, and related properties. It probes the models' medical knowledge retrieval and reasoning capabilities in a specialized, low-resource domain. Use when the user wants to benchmark on ReDis-QA, or asks about evaluating this task. Reports accuracy.
- ▌ Refcocom Eval · qhjqhj00Evaluates a model's ability to perform referring expression segmentation at both object and part levels. It probes fine-grained cross-modal alignment and pixel-level semantic understanding by requiring precise mask prediction for diverse textual references. Use when the user wants to benchmark on RefCOCOm, or asks about evaluating this task. Reports mIoU.
- ▌ Relbench Eval · qhjqhj00Evaluates the ability of graph neural networks to learn from relational databases by predicting entity attributes (classification and regression) and discovering relationships between entities (link prediction) using primary-foreign key graph structures. Use when the user wants to benchmark on RELBENCH, or asks about evaluating this task. Reports AUROC.
- ▌ Replaydf Eval · qhjqhj00Evaluates the robustness of audio deepfake detection models against replay attacks where deepfake audio is played back and re-recorded through real-world hardware, introducing acoustic distortions and room impulse responses. Use when the user wants to benchmark on ReplayDF, or asks about evaluating this task. Reports EER (%).
- ▌ Reviewmt Eval · qhjqhj00Evaluates LLMs on simulating the academic peer review process across multi-turn dialogues. It probes the model's ability to generate relevant paper summaries, write comprehensive reviews, and make accurate acceptance or rejection decisions based on long-context interactions. Use when the user wants to benchmark on ReviewMT, or asks about evaluating this task. Reports F1-score.
- ▌ Rexbench Eval · qhjqhj00This benchmark evaluates the ability of LLM-based coding agents to autonomously implement research extensions by modifying existing AI/ML codebases based on domain-expert instructions. It probes complex, multi-step software engineering capabilities, including codebase navigation, hypothesis-driven implementation, and producing executable patches. Use when the user wants to benchmark on REXBench, or asks about evaluating this task. Reports final success rate.
- ▌ Rm Bench Eval · qhjqhj00Evaluates reward models' ability to correctly identify preferred responses based on substantive content rather than superficial stylistic cues. It probes sensitivity to subtle correctness differences, resistance to verbosity/style bias, and performance across diverse domains like math, code, and safety. Use when the user wants to benchmark on RM-Bench, or asks about evaluating this task. Reports Average Accuracy.
- ▌ Robocoin Eval · qhjqhj00Evaluates the effectiveness of a large-scale bimanual manipulation dataset and a hierarchical annotation framework on Vision-Language-Action models across different robotic platforms. It probes the models' ability to generalize across task complexities, leverage multi-resolution annotations, and benefit from trajectory quality filtering. Use when the user wants to benchmark on RoboCOIN, or asks about evaluating this task. Reports success_rate.
- ▌ Roc Auc Score · qhjqhj00Compute the roc_auc_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute roc_auc_score, or asks how to score with roc_auc_score.
- ▌ Rood Mri Eval · qhjqhj00Evaluates the robustness of deep learning segmentation models to out-of-distribution MRI data and synthetic corruptions (noise, contrast, resolution, spatial shifts, motion artifacts) across multiple severity levels. It measures performance degradation on anatomical and lesion segmentation tasks compared to clean data. Use when the user wants to benchmark on ROOD-MRI Benchmark (Hippocampus, Ventricle, WMH), or asks about evaluating this task. Reports DSC.
- ▌ Routenlp Eval · qhjqhj00Evaluates a closed-loop LLM routing system's ability to dynamically select between a four-tier model portfolio based on task difficulty, balancing inference cost, response quality, and latency. The benchmark probes how well a router can escalate queries to more capable models only when necessary, while using distillation and conformal cascading to maintain performance at lower cost tiers. Use when the user wants to benchmark on EDGAR (NER), EDGAR (Summarization), BANKING77* (Intent Classification), BANKING77* (Response Generation), CUAD* (Clause Extraction), CUAD* (Risk Assessment), or asks about evaluating this task. Reports Quality Ratio.
- ▌ Ruozhiba Eval · qhjqhj00Evaluates large language models' ability to solve complex, logic-heavy Chinese natural language puzzles and perform multi-step reasoning on diverse benchmark tasks. It probes the model's capacity for progressive reasoning, self-verification, and adaptability to structured prompt frameworks without manual tuning. Use when the user wants to benchmark on Ruozhiba, BIG-Bench-Hard, or asks about evaluating this task. Reports accuracy.
- ▌ Rvcbench Eval · qhjqhj00Evaluates the robustness of modern voice cloning models under realistic deployment conditions, including input variations (accents, text shifts, long context), cross-lingual synthesis, post-processing degradation, and adversarial perturbations. It probes the trade-offs between generation quality, content fidelity, speaker similarity, and deepfake detectability across diverse acoustic and linguistic stressors. Use when the user wants to benchmark on LibriTTS, VCTK, LibriSpeech, RVCBench, or asks about evaluating this task. Reports SIM, WER, MCD.
- ▌ Saebench Eval · qhjqhj00Evaluates sparse autoencoder (SAE) architectures across multiple dimensions including reconstruction fidelity, feature disentanglement, concept detection, and practical interpretability tasks. It systematically compares how different SAE designs, dictionary sizes, and sparsity levels impact these capabilities. Use when the user wants to benchmark on Gemma-2-2B, Pythia-160M, or asks about evaluating this task. Reports Loss Recovered.
- ▌ Safe Pro Eval · qhjqhj00This benchmark probes the safety judgment and alignment capabilities of professional-level AI agents. It evaluates whether agents can resist executing harmful or risky actions when given complex, domain-specific instructions in fields like finance, law, and healthcare. Use when the user wants to benchmark on SafePro, or asks about evaluating this task. Reports unsafe rate.
- ▌ Sard Ocr Eval · qhjqhj00Evaluates the robustness and accuracy of OCR models on synthetic, book-style Arabic documents with high typographic diversity across 10 fonts. It measures character-level precision, word-level accuracy, and overall sequence fluency to benchmark vision-language and traditional OCR systems. Use when the user wants to benchmark on SARD, or asks about evaluating this task. Reports CER.
- ▌ Sbr Srgi Eval · qhjqhj00Evaluates session-based recommendation models by predicting the next item in a user's session sequence. It probes the model's ability to capture sequential item transitions and leverage global item-transition patterns across sessions to improve ranking accuracy. Use when the user wants to benchmark on Diginetica, Tmall, Nowplaying, or asks about evaluating this task. Reports P@20.
- ▌ Scannerf Eval · qhjqhj00Evaluates the rendering quality and novel-view synthesis capability of Neural Radiance Field (NeRF) methods on real-world inward-facing object scans. It probes how well models generalize to unseen camera poses when trained with varying image densities and localized acquisition patterns. Use when the user wants to benchmark on ScanNeRF, or asks about evaluating this task. Reports PSNR.
- ▌ Scibench Eval · qhjqhj00This benchmark evaluates large language models' ability to solve college-level scientific problems across mathematics, chemistry, and physics. It probes multi-step quantitative reasoning, unit conversion, physical derivations, and the efficacy of prompting strategies and external computational tools. Use when the user wants to benchmark on SciBench, or asks about evaluating this task. Reports accuracy.
- ▌ Scicm Scieval · qhjqhj00Evaluates cross-modality scientific information extraction by jointly predicting named entities, result entities, and relations from both full-text paragraphs and scientific tables. It probes a model's ability to handle long documents, align entities across modalities, and generalize across different scientific domains. Use when the user wants to benchmark on ScICM, or asks about evaluating this task. Reports F1.
- ▌ Scp 116k Eval · qhjqhj00Evaluates the scientific reasoning and problem-solving capabilities of LLMs on graduate-level higher education science problems. It measures how well models can parse and solve complex scientific questions involving formulas and equations. Use when the user wants to benchmark on SCP-116K, or asks about evaluating this task. Reports Accuracy.
- ▌ Screenpr Eval · qhjqhj00Evaluates a model's ability to read and describe the content and layout of a GUI screenshot at a specific pointed location. It probes layout-aware screen reading, spatial reasoning, and the capacity to generate focused descriptions for mobile agent navigation. Use when the user wants to benchmark on ScreenPR, or asks about evaluating this task. Reports Content Acc.
- ▌ Sea Helm Eval · qhjqhj00Evaluates Thai language models across eight competencies including instruction following, multi-turn dialogue stability, natural language understanding, generation, reasoning, safety, and code-switching resistance. Use when the user wants to benchmark on SEA-HELM, or asks about evaluating this task. Reports SEA-HELM Average Score.
- ▌ Seabench Eval · qhjqhj00Evaluates LLMs' ability to handle open-ended, daily interaction scenarios in Southeast Asian languages. It probes contextual adaptation, instruction following, and safety in real-world multilingual usage. Use when the user wants to benchmark on SeaBench, or asks about evaluating this task. Reports LLM-as-a-Judge Score.
- ▌ Secbench Eval · qhjqhj00Evaluates large language models' cybersecurity knowledge retention and logical reasoning capabilities across multiple subdomains, languages, and difficulty levels using multiple-choice and short-answer questions. Use when the user wants to benchmark on SecBench, or asks about evaluating this task. Reports correctness percentage.
- ▌ Seed Tts Eval · qhjqhj00Evaluates zero-shot voice conversion systems on linguistic preservation, speaker identity retention, and audio naturalness across English, Chinese, and cross-lingual settings. It also measures computational efficiency and latency for both streaming and offline inference modes. Use when the user wants to benchmark on Seed-TTS-Eval, or asks about evaluating this task. Reports WER (%).
- ▌ Seisclip Eval · qhjqhj00Evaluates a seismology foundation model's ability to classify seismic event types, localize epicenters and depths, and determine focal mechanisms using multi-modal seismic data. It probes cross-dataset generalization and compares fine-tuned, frozen, and scratch-trained variants against spectrum-based baselines. Use when the user wants to benchmark on PNW dataset, SCSN dataset, or asks about evaluating this task. Reports AUC.
- ▌ Senteval Eval · qhjqhj00Evaluates the transferability and quality of universal sentence embeddings across a standardized suite of downstream tasks. It probes capabilities in sentiment classification, natural language inference, semantic textual similarity, and cross-modal image-caption retrieval using fixed hyperparameters and consistent preprocessing. Use when the user wants to benchmark on MR, CR, SUBJ, MPQA, TREC, SST-2, SST-5, SNLI, SICK-E, SICK-R, STS14, MRPC, COCO, or asks about evaluating this task. Reports accuracy, pearson.
- ▌ Sh Bench Eval · qhjqhj00Evaluates audio LLMs' ability to comprehend multi-speaker conversations while selectively focusing on a target speaker and ignoring bystanders for privacy. It measures both general audio understanding and selective hearing capability under different instruction modes. Use when the user wants to benchmark on SH-Bench, or asks about evaluating this task. Reports Selective Efficacy (SE).
- ▌ Silicone Eval · qhjqhj00Evaluates a model's ability to perform sequence labelling on spoken dialogues, specifically predicting dialog acts (DA) and emotion/sentiment (E/S) labels per utterance within multi-utterance conversations. Use when the user wants to benchmark on SILICONE, or asks about evaluating this task. Reports accuracy.
- ▌ Simmc2 0 Eval · qhjqhj00Evaluates multimodal task-oriented dialogue capabilities, specifically focusing on dialogue state tracking, disambiguation, coreference resolution, and response generation using visual scene representations. Use when the user wants to benchmark on SIMMC 2.0, or asks about evaluating this task. Reports Intent-F1.
- ▌ Skillret Eval · qhjqhj00Evaluates the ability of embedding models and rerankers to accurately retrieve relevant software skills from a large, noisy library based on long, scenario-rich user queries. It probes ranking quality, recall, and completeness in a two-stage retrieve-then-rerank pipeline, highlighting the need for domain-specific fine-tuning over general semantic matching. Use when the user wants to benchmark on SkillRet, or asks about evaluating this task. Reports NDCG@k.
- ▌ Slt Pose Eval · qhjqhj00This benchmark evaluates how different pose estimation models impact the quality of sign language translation. It probes the robustness of pose estimators to occlusion, temporal instability, and missing hand keypoints, measuring their downstream effect on translation metrics. Use when the user wants to benchmark on RWTH-PHOENIX-Weather 2014, Signsuisse, or asks about evaluating this task. Reports BLEU.
- ▌ Soar Rna Eval · qhjqhj00Evaluates large language models on zero-shot and chain-of-thought cell type annotation tasks using single-cell RNA-seq gene expression profiles. It probes the models' ability to translate structured genomic data into textual descriptions and accurately predict cell type labels without fine-tuning. Use when the user wants to benchmark on SOAR-RNA, or asks about evaluating this task. Reports Average BLEU.
- ▌ Soberdse Eval · qhjqhj00Evaluates a learning-based algorithm selection framework for High-Level Synthesis Design Space Exploration (DSE). It measures how accurately the model recommends the best-performing DSE algorithm for a given benchmark, and assesses the resulting optimization performance (ADRS) and runtime compared to heuristic and reinforcement learning baselines. Use when the user wants to benchmark on MachSuite & Polyhedral Benchmarks, or asks about evaluating this task. Reports recommendation_accuracy.
- ▌ Somd2025 Eval · qhjqhj00Probes the capability of joint entity and relation extraction for identifying software mentions and their attributes (URLs, versions, licenses) in scholarly articles. It specifically tests in-distribution performance and out-of-distribution generalization across two competition phases. Use when the user wants to benchmark on SOMD 2025, or asks about evaluating this task. Reports F1-score.
- ▌ Songbsab Eval · qhjqhj00Evaluates the effectiveness of an adversarial perturbation method (SongBsAb) designed to prevent illegal singing voice conversion. It probes the method's ability to disrupt singer identity and lyrical fidelity in converted audio while maintaining high audio quality and imperceptibility. Use when the user wants to benchmark on OpenSinger, NUS-48E, or asks about evaluating this task. Reports Lyric Word Error Rate (WER).
- ▌ Sparc Cg Eval · qhjqhj00This benchmark probes a model's ability to perform compositional generalization in context-dependent Text-to-SQL. It evaluates whether models can correctly combine previously seen SQL query structures with novel modification patterns (e.g., new WHERE or ORDER BY clauses) in multi-turn dialogues. Use when the user wants to benchmark on SPARC-CG, or asks about evaluating this task. Reports question match (QM).
- ▌ Sportmot Eval · qhjqhj00Evaluates multi-object tracking performance in sports scenes, specifically probing a model's ability to maintain track identities under fast, variable-speed motion and highly similar player appearances. Use when the user wants to benchmark on SportsMOT, or asks about evaluating this task. Reports HOTA.
- ▌ Squad2 0 Eval · qhjqhj00Probes a model's ability to perform extractive reading comprehension while correctly identifying when a question cannot be answered from the provided context. It forces models to distinguish between answerable and unanswerable questions, testing knowledge gap detection and resistance to semantically relevant distractors. Use when the user wants to benchmark on SQuAD 2.0, or asks about evaluating this task. Reports F1.
- ▌ Squality Eval · qhjqhj00Evaluates long-document, question-focused summarization quality through structured human ratings and automatic metric correlation. It probes a model's ability to generate accurate, comprehensive, and high-quality summaries that align with human preferences rather than relying on surface-level n-gram overlap. Use when the user wants to benchmark on SQuALITY, or asks about evaluating this task. Reports Human Rating (1-100).
- ▌ Ssrbench Eval · qhjqhj00Evaluates vision-language models on spatial understanding and general question answering using image-text pairs. It probes capabilities such as object existence, attribute recognition, action identification, counting, positional reasoning, and object identification, specifically testing how well models leverage depth information and spatial reasoning. Use when the user wants to benchmark on SSRBench, or asks about evaluating this task. Reports accuracy.
- ▌ Stage Es Eval · qhjqhj00Assesses an LLM's capacity to abstract scene-level events into concise, free-form descriptions without schema constraints. It evaluates whether the generated events form a coherent, non-redundant structure and remain factually grounded in the screenplay text. Use when the user wants to benchmark on STAGE-ES, or asks about evaluating this task. Reports Event-Structure Consistency.
- ▌ Stage Kg Eval · qhjqhj00Evaluates an LLM's ability to extract and structure narrative knowledge from full-length movie screenplays into canonical graphs. It probes entity and relation recognition while filtering out event-centric noise to focus on stable world-building elements. Use when the user wants to benchmark on STAGE-KG, or asks about evaluating this task. Reports Entity F1.
- ▌ Stage QA Eval · qhjqhj00Tests retrieval-augmented question answering over long screenplay contexts. It probes the model's ability to locate relevant information across chunked text and synthesize accurate answers under different retrieval architectures. Use when the user wants to benchmark on STAGE-QA, or asks about evaluating this task. Reports Question Correctness.
- ▌ Star SQL Eval · qhjqhj00Evaluates an LLM's ability to generate correct SQL queries from natural language questions over complex, multi-table database schemas. It probes the model's reasoning capabilities and schema generalization by requiring step-by-step rationales and testing on unseen databases. Use when the user wants to benchmark on Spider, or asks about evaluating this task. Reports execution accuracy (EX).
- ▌ Starflow Eval · qhjqhj00Evaluates vision-language models' ability to parse free-form workflow sketch images and generate structured JSON workflow definitions. It probes structural fidelity, trigger and component recognition, and hierarchical consistency in diagram-to-code translation. Use when the user wants to benchmark on StarFlow Dataset, or asks about evaluating this task. Reports FlowSim.
- ▌ Step Gui Eval · qhjqhj00Evaluates a vision-language model's ability to perceive, locate, and interact with graphical user interfaces across desktop and mobile environments. It also measures general multimodal reasoning and OCR capabilities to ensure the model retains broad foundational skills after GUI-specific training. Use when the user wants to benchmark on ScreenSpot-Pro, ScreenSpot-v2, OSWorld-G, MMBench-GUI-L2, VisualWebBench, OSWorld-Verified, AndroidWorld, AndroidDaily, or asks about evaluating this task. Reports accuracy, pass@3.
- ▌ Summeval Eval · qhjqhj00This benchmark evaluates how well automatic summarization metrics align with human judgments across multiple quality dimensions. It probes whether standard n-gram, embedding-based, and reference-less metrics reliably predict human-perceived coherence, consistency, fluency, and relevance of generated summaries. Use when the user wants to benchmark on SummEval, or asks about evaluating this task. Reports Kendall’s tau.
- ▌ Svdquant Eval · qhjqhj00Evaluates the visual fidelity and text-image alignment of quantized diffusion models by generating images from text prompts and comparing them against reference outputs. It probes whether low-bit quantization preserves distributional similarity, perceptual quality, and human-preferred aesthetics compared to full-precision baselines. Use when the user wants to benchmark on MJHQ-30K, sDCI, or asks about evaluating this task. Reports FID.