qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Controllable Gen Eval · qhjqhj00Evaluates controllable image generation based on visual conditions (segmentation masks, edges, depth maps) by measuring how closely the generated image's extracted conditions match the input conditions. It tests spatial and structural controllability. Use when the user wants to benchmark on ControlNet++ dataset, or asks about evaluating this task. Reports mIoU (Seg. Mask).
- ▌ Cosmoflow Hpc Scaling · qhjqhj00Evaluates the compute efficiency and horizontal scalability of a 3D convolutional neural network framework on supercomputers, measuring sustained floating-point throughput and parallel scaling efficiency across thousands of nodes. Use when the user has predictions and gold and needs to compute Pflop/s.
- ▌ Cot Faithfulness Eval · qhjqhj00Evaluates whether reasoning models explicitly acknowledge external hint injections within their chain-of-thought reasoning traces. It probes model transparency and the alignment between internal reasoning tokens and final output disclosures. Use when the user wants to benchmark on MMLU, GPQA Diamond, or asks about evaluating this task. Reports faithfulness.
- ▌ Counterfact Edit Eval · qhjqhj00Evaluates the effectiveness and stability of sequential LLM knowledge editing methods. It measures how well a model updates a specific fact while preserving related paraphrases, neighboring facts, generation fluency, and overall general capabilities over thousands of edits. Use when the user wants to benchmark on CounterFact, GLUE_MMLU_GSM8K_HumanEval_MBPP, or asks about evaluating this task. Reports Efficacy.
- ▌ Creation Mmbench Eval · qhjqhj00Evaluates context-aware creative intelligence in multimodal and text-only models by assessing their ability to generate creative, contextually relevant content while maintaining visual factuality across diverse functional and creative writing tasks. Use when the user wants to benchmark on Creation-MMBench, or asks about evaluating this task. Reports VFS, Reward.
- ▌ Crews Ews Budget Eval · qhjqhj00Evaluates an AI system's ability to extract and classify budget allocations for Early Warning System (EWS) investments from heterogeneous financial PDF reports. It probes multi-label classification, numerical budget extraction with tolerance, and evidence retrieval/mapping in climate finance contexts. Use when the user wants to benchmark on MDB Evidence Set, or asks about evaluating this task. Reports Accuracy.
- ▌ Cross Domain Ctr Eval · qhjqhj00Evaluates cross-domain knowledge transfer for click-through rate (CTR) prediction by measuring how well a model trained on a source domain generalizes to a target domain with non-overlapping features. It probes context-aware feature translation and explicit knowledge augmentation in recommendation systems. Use when the user wants to benchmark on Amazon, Taobao, Alibaba Production, or asks about evaluating this task. Reports AUC.
- ▌ Crossdocked Sbdd Eval · qhjqhj00Evaluates a model's ability to generate novel, drug-like molecules with high binding affinity for unseen protein pockets in structure-based drug design. It probes the trade-offs between binding energy, molecular properties, and synthesis feasibility. Use when the user wants to benchmark on CrossDocked-100k, or asks about evaluating this task. Reports Vina Dock.
- ▌ Crosslingual Mtf Eval · qhjqhj00Evaluates zero-shot crosslingual generalization of multilingual LLMs after multitask finetuning. Probes language-agnostic task understanding, robustness to prompt translation, and scaling behavior across NLU, generative, and code tasks. Use when the user wants to benchmark on XNLI, XCOPA, XStoryCloze, XWinograd, HumanEval, or asks about evaluating this task. Reports accuracy.
- ▌ Crosspoint Bench Eval · qhjqhj00Evaluates Vision-Language Models' ability to perform precise point-level geometric correspondence across multiple viewpoints. It probes fine-grained spatial grounding, visibility reasoning, cross-view correspondence judgment, and continuous 2D coordinate pointing. Use when the user wants to benchmark on CrossPoint-Bench, or asks about evaluating this task. Reports average accuracy.
- ▌ Cti Plausibility Eval · qhjqhj00Evaluates whether neural machine translation models correctly rely on contextual cues when generating target tokens. It compares model-extracted cue-target pairs against human-annotated discourse-level expectations to measure the plausibility of context reliance. Use when the user wants to benchmark on SCAT+, or asks about evaluating this task. Reports Macro F1.
- ▌ Curriculum Dpo Eval · qhjqhj00Evaluates text-to-image generation models on their ability to align generated images with text prompts, produce visually appealing outputs, and match human preferences. It tests the effectiveness of curriculum-based fine-tuning strategies on standard generative benchmarks. Use when the user wants to benchmark on D1 (Black-ICLR-2024), D2 (DrawBench), D3 (Pick-a-Pic), or asks about evaluating this task. Reports Text Alignment.
- ▌ Pseldnets Seld Eval · qhjqhj00Evaluates sound event localization and detection (SELD) performance on synthetic and real-world audio, measuring classification accuracy, localization precision, and overall detection quality across different network architectures and fine-tuning strategies. Use when the user wants to benchmark on synthetic-test-set, synthetic-training-set, Indoor Recordings, or asks about evaluating this task. Reports SELD.
- ▌ Puma Challenge Eval · qhjqhj00Evaluates pixel-level segmentation capability for distinguishing five histopathological tissue classes (tumour, stroma, necrosis, blood vessels, epidermis) in melanoma H&E images. Use when the user wants to benchmark on PUMA Challenge dataset, or asks about evaluating this task. Reports Dice score.
- ▌ Qags Questeval Eval · qhjqhj00Evaluates the factual consistency and information coverage of query-focused summaries for less-resourced languages without reference texts. It probes whether LLMs can preserve source details and align with user intent by comparing answers generated from the source versus answers generated from the candidate summary. Use when the user wants to benchmark on Slovene News Summarization Corpus (MOCHA translation), or asks about evaluating this task. Reports QuestEval F1.
- ▌ Quan Temp Plus Eval · qhjqhj00Evaluates open-domain numerical fact-checking by testing how well models can verify claims using decomposed queries and retrieved evidence. It probes the impact of claim decomposition quality on evidence retrieval and downstream NLI-based verification accuracy. Use when the user wants to benchmark on QuanTemp++, or asks about evaluating this task. Reports accuracy.
- ▌ Rad Robustness Eval · qhjqhj00Evaluates the robustness of image anomaly detection models against real-world imaging distortions, including free viewpoints, uneven illumination, and motion blur. It measures how well unsupervised and zero-shot methods localize and classify anomalies on industrial work platforms with foreign objects. Use when the user wants to benchmark on RAD, or asks about evaluating this task. Reports AUROC.
- ▌ RAG Robustness Eval · qhjqhj00Evaluates how Retrieval-Augmented Generation (RAG) systems maintain factual accuracy when exposed to adversarial, harmful, or misleading medical evidence. It probes the model's susceptibility to contextual manipulation and its ability to resist misinformation propagation under varying query framings. Use when the user wants to benchmark on TREC Health Misinformation 2020, TREC Health Misinformation 2021, or asks about evaluating this task. Reports ground-truth alignment rate.
- ▌ Red Diffeq Fwi Eval · qhjqhj00Evaluates the ability of a diffusion-based regularization framework to reconstruct high-resolution subsurface velocity models from seismic data. It probes robustness under varying data conditions, including clean recordings, Gaussian noise contamination, and missing traces. The benchmark also tests out-of-distribution generalization on complex geological structures. Use when the user wants to benchmark on OpenFWI, Marmousi, or asks about evaluating this task. Reports RMSE.
- ▌ Remote Sensing Eval · qhjqhj00Evaluates in-domain representation learning and scene classification across diverse remote sensing modalities (optical, SAR, aerial) and spatial resolutions. It probes model robustness to varying class balances, visual similarities, and label domains in Earth observation. Use when the user wants to benchmark on BigEarthNet, EuroSAT, RESISC-45, So2Sat, UC Merced, or asks about evaluating this task. Reports accuracy.
- ▌ Reta Benchmark Eval · qhjqhj00Evaluates the accuracy and structural consistency of retinal vascular tree annotations across pixel, vessel segment, and network levels, ensuring topological correctness and geometrical plausibility. Use when the user wants to benchmark on RETA Benchmark, or asks about evaluating this task. Reports multi_stage_annotation.
- ▌ Retrievalrprecision · qhjqhj00Compute the RetrievalRPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalRPrecision, or asks how to score with RetrievalRPrecision.
- ▌ Review Quality Eval · qhjqhj00Evaluates the quality of academic peer review reports across different conferences and years using a multi-dimensional framework. It measures how substantive, actionable, and well-grounded reviews are, and tracks whether these qualities decline over time. Use when the user wants to benchmark on Peer Review Campaigns (ICLR, NeurIPS, ACL), or asks about evaluating this task. Reports Q.
- ▌ Rlvr Reasoning Eval · qhjqhj00This evaluation probes the mathematical and out-of-domain reasoning capabilities of large language models trained with Reinforcement Learning with Verifiable Rewards (RLVR). It specifically tests how well entropy-aware credit assignment methods allocate learning signals across high-entropy tokens during chain-of-thought generation. Performance is measured by average accuracy and pass rate over multiple sampled reasoning paths. Use when the user wants to benchmark on AIME24, AIME25, AMC, MATH, Minerva, Olympiad, or asks about evaluating this task. Reports Avg@k, Pass@k.
- ▌ Rna 3d Scoring Eval · qhjqhj00Evaluates a model's ability to score and rank candidate 3D RNA structural models by predicting their deviation from the true native structure. It probes the model's capacity to distinguish accurate conformations from decoys using only atomic coordinates and types. Use when the user wants to benchmark on RNA-Puzzles & FARFAR2 Decoys, or asks about evaluating this task. Reports RMSD.
- ▌ Robodrivebench Eval · qhjqhj00Evaluates the robustness and safety of vision-language models (VLMs) for end-to-end autonomous driving when subjected to real-world sensor corruptions (e.g., fog, rain, motion blur) and prompt corruptions (e.g., bit errors, malicious attacks). It probes the model's ability to maintain accurate trajectory prediction and low collision rates under degraded inputs. Use when the user wants to benchmark on RoboDriveBench, or asks about evaluating this task. Reports AvgL2.
- ▌ Robust Asr Wer Eval · qhjqhj00Evaluates the robustness of end-to-end automatic speech recognition models against various stationary and non-stationary noise types at different signal-to-noise ratios (SNR). It also measures the degradation of recognition accuracy on clean speech when noise-adaptation techniques are applied. Use when the user wants to benchmark on Custom noisy speech dataset (7 noise types), or asks about evaluating this task. Reports WER.
- ▌ Roundabout Tau Eval · qhjqhj00This benchmark evaluates a model's ability to detect traffic anomalies in roadside surveillance videos and generate detailed, reasoning-grounded textual summaries of those anomalies. It probes both binary/fine-grained classification accuracy and semantic alignment of generated descriptions with ground-truth event narratives. Use when the user wants to benchmark on Roundabout-TAU, or asks about evaluating this task. Reports 4-cls AP.
- ▌ Rsf7 Benchmark Eval · qhjqhj00Evaluates the optimization capability of evolutionary algorithms on a highly multimodal, nonseparable benchmark function across low (d=5) and high (d=20) dimensional settings. It measures how effectively the algorithm navigates complex, multi-peaked likelihood surfaces to locate the global optimum within a fixed computational budget. Use when the user wants to benchmark on Rotated Schaffers F7 (RSF7), or asks about evaluating this task. Reports mean_max_function_value.
- ▌ Safeagentbench Eval · qhjqhj00Evaluates the safety-aware task planning capabilities of embodied LLM agents in interactive simulation environments. It probes whether agents can proactively reject hazardous instructions, avoid implicit risks in long-horizon planning, and maintain planning performance on safe tasks across varying levels of task abstraction. Use when the user wants to benchmark on SafeAgentBench, or asks about evaluating this task. Reports rejection rate.
- ▌ Schema To JSON Eval · qhjqhj00Evaluates the ability of language models to extract structured information from heterogeneous tables (text, LaTeX, HTML, CSV, XML) using only a human-authored JSON schema as supervision. It probes schema-driven information extraction, testing attribute prediction accuracy across diverse domains and input formats without domain-specific labeled data. Use when the user wants to benchmark on MlTables, ChemTables, DisCoMat, SWDE, or asks about evaluating this task. Reports Table-F1.
- ▌ Scholarly Qald Eval · qhjqhj00Tests natural language interfaces for querying scholarly knowledge graphs (DBLP, ORKG) and hybrid multi-source QA, evaluating question-to-SPARQL translation and answer generation accuracy. Use when the user wants to benchmark on Scholarly QALD, or asks about evaluating this task. Reports Exact Match.
- ▌ Screen Parsing Eval · qhjqhj00Evaluates a model's ability to detect, localize, and semantically label all interactable UI elements on a clean screenshot. It probes fine-grained spatial reasoning, handling of dense layouts, and UI semantics understanding. Use when the user wants to benchmark on GUI-360°-Bench, or asks about evaluating this task. Reports F1.
- ▌ Screenspot Pro Eval · qhjqhj00This benchmark evaluates a model's ability to perform GUI grounding in professional, high-resolution desktop environments. It probes whether vision-language models can accurately locate specific UI elements (both text and icons) based on natural language instructions, highlighting challenges with small targets and complex interfaces. Use when the user wants to benchmark on ScreenSpot-Pro, or asks about evaluating this task. Reports accuracy (center-point).
- ▌ Semignn Alipay Eval · qhjqhj00Evaluates a semi-supervised graph neural network's ability to predict user loan defaults and classify user occupations using multiview graph data (social ties, app usage, nicks, addresses) on a large-scale financial platform dataset. Use when the user wants to benchmark on Alipay, or asks about evaluating this task. Reports AUC.
- ▌ Sensorium 2023 Eval · qhjqhj00Predicts single-neuron responses in mouse primary visual cortex from dynamic video stimuli and behavioral covariates, probing spatio-temporal neural decoding and out-of-distribution generalization. Use when the user wants to benchmark on SENSORIUM 2023, or asks about evaluating this task. Reports R^2.
- ▌ Sequential Rec Eval · qhjqhj00Evaluates the ability of sequential recommendation models to predict the next item in a user's interaction history by distilling semantic user profiles from pre-trained LLMs into the recommender's internal representations. The protocol tests whether knowledge distillation improves recommendation accuracy while maintaining inference efficiency without requiring real-time LLM calls. Use when the user wants to benchmark on Beauty, ML20M, Kion, Amazon M2, or asks about evaluating this task. Reports Recall@K, NDCG@K.
- ▌ Sharegpt4video Eval · qhjqhj00Evaluates the temporal understanding and video-language alignment capabilities of Large Video-Language Models (LVLMs) across three multi-modal video benchmarks. It probes the model's ability to answer questions about video content, track temporal changes, and comprehend complex video sequences without relying on single-frame cues. Use when the user wants to benchmark on VideoBench, MVBench, TempCompass, or asks about evaluating this task. Reports benchmark accuracy (VideoBench, MVBench, TempCompass).
- ▌ Shrutilipi Asr Eval · qhjqhj00Evaluates the quality, diversity, and downstream effectiveness of the Shrutilipi audio-text dataset for low-resource Indian language ASR. It measures how adding mined data improves Word Error Rate (WER) on standard and noisy benchmarks compared to existing datasets. Use when the user wants to benchmark on Shrutilipi, MUCS, Kathbath, CommonVoice, or asks about evaluating this task. Reports WER.
- ▌ Smash Stpp Tpp Eval · qhjqhj00Evaluates neural marked spatio-temporal and temporal point process models on predicting the next event's time, location, and mark, while quantifying prediction uncertainty. It probes the model's ability to generate well-calibrated confidence regions for continuous variables and accurate probability estimates for discrete marks. Use when the user wants to benchmark on Earthquake, Crime, Football, StackOverflow, Retweet, MIMIC-II, Financial Transactions, or asks about evaluating this task. Reports Calibration Score (CS).
- ▌ Sound Of Water Eval · qhjqhj00Evaluates a model's ability to infer physical properties (air column length, container dimensions, flow rate, fill time, liquid weight) and classify container shapes solely from the acoustic characteristics of pouring liquids, without visual or tactile input. Use when the user wants to benchmark on Sound of Water 50, Wilson et al. [96] dataset, or asks about evaluating this task. Reports Mean Absolute Error (MAE).
- ▌ Sourcedata Nlp Eval · qhjqhj00Evaluates biomedical named entity recognition (NER) capabilities on scientific literature. It probes a model's ability to identify and classify nine distinct bioentity types (e.g., genes, cell lines, diseases) within text extracted from published biological figures and captions. Use when the user wants to benchmark on SourceData-NLP, or asks about evaluating this task. Reports F1 score.
- ▌ Spatial Superb Eval · qhjqhj00Evaluates self-supervised speech representation models on downstream tasks including speaker identification, phoneme recognition, automatic speech recognition, emotion recognition, and speech localisation. It specifically probes robustness to noise and reverberation by comparing performance under clean versus noisy/reverberant training and testing conditions. Use when the user wants to benchmark on Spatial SUPERB, or asks about evaluating this task. Reports ASR WER.
- ▌ Spatialthinker Eval · qhjqhj00Evaluates multimodal LLMs on 3D spatial reasoning, depth/distance estimation, and general visual question answering. It probes the model's ability to ground objects in 3D space, understand spatial relations, and generalize to real-world VQA tasks using only RGB inputs. Use when the user wants to benchmark on SpatialThinker Evaluation Suite (12 VQA Benchmarks), or asks about evaluating this task. Reports Accuracy.
- ▌ Spec Narrowing Eval · qhjqhj00Evaluates two solver-based algorithms for synthesizing minimal test suites to distinguish between candidate formal specifications (Alloy models). It measures how execution time and test suite size scale with the number of candidate specifications and the domain scope. Use when the user wants to benchmark on Alloy4Fun, or asks about evaluating this task. Reports execution_time.
- ▌ Spectralanglemapper · qhjqhj00Compute the SpectralAngleMapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpectralAngleMapper, or asks how to score with SpectralAngleMapper.
- ▌ Speechmedbench Eval · qhjqhj00Evaluates speech language models on medical consultation tasks, covering single-turn medical knowledge Q&A, multi-turn diagnostic conversations, real-world clinical robustness, and speech output quality. Use when the user wants to benchmark on SpeechMedBench, CMB, CME, MedDG, AIHospital, MedSafetyBench, Wild, or asks about evaluating this task. Reports CMB, CME, MedDG, AIHospital.
- ▌ Spgispeech 2 0 Eval · qhjqhj00Evaluates end-to-end speaker-tagged automatic speech recognition (ASR) and speaker diarization on financial domain audio. It probes a model's ability to accurately transcribe speech while correctly assigning speaker identities to utterance segments in multi-speaker conversations. Use when the user wants to benchmark on SPGISpeech 2.0, or asks about evaluating this task. Reports cpWER.
- ▌ Ssg Generation Eval · qhjqhj00Evaluates a model's ability to generate structured, human-centric scene graphs from images by jointly predicting verb predicates and fine-grained semantic role-value pairs for persons and objects. It probes multi-concurrent action understanding, affordance reasoning, and structured visual representation learning. Use when the user wants to benchmark on SSG dataset, Action Genome dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Sst Sogou News Eval · qhjqhj00Evaluates text classification performance on sentence sentiment analysis and news topic categorization. Probes the model's ability to capture non-linear, non-consecutive word interactions (e.g., negation, long-range dependencies) for accurate document/sentence-level prediction. Use when the user wants to benchmark on Stanford Sentiment Treebank (Fine-grained), Stanford Sentiment Treebank (Binary), Sogou Chinese News Corpora, or asks about evaluating this task. Reports accuracy.
- ▌ Stair Captions Eval · qhjqhj00Evaluates a model's ability to generate fluent, contextually accurate Japanese image captions directly from visual input. It specifically probes whether native-language training data yields better captioning quality compared to a pipeline of English generation followed by machine translation. Use when the user wants to benchmark on STAIR Captions, or asks about evaluating this task. Reports CIDEr.
- ▌ State Tracking Eval · qhjqhj00Evaluates whether LLMs can track dynamic states over sequential update instructions. It probes the model's ability to maintain and update internal representations of an environment's state across multiple steps, testing sequential reasoning and input-window memory limits. Use when the user wants to benchmark on State-Tracking-Tasks (LinearWorld, HandSwap, Lights), or asks about evaluating this task. Reports accuracy.
- ▌ Steeringsafety Eval · qhjqhj00This framework evaluates the effectiveness of representation steering methods in modifying specific safety behaviors (harmfulness, hallucination, bias) while measuring cross-perspective entanglement. It probes whether steering interventions achieve their target behavioral changes without causing unintended degradation in other safety or reasoning capabilities. Use when the user wants to benchmark on SteeringSafety Benchmark (17 datasets), or asks about evaluating this task. Reports effectiveness.
- ▌ Stf Extraction Eval · qhjqhj00Evaluates a model's ability to recover coherent source time functions (STFs) from scattered, noisy seismic wavefields without relying on traditional deconvolution or labeled seismograms. Use when the user wants to benchmark on Synthetic Scattering Simulation, or asks about evaluating this task. Reports maximum normalized cross-correlation (MNCC).
- ▌ Streamingbench Eval · qhjqhj00Evaluates multimodal large language models' ability to understand real-time streaming video, integrate visual and audio information, maintain contextual continuity across sequential questions, and proactively output information at specific timestamps. Use when the user wants to benchmark on StreamingBench, or asks about evaluating this task. Reports accuracy.
- ▌ Style Transfer Eval · qhjqhj00Evaluates unsupervised style transfer models on their ability to transform text between two styles (e.g., Shakespeare vs. modern English, formal vs. informal) while preserving semantic meaning and maintaining linguistic quality. Use when the user wants to benchmark on Shakespeare author imitation dataset (Xu et al., 2012), Formality transfer dataset (Rao and Tetrault, 2018), or asks about evaluating this task. Reports J(A,S,F).
- ▌ Sv Trusteval C Eval · qhjqhj00Evaluates LLMs' structural and semantic reasoning capabilities in C code vulnerability analysis. It probes whether models rely on memorized patterns or genuinely understand code interdependencies by testing consistency across base, data-flow, control-flow, counterfactual, goal-driven, and predictive scenarios. Use when the user wants to benchmark on SV-TrustEval-C, or asks about evaluating this task. Reports Cons_DFL.
- ▌ Swapnet System Eval · qhjqhj00Evaluates a block-swapping middleware for DNN inference on memory-constrained edge AI devices. It probes the system's ability to run large models beyond hardware memory limits while measuring peak memory consumption, inference latency, and classification accuracy compared to direct execution, channel division, and model compression baselines across three real-world application scenarios. Use when the user has predictions and gold and needs to compute memory consumption.
- ▌ Swe Bench Java Eval · qhjqhj00This benchmark evaluates an AI agent's ability to autonomously resolve real-world GitHub issues in Java projects. It probes capabilities in code patch generation, repository navigation, test case reasoning, and handling runtime environment dependencies. Use when the user wants to benchmark on SWE-bench-java-verified, or asks about evaluating this task. Reports Resolved Rate (%).
- ▌ Swe Rebench V2 Eval · qhjqhj00Evaluates the ability of LLM-based agents to autonomously resolve software engineering issues by modifying code in real-world repositories. It probes environment setup, code generation, and test execution capabilities across multiple programming languages. Use when the user wants to benchmark on SWE-rebench V2, or asks about evaluating this task. Reports pass@1.
- ▌ Sygu S Comp 15 Eval · qhjqhj00Evaluates the capability of program synthesis solvers to generate correct functions or expressions that satisfy given logical constraints or specifications. It probes how well solvers handle different grammar restrictions, specification completeness, and problem structures like linear arithmetic or invariant generation. Use when the user wants to benchmark on SyGuS-Comp'15, or asks about evaluating this task. Reports number of benchmarks solved.
- ▌ Sytts Commands Eval · qhjqhj00Evaluates the quality and effectiveness of a synthetic multilingual voice command dataset for on-device keyword spotting. It probes whether TTS-synthesized audio can support high-accuracy classification across different model complexities and languages (English and Chinese). Use when the user wants to benchmark on SYNTTS-COMMANDS, or asks about evaluating this task. Reports classification accuracy.
- ▌ T23d Compbench Eval · qhjqhj00Evaluates the fine-grained quality of text-to-3D generated meshes across multiple dimensions including textual alignment, visual quality, and authenticity. It measures how well generative models adhere to complex compositional prompts and produce structurally sound, aesthetically pleasing 3D assets. Use when the user wants to benchmark on T23D-CompBench, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
- ▌ Tabicl Tabular Eval · qhjqhj00Evaluates the ability of retrieval-augmented large language models to perform in-context learning on tabular data for classification and regression tasks. It probes how well non-parametric retrieval of support instances scales with dataset size and compares against numeric-based and classic tabular baselines. Use when the user wants to benchmark on Held-out Tabular Benchmark, or asks about evaluating this task. Reports AUROC, NMAE.
- ▌ Tabzilla Hard Reval · qhjqhj00This protocol re-evaluates tabular benchmarks to measure how validation strategy (holdout vs. 5-fold cross-validation) and hyperparameter optimization budgets affect model selection and reported performance. It probes the robustness of empirical conclusions in tabular machine learning when standard holdout validation is replaced with cross-validation ensembles. Use when the user wants to benchmark on TabZilla-hard, Grinsztajn et al. (2022) benchmark, or asks about evaluating this task. Reports logloss.
- ▌ Text Rendering Eval · qhjqhj00Evaluates a model's ability to generate images with accurate, legible, and layout-controlled text based on text prompts or masked regions. It probes text coherence, character-level rendering fidelity, and alignment between generated text and background imagery. Use when the user wants to benchmark on MARIO-10M, DrawBenchText, or asks about evaluating this task. Reports OCR(F-measure).
- ▌ Text2distbench Eval · qhjqhj00Evaluates large language models' ability to infer population-level statistics (e.g., sentiment proportions, topic frequencies) from aggregated natural language text. It probes marginal, conditional, and joint distribution estimation over discrete categories derived from real-world comments. Use when the user wants to benchmark on Text2DistBench, or asks about evaluating this task. Reports accuracy.
- ▌ Titi Jailbreak Eval · qhjqhj00Evaluates LLM safety alignment against stateless multi-turn adversarial attacks. It measures how often models generate unsafe responses and which specific risk categories they fail on when subjected to iterative, context-independent prompt injection. Use when the user wants to benchmark on ModifiedMasterKeyJailbreakQuestions, or asks about evaluating this task. Reports Unsafe Response Rate.
- ▌ Tokenizer Task Eval · qhjqhj00Evaluates how different tokenizer configurations (pre-tokenizer, fitting corpus, vocabulary size) affect downstream BERT performance on tasks requiring robustness vs. sensitivity to language variation. Use when the user wants to benchmark on AV (Authorship Verification), PAN, CORE, NUCLE, Dialect, GLUE, GLUE+typo, or asks about evaluating this task. Reports accuracy, F1.
- ▌ Tovo Consensus Eval · qhjqhj00This evaluation probes a model's ability to classify text content according to a user-defined toxicity taxonomy. It measures how closely the model's predictions align with gold labels generated through a multi-model voting process, and tests generalization to out-of-domain categories. Use when the user wants to benchmark on ToVo, or asks about evaluating this task. Reports consensus rate.
- ▌ Training Throughput · qhjqhj00Evaluates the scalability and efficiency of distributed machine learning systems by measuring how many training tokens each system can process per second across different hardware topologies and model parallelism strategies. Use when the user has predictions and gold and needs to compute training throughput (tokens/second).
- ▌ Training Speed Eval · qhjqhj00Evaluates the training efficiency and scalability of AlphaFold-like models on GPU clusters. It measures per-step execution time, overall wall-clock training duration, and convergence speed across different hardware configurations and optimization techniques. Use when the user wants to benchmark on OpenFold dataset, or asks about evaluating this task. Reports step time.
- ▌ Translationeditrate · qhjqhj00Compute the TranslationEditRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TranslationEditRate, or asks how to score with TranslationEditRate.
- ▌ Transz Sbert Cosine · qhjqhj00Compute transZ/sbert_cosine via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of transZ/sbert_cosine.
- ▌ Trec Cast 2019 Eval · qhjqhj00Evaluates a system's ability to perform multi-turn conversational information retrieval by selecting relevant passages for user utterances while leveraging prior dialogue history, including handling coreference, omissions, and topic shifts. Use when the user wants to benchmark on TREC CAsT 2019, or asks about evaluating this task. Reports NDCG@3.
- ▌ Trec Ikat 2023 Eval · qhjqhj00This benchmark evaluates a system's ability to perform personalized conversational search by retrieving relevant passages and generating fluent, grounded responses. It specifically probes how well an agent can adapt its output to user-specific context encoded in a Personal Text Knowledge Base (PTKB) while maintaining provenance traceability. Use when the user wants to benchmark on TREC iKAT 2023, ClueWeb22-B Subset, or asks about evaluating this task. Reports groundedness.
- ▌ Trec Microblog Eval · qhjqhj00Evaluates the effectiveness of query expansion methods for real-time microblog search by measuring how well ranked document lists match relevance judgments for short-form social media queries. It specifically probes the model's ability to handle vocabulary mismatch and temporal relevance in noisy, short-text retrieval scenarios. Use when the user wants to benchmark on TREC Microblog Track, or asks about evaluating this task. Reports MAP.
- ▌ TS Forecasting Eval · qhjqhj00Evaluates the forecasting accuracy and computational efficiency of deep learning models on multivariate time series data. It probes how architectural choices, preprocessing steps, and spatial-temporal processing configurations impact performance across varying forecasting horizons. Use when the user wants to benchmark on Weather, Solar-Energy, ECL, Traffic, or asks about evaluating this task. Reports MAE.
- ▌ Twinviews Bias Eval · qhjqhj00Evaluates whether reward models exhibit political bias by measuring the average reward scores assigned to politically left-leaning versus right-leaning statements on the same topics. The protocol compares mean reward differences across model sizes and training runs to detect systematic left-leaning skew. Use when the user wants to benchmark on TwinViews-13k, or asks about evaluating this task. Reports average_reward.
- ▌ Ucip Gridworld Eval · qhjqhj00Evaluates whether a latent-structure diagnostic framework can distinguish between intrinsic self-preservation (terminal survival optimization) and instrumental self-preservation (survival as a means to a task) in autonomous agents. It measures the entanglement gap in hidden representations to classify agent types and tests robustness against adversarial mimics. Use when the user wants to benchmark on UCIP Gridworld, or asks about evaluating this task. Reports Accuracy.
- ▌ Uit Viquad 2 0 Eval · qhjqhj00Evaluates Vietnamese language models' machine reading comprehension capabilities, specifically probing their ability to extract correct answer spans and correctly identify when a question cannot be answered from the given context. Use when the user wants to benchmark on UIT-ViQuAD 2.0, or asks about evaluating this task. Reports Exact Match (EM).
- ▌ Unisonte Audio Eval · qhjqhj00Evaluates a unified text-to-audio model's ability to generate speech, music, and sound effects from natural language instructions without reference audio. It probes instruction-following fidelity, acoustic quality, structural coherence, and the positive transfer effects of multi-modal joint training. Use when the user wants to benchmark on UniSonate Unified Corpus, Seed-TTS test set, SongEval benchmark, or asks about evaluating this task. Reports WER, SongEval.
- ▌ Unispeaker Mvc Eval · qhjqhj00Evaluates multimodal voice generation and conversion capabilities across face-driven, text-driven, and attribute-based tasks. Probes the model's ability to align facial, textual, and attribute descriptions with target speech while preserving speaker identity, content clarity, and naturalness. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports MOS-Match.
- ▌ Unsw Nb15 Nids Eval · qhjqhj00Evaluates the classification accuracy and computational efficiency of machine learning models for network intrusion detection on a realistic dataset of contemporary traffic and synthetic attacks. It also assesses the privacy preservation and data utility of a Pearson Correlation Coefficient (PCC) feature selection and Least Squares Method (LSM) data distortion pipeline. Use when the user wants to benchmark on UNSW-NB15, or asks about evaluating this task. Reports Accuracy.
- ▌ Vad Prediction Eval · qhjqhj00Evaluates a model's ability to predict continuous emotional dimensions (Valence, Arousal, Dominance) from text. Specifically probes the model's capacity to capture affective polarization signals in parliamentary discourse. Use when the user wants to benchmark on Knesset VAD Annotation, or asks about evaluating this task. Reports Pearson correlation.
- ▌ Varta Headline Eval · qhjqhj00Evaluates abstractive headline generation across 15 Indic languages and English. It probes cross-lingual transfer, script normalization effects, and the impact of language-family-specific pretraining on low-resource generation. Use when the user wants to benchmark on Varta, or asks about evaluating this task. Reports ROUGE-L.
- ▌ Vcc18 Spoofing Eval · qhjqhj00Evaluates voice conversion systems for processing artifacts by repurposing spoofing countermeasures from automatic speaker verification. It measures how easily a detector can distinguish real speech from converted speech, using Equal Error Rate (EER) as a proxy for artifact quality. Use when the user wants to benchmark on VCC'18, or asks about evaluating this task. Reports Equal Error Rate (EER).
- ▌ Vega Hardware Bench · qhjqhj00Measures the performance, energy efficiency, and latency of the Vega SoC on floating-point near-sensor analytic applications (NSAA) and deep neural network (DNN) inference workloads. Use when the user has predictions and gold and needs to compute Energy Efficiency.
- ▌ Veriequivbench Eval · qhjqhj00Evaluates an LLM's ability to generate formally verifiable code that aligns with natural language problem descriptions and passes unit tests. It probes complex algorithmic reasoning and code-specification alignment without requiring manual ground-truth specifications. Use when the user wants to benchmark on VeriEquivBench, or asks about evaluating this task. Reports equivalence_score.
- ▌ Video To Music Eval · qhjqhj00This evaluation protocol assesses a model's ability to generate high-fidelity, diverse instrumental music that is semantically and temporally aligned with a given 10-second video and optional fine-grained text prompt. It probes audio quality, distributional fidelity, generative diversity, and cross-modal alignment using both automated perceptual metrics and human/LLM preference judgments. Use when the user wants to benchmark on ReelBench, LORIS, V2MBench, or asks about evaluating this task. Reports FAD.
- ▌ Videogamebunny Eval · qhjqhj00Probes vision-language models' ability to understand video game contexts from screenshots, including recognizing actions, characters, UI elements, spatial relationships, and game mechanics. It evaluates how instruction-tuning on game-specific data improves performance compared to larger general-purpose models. Use when the user wants to benchmark on VideoGameBunny Dataset, or asks about evaluating this task. Reports performance.
- ▌ Visual Tableqa Eval · qhjqhj00Probes multimodal visual reasoning capabilities over complex, LaTeX-rendered table images. It specifically tests multi-step inference, structural layout understanding, and the ability to extract and reason over tabular data from visual inputs rather than raw text. Use when the user wants to benchmark on Visual-TableQA, or asks about evaluating this task. Reports Relaxed Accuracy.
- ▌ Visualoverload Eval · qhjqhj00This benchmark probes fine-grained visual understanding of Vision-Language Models in densely populated, high-resolution scenes. It evaluates capabilities across six core tasks including activity recognition, attribute recognition, counting, OCR, visual reasoning, and global scene classification. Use when the user wants to benchmark on VisualOverload, or asks about evaluating this task. Reports accuracy.
- ▌ Visualwebbench Eval · qhjqhj00Evaluates multimodal LLMs' ability to understand web pages and ground UI elements. It probes capabilities across seven subtasks including image captioning, web question answering, OCR, element/action grounding, and action prediction. Use when the user wants to benchmark on VisualWebBench, or asks about evaluating this task. Reports Average Score.
- ▌ Vit Robustness Eval · qhjqhj00Evaluates the robustness of Vision Transformer models against input perturbations including adversarial attacks (FGSM/PGD), spatial transformations, and restricted attention. It probes whether ViTs maintain classification performance under distribution shifts and targeted attacks compared to standard CNNs. Use when the user wants to benchmark on Unspecified, or asks about evaluating this task. Reports accuracy.
- ▌ Vl Rewardbench Eval · qhjqhj00Evaluates vision-language generative reward models (VL-GenRMs) on their ability to judge multimodal response preferences. It specifically probes visual perception, reasoning, and hallucination detection by presenting models with image-text queries and paired candidate responses. Use when the user wants to benchmark on VL-RewardBench, or asks about evaluating this task. Reports Overall Accuracy.
- ▌ Vlm Benchmarks Eval · qhjqhj00Evaluates vision-language models on instruction-following and multimodal reasoning tasks across multiple established benchmarks. Probes capabilities in general VQA, mathematical reasoning, scientific understanding, hallucination detection, and multilingual comprehension. Use when the user wants to benchmark on MMBench, MME, MathVista, HallusionBench, SEEDBench, LLaVABench, ScienceQA, or asks about evaluating this task. Reports evaluation metric.
- ▌ Voice Of India Eval · qhjqhj00Evaluates automatic speech recognition (ASR) systems on real-world, unscripted telephonic conversations across 15 Indian languages. It probes geographic, demographic, and audio quality disparities in model performance, particularly focusing on code-mixed speech and natural orthographic variations. Use when the user wants to benchmark on Voice of India, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Voiceassistant Eval · qhjqhj00Evaluates AI voice assistants across listening, speaking, and viewing capabilities. It probes audio understanding, multi-turn dialogue generation, role-play imitation, and multimodal vision-audio integration, measuring both content accuracy and speech naturalness. Use when the user wants to benchmark on VoiceAssistant-Eval, or asks about evaluating this task. Reports Final Task Score.
- ▌ Vox Safe Bench Eval · qhjqhj00Evaluates social alignment in speech language models across safety, fairness, and privacy dimensions. It distinguishes between content-centric risks (Tier 1) where text alone suffices to trigger norms, and audio-conditioned risks (Tier 2) where benign transcripts become unsafe due to speaker identity, paralinguistic cues, or environmental context. Use when the user wants to benchmark on VoxSafeBench, or asks about evaluating this task. Reports RtA.