qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Ms Marco V2 Ranking Eval · qhjqhj00Evaluates passage and document ranking systems on a large-scale, document-native corpus. It probes a model's ability to retrieve relevant content from millions of documents using sparse, crowd-sourced relevance judgments, while handling realistic corpus drift and query-independent passage extraction. Use when the user wants to benchmark on MS MARCO v2, or asks about evaluating this task. Reports NDCG@10.
- ▌ Mtzig Cross Entropy Loss · qhjqhj00Compute mtzig/cross_entropy_loss via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mtzig/cross_entropy_loss.
- ▌ Multi Agent Latency Eval · qhjqhj00Evaluates the end-to-end latency, success rate, and cost-efficiency of a multi-agent LLM tutoring system under varying concurrency levels across different cloud inference throughput tiers. It isolates the impact of shared vs. priority vs. provisioned inference pools on response time variance and system reliability. Use when the user wants to benchmark on ITAS Student Query Corpus, or asks about evaluating this task. Reports end-to-end latency.
- ▌ Multilingual Safety Eval · qhjqhj00Evaluates a parameter-efficient multilingual safety guardrail's ability to classify content as safe or unsafe across high-resource and low-resource languages. It probes cross-lingual generalization and robustness against diverse harm categories using cluster-guided transfer. Use when the user wants to benchmark on Aegis-Content-Safety-2.0-Test (Aegis-CS2), HarmBench, Redteam2k, JBB-Behaviors, StrongReject, or asks about evaluating this task. Reports accuracy.
- ▌ Multimodal Rec Benchmark · qhjqhj00Evaluates the performance of multimodal deep learning recommender systems across five Amazon product categories. It probes both standard recommendation accuracy and beyond-accuracy dimensions such as novelty, diversity, popularity bias, and catalog coverage. Use when the user wants to benchmark on Amazon (Office, Toys, Beauty, Sports, Clothing), or asks about evaluating this task. Reports Recall@k.
- ▌ Multimodal Tool Use Eval · qhjqhj00Evaluates the ability of agentic multimodal models to perform visual perception, document understanding, and mathematical reasoning. It specifically probes whether models can strategically decide when to invoke external tools (e.g., image cropping, web search, Python code execution) versus answering directly, balancing task accuracy with tool efficiency. Use when the user wants to benchmark on V-Bench, HRBench-4K/8K, TreeBench, MME-RealWorld, SEEDBench2-Plus, CharXiv, MathVista_mini, MathVerse_mini, WeMath, DynaMath, LogicVista, or asks about evaluating this task. Reports accuracy.
- ▌ Multiple Choice Vqa Eval · qhjqhj00This evaluation probes the true multimodal reasoning capability of vision-language models on multiple-choice question answering tasks. It specifically measures whether models rely on actual question understanding or exploit visual relevance imbalances between correct answers and distractors. Performance is assessed under both standard (vision, question, options) and question-omitted (vision, options) settings to detect easy-option bias. Use when the user wants to benchmark on NExT-QA, MMStar, or asks about evaluating this task. Reports accuracy.
- ▌ Music Audio Tagging Eval · qhjqhj00This evaluation probes a model's ability to perform large-scale music audio tagging by predicting a fixed set of semantic labels (e.g., genre, mood, instrumentation) from raw 30-second audio clips. It measures how well architectures generalize across varying dataset sizes and label granularities. Use when the user wants to benchmark on MagnaTagATune (MTT), Million Song Dataset (MSD), Private Dataset, or asks about evaluating this task. Reports top-50 tag prediction.
- ▌ Nemotron Nano V2 Vl Eval · qhjqhj00Evaluates a 12B vision-language model's capabilities across multimodal understanding, long-context reasoning, document/OCR processing, video comprehension, and pure text reasoning. It probes the model's ability to handle diverse visual inputs, follow instructions, and perform complex STEM and code reasoning under varying decoding and reasoning budget constraints. Use when the user wants to benchmark on MMBench V1.1, MMMU, OCRBench, DocVQA, LongVideoBench, MATH-500, GPQA-Diamond, or asks about evaluating this task. Reports accuracy.
- ▌ Nerf View Synthesis Eval · qhjqhj00Evaluates the ability of a neural radiance field to synthesize photorealistic novel views of 3D scenes from a sparse set of input images. It probes geometric reconstruction fidelity, appearance modeling (including non-Lambertian materials), and multi-view consistency across synthetic and real-world captures. Use when the user wants to benchmark on Diffuse Synthetic 360° (DeepVoxels), Realistic Synthetic 360°, Real Forward-Facing, or asks about evaluating this task. Reports PSNR.
- ▌ News Bias Detection Eval · qhjqhj00Evaluates transformer models' ability to classify news articles as biased or unbiased, comparing standard fine-tuning against domain-adapted training. It further probes model decision-making by analyzing word-level SHAP attribution magnitudes and lexical feature importance across true and false predictions. Use when the user wants to benchmark on BABE, or asks about evaluating this task. Reports Binary F1.
- ▌ Nl Code Pair Mining Eval · qhjqhj00Evaluates a machine learning model's ability to automatically extract high-quality, aligned natural language intent and code snippet pairs from Stack Overflow posts. It probes the system's ranking capability, precision, and recall across different programming languages and feature combinations. Use when the user wants to benchmark on Stack Overflow NL-Code Pairs, or asks about evaluating this task. Reports AUC.
- ▌ Nl Object Retrieval Eval · qhjqhj00Evaluates a model's ability to ground natural language queries to specific regions within images by scoring candidate bounding boxes. It probes spatial reasoning, contextual understanding, and cross-modal alignment between text and visual features. Use when the user wants to benchmark on ReferIt, Kitchen, or asks about evaluating this task. Reports P@1.
- ▌ Nllb Clip Retrieval Eval · qhjqhj00Probes multilingual image-text retrieval capability across low-resource languages. It evaluates how effectively the model aligns visual and textual representations when trained with limited data and frozen encoders. Use when the user wants to benchmark on XTD200, Flickr30k-200, or asks about evaluating this task. Reports R@10.
- ▌ Nllp 2024 Legal Nli Eval · qhjqhj00Probes an LLM's ability to perform Natural Language Inference (NLI) within the legal domain. Specifically, it tests whether the model can correctly classify the logical relationship (entailment, neutral, or contradiction) between a formal legal case summary and an informal social media review. Use when the user wants to benchmark on NLLP 2024 Legal NLI, or asks about evaluating this task. Reports F1.
- ▌ Node Classification Eval · qhjqhj00This evaluation probes a model's ability to perform node classification on graphs across a spectrum of homophily regimes, from strongly heterophilic to strongly homophilic. It specifically tests the framework's adaptive capability to switch between a combinatorial predictor and a neural refinement stage based on validation performance. Use when the user wants to benchmark on Texas, Cornell, Actor, CiteSeer, Cora, Pubmed, or asks about evaluating this task. Reports classification accuracy.
- ▌ Optimam Mammography Eval · qhjqhj00Binary classification of high-resolution mammography images to detect malignant breast tissue. It probes the model's ability to distinguish between malignant and non-malignant cases using both localized patches and full-resolution inputs. Use when the user wants to benchmark on OPTIMAM, or asks about evaluating this task. Reports AUC.
- ▌ Ota Firmware Update Eval · qhjqhj00Evaluates the energy efficiency and update latency of Over-The-Air (OTA) firmware update strategies on flash-based, batteryless IoT devices under simulated energy-harvesting conditions. Use when the user wants to benchmark on OTA Firmware Update Benchmarks, or asks about evaluating this task. Reports Total Update Energy Consumption.
- ▌ Padt Unified Vision Eval · qhjqhj00Evaluates a multimodal large language model's ability to perform visual grounding, segmentation, open-vocabulary detection, and referring image captioning by predicting structured visual outputs directly from interleaved visual reference tokens and text. Use when the user wants to benchmark on RefCOCO/+/g, COCO 2017, RIC, or asks about evaluating this task. Reports IoU@0.5 accuracy.
- ▌ Paired T Test Evaluation · qhjqhj00Tests whether a peer prediction mechanism's scoring function is sensitive to report quality by verifying that replacing high-quality reports with degraded or LLM-generated low-quality reports leads to a statistically significant decrease in expected scores. Use when the user has predictions and gold and needs to compute paired difference t-test (p-value).
- ▌ Parrot Multilingual Eval · qhjqhj00Evaluates the multilingual visual-language understanding capabilities of multimodal large language models (MLLMs) across six languages (English, Chinese, Portuguese, Arabic, Turkish, Russian). It probes how well models align visual features with non-English textual instructions and handle cross-lingual multimodal tasks without relying on naive translation. Use when the user wants to benchmark on MMMB, MMBench, or asks about evaluating this task. Reports Accuracy.
- ▌ Pencil Puzzle Bench Eval · qhjqhj00Evaluates multi-step verifiable reasoning and agentic iteration on constraint-satisfaction puzzles. It probes a model's ability to plan, execute moves, check constraints step-by-step, and course-correct over long contexts. Use when the user wants to benchmark on Pencil Puzzle Bench, or asks about evaluating this task. Reports success rate.
- ▌ Phonebit Inference Bench · qhjqhj00Evaluates the inference speed, energy efficiency, and classification accuracy of a GPU-accelerated binary neural network engine on mobile devices against standard frameworks. Use when the user has predictions and gold and needs to compute runtime.
- ▌ Platinum Benchmarks Eval · qhjqhj00Evaluates LLM reliability on curated, low-noise subsets of standard benchmarks (VQA v2.0, SQuAD 2.0, HotPotQA, DROP, BIG-bench) by removing ambiguous examples and re-labeling to minimize ground-truth errors, revealing true model failures on elementary reasoning tasks. Use when the user wants to benchmark on VQA v2.0, SQuAD 2.0, HotPotQA, DROP, BIG-bench, or asks about evaluating this task. Reports accuracy.
- ▌ Polarmem Multimodal Eval · qhjqhj00Evaluates training-free multimodal agents on retrieval-augmented generation, general reasoning, and hallucination robustness by testing a polarized latent graph memory that injects logical constraints at inference time. Use when the user wants to benchmark on MRAMG-Bench, MRAG-Bench, Visual-RAG, MMMU, MMStar, HallusionBench, or asks about evaluating this task. Reports performance.
- ▌ Principle Alignment Eval · qhjqhj00Probes an LLM's ability to align generated responses with a set of natural language constitutional principles without parameter fine-tuning. It measures both overall conformance quality and the reduction of critical principle violations through an inference-time self-correction pipeline. Use when the user wants to benchmark on SafeRLHF, HH-RLHF, or asks about evaluating this task. Reports 5-Point Likert Score Ranking.
- ▌ Prior Loss Seg Benchmark · qhjqhj00Evaluates the effectiveness of various prior-based loss functions (low-level boundary/distance and high-level shape/size constraints) for medical image segmentation across diverse anatomical structures and imaging modalities. Use when the user wants to benchmark on WMH, ISLES, Atrium, Colon, Spleen, Hippocampus, Prostate, ACDC, or asks about evaluating this task. Reports Dice score.
- ▌ Prism Hallucination Eval · qhjqhj00Probes LLM hallucinations across four dimensions (knowledge missing, knowledge errors, reasoning errors, and instruction-following errors) by isolating error sources through controlled, task-specific queries. It evaluates model reliability and guides optimization by measuring error rates across memory, instruction, and reasoning generation stages. Use when the user wants to benchmark on PRISM, or asks about evaluating this task. Reports H-Score.
- ▌ Procgen Competition Eval · qhjqhj00Evaluates reinforcement learning agents on their ability to learn efficiently and generalize to unseen, procedurally generated environments. It measures how well algorithms adapt to novel level distributions under strict computational and timestep constraints. Use when the user wants to benchmark on Procgen Benchmark, or asks about evaluating this task. Reports mean normalized return.
- ▌ Pyroclast Benchmark Eval · qhjqhj00Evaluates the numerical correctness, stability, and parallel performance (CPU/GPU strong scaling and distributed weak scaling) of a matrix-free finite difference geodynamic solver. Use when the user wants to benchmark on Pyroclast Stokes & Advection Benchmarks, or asks about evaluating this task. Reports parallel speedup.
- ▌ Quadsentinel Safety Eval · qhjqhj00This evaluation probes the ability of multi-agent guardrail systems to enforce machine-checkable safety policies over agent trajectories in real-time. It measures how effectively a system detects and blocks unsafe actions while minimizing false positives on enterprise web-agent and malicious behavior benchmarks. Use when the user wants to benchmark on ST-WebAgentBench, AgentHarm, or asks about evaluating this task. Reports Accuracy.
- ▌ R1 Code Interpreter Eval · qhjqhj00Probes LLMs' ability to autonomously generate and execute code for complex reasoning and planning tasks across logic, spatial, order, optimization, search, and math domains. It evaluates how well models can iteratively explore, optimize, and self-check solutions in multi-turn code execution environments. Use when the user wants to benchmark on SymBench, Big-Bench-Hard, Reasoning-Gym, or asks about evaluating this task. Reports exact match or constraint check.
- ▌ Radgpt Tumor Report Eval · qhjqhj00Evaluates an AI pipeline's ability to generate clinically accurate radiology reports from 3D CT scans by detecting tumors, measuring their size, localizing them within organ sub-segments, and staging cancers. It also assesses the textual similarity and diagnostic utility of generated reports compared to ground-truth clinical notes. Use when the user wants to benchmark on AbdomenAtlas 3.0, or asks about evaluating this task. Reports Tumor Detection Sensitivity & Specificity.
- ▌ Reconstruction Pose Eval · qhjqhj00Evaluates how well different 3D reconstruction methods perform in a downstream object pose estimation task, rather than measuring standalone geometric reconstruction accuracy. It compares pose estimation results using reconstructed 3D models against those using ground-truth CAD models. Use when the user wants to benchmark on YCB-V, or asks about evaluating this task. Reports accuracy of the estimated poses.
- ▌ Refugee Law Outcome Eval · qhjqhj00Evaluates NLP models' ability to perform legal information extraction (NER) and predict refugee claim decision outcomes from Canadian legal documents. It probes the model's capacity to handle domain-specific terminology, extract structured entities from unstructured text, and classify case outcomes based on judicial reasoning. Use when the user wants to benchmark on Canadian Refugee Status Determination (RSD) Cases, or asks about evaluating this task. Reports accuracy.
- ▌ Relative Decrease In Mae · qhjqhj00Evaluates the effectiveness of machine-learned preconditioners for linear solvers in semi-implicit shallow-water models by measuring error reduction in the first solver iteration and overall convergence rates during free-running simulations. Use when the user has predictions and gold and needs to compute Relative decrease in MAE.
- ▌ Resel Scientific Ie Eval · qhjqhj00Evaluates a model's ability to perform N-ary relation extraction from scientific documents by first retrieving relevant text/table components and then selecting the correct entities within those components. Use when the user wants to benchmark on SciREX, PubMed, NLP-TDMS (Full), or asks about evaluating this task. Reports Accuracy (Acc).
- ▌ Roboverse Imitation Eval · qhjqhj00Evaluates robot manipulation policies on a unified set of contact-rich pick-and-place and articulation tasks across multiple simulators. It probes both specialist and generalist vision-language-action models on success rates under standard and progressively challenging generalization levels. Use when the user wants to benchmark on ROBOVERSE Imitation Learning Benchmark, or asks about evaluating this task. Reports success rate.
- ▌ Robustbench Cifar10 Eval · qhjqhj00Evaluates adversarial robustness of image classifiers under multi-norm threat models, testing whether pre-screening diagnostics (FOSC, RDI) reliably predict full attack performance and expose worst-case vulnerabilities masked by single-norm evaluations. Use when the user wants to benchmark on RobustBench CIFAR-10, or asks about evaluating this task. Reports robust_accuracy.
- ▌ Roughness Index Distance · qhjqhj00Evaluates surface irregularities and topological consistency between predicted and ground-truth 3D medical segmentation masks. It quantifies local surface roughness, relative roughness differences, and average surface distance to detect spikes, holes, and smoothing artifacts. Use when the user has predictions and gold and needs to compute Roughness Index (RI).
- ▌ Rowen Hallucination Eval · qhjqhj00This evaluation probes an LLM's ability to generate factually correct responses and mitigate hallucinations under adaptive retrieval-augmented generation. It measures factual accuracy via LLM-based scoring on open-ended questions and exact-match accuracy on multi-step reasoning yes/no questions. Use when the user wants to benchmark on TruthfulQA, StrategyQA, or asks about evaluating this task. Reports FactScore.
- ▌ Rscc Change Caption Eval · qhjqhj00Evaluates vision-language models' ability to generate detailed, semantically accurate captions describing changes between bi-temporal remote sensing image pairs, particularly in disaster scenarios. It probes spatiotemporal reasoning, fine-grained environmental change detection, and long-text generation quality. Use when the user wants to benchmark on RSCC, or asks about evaluating this task. Reports ST5-SCS.
- ▌ Saga 3d Mesh Attack Eval · qhjqhj00Evaluates the effectiveness of a spectral geometric adversarial attack on 3D mesh autoencoders by measuring how well perturbed meshes deceive a downstream classifier and evade detection. Use when the user wants to benchmark on CoMA, SMAL, or asks about evaluating this task. Reports Targeted classification accuracy.
- ▌ Saplma Truthfulness Eval · qhjqhj00Evaluates whether an LLM's internal hidden layer activations can predict the veracity of a given statement. It probes the model's implicit knowledge of truthfulness by training a classifier on neural activations rather than relying on explicit prompting or output probabilities. Use when the user wants to benchmark on True-False Dataset, LLM-Generated Statements, or asks about evaluating this task. Reports accuracy.
- ▌ Scandeval Benchmark Eval · qhjqhj00Evaluates the performance of monolingual and multilingual language models across five Scandinavian languages (Danish, Norwegian, Swedish, Icelandic, Faroese) on question answering, linguistic acceptability, and named entity recognition. It also probes cross-lingual transfer capabilities between these languages by measuring performance variance across language groups. Use when the user wants to benchmark on ScandiQA, ScaLA, MIM-GOLD-NER, WikiANN, or asks about evaluating this task. Reports accuracy.
- ▌ Schema Guided Dstc8 Eval · qhjqhj00Evaluates zero-shot dialogue state tracking across single and multi-domain conversations. It measures the model's ability to predict intents, extract slot values, and maintain accurate dialogue states over long contexts without prior exposure to unseen service domains. Use when the user wants to benchmark on Schema-Guided Dialogue (DSTC8 Track 4), or asks about evaluating this task. Reports Joint Goal Accuracy.
- ▌ Scientific Ideation Eval · qhjqhj00Evaluates the ability of LLMs to generate novel, feasible, and effective scientific research ideas given a research question. Probes open-ended scientific reasoning and ideation quality under compute-matched inference budgets. Use when the user wants to benchmark on ICLR 2024 & NeurIPS 2025, or asks about evaluating this task. Reports Absolute Novelty.
- ▌ Scmamba Integration Eval · qhjqhj00Evaluates computational methods for single-cell multi-omics integration by measuring their ability to preserve biological variation, align different omics layers at the cell and single-cell levels, and enable accurate cell type annotation. Use when the user wants to benchmark on SHARE-seq BMMC, SNARE-seq, 10X Genomics Multiome, Human fetal atlas, CITE-seq BMMC S1, CITE-seq BMMC S4, Human brain multi-omics, Human brain 3k, or asks about evaluating this task. Reports biological variation conservation score, omics alignment score, overall integration score.
- ▌ Secure Inference Latency · qhjqhj00Measures the online and offline computation latency and communication bandwidth for cryptographic primitives and neural network operations under secure two-party computation. It evaluates how efficiently packed homomorphic encryption and garbled circuits handle matrix-vector products, convolutions, and activation functions without revealing inputs or model parameters. Use when the user has predictions and gold and needs to compute t_online.
- ▌ Semeval 2023 Task12 Eval · qhjqhj00Sentiment classification across twelve low-resource African languages and Creoles, evaluating model robustness to code-switching and varying degrees of lexical similarity to pretraining data. Use when the user wants to benchmark on SemEval-2023 Task 12, or asks about evaluating this task. Reports macro-F1.
- ▌ Semigda Medical Seg Eval · qhjqhj00Evaluates semi-supervised medical image segmentation performance under limited labeled data ratios (10% and 30%). It measures segmentation accuracy and boundary precision across multiple medical domains including colonoscopy, dermoscopy, pathology, and ultrasound. Use when the user wants to benchmark on Colonoscopy (CVC-ClinicDB, Kvasir, CVC-300), ISIC-2018, BCSS, BUSI, or asks about evaluating this task. Reports Dice coefficient (Dice).
- ▌ Sensitivityatspecificity · qhjqhj00Compute the SensitivityAtSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SensitivityAtSpecificity, or asks how to score with SensitivityAtSpecificity.
- ▌ Shalakasatheesh Squad V2 · qhjqhj00Compute shalakasatheesh/squad_v2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of shalakasatheesh/squad_v2.
- ▌ Shirt Pose Tracking Eval · qhjqhj00Evaluates the accuracy and robustness of a neural network-integrated Unscented Kalman Filter for monocular pose tracking of tumbling noncooperative spacecraft. It probes the system's ability to maintain steady-state position and orientation accuracy under domain gaps between synthetic training data and real hardware-in-the-loop test images. Use when the user wants to benchmark on SHIRT, or asks about evaluating this task. Reports e_pose.
- ▌ Shpi Recommendation Eval · qhjqhj00Evaluates offline reinforcement learning methods for session-based recommendation systems in optimizing long-term user retention versus short-term clicks. It tests the ability of algorithms to learn from fixed logging policies and generalize to online rollouts across synthetic, simulated, and real-world recommendation environments. Use when the user wants to benchmark on Synthetic recommendation problem, RecoGym, HIV treatment simulator, Private dataset X, or asks about evaluating this task. Reports undiscounted test performance on true environment rewards.
- ▌ Simmotion Retrieval Eval · qhjqhj00This benchmark evaluates a model's ability to retrieve videos based on semantic motion similarity, disentangling dynamic behavior from static appearance, camera viewpoint, and scene context. It tests robustness to appearance variations in controlled synthetic settings and unsynchronized, in-the-wild video pairs. Use when the user wants to benchmark on SimMotion-Synthetic, SimMotion-Real-1K, Jester, or asks about evaluating this task. Reports Retrieval accuracy.
- ▌ Slamming Additional Eval · qhjqhj00Evaluates the speech language model's performance across multiple benchmarks including linguistic acceptability, story completion, audio generation quality, and cross-domain text generation perplexity. Use when the user wants to benchmark on sBLIMP, StoryCloze, People Speech, or asks about evaluating this task. Reports MOSnet.
- ▌ Soi Id Ood Accuracy Eval · qhjqhj00Evaluates pretrained language models' in-distribution (ID) and out-of-distribution (OOD) classification accuracy under single-setting and multi-setting fine-tuning configurations. It probes how training dynamics and subset selection affect robustness and generalization across languages, sources, and tasks. Use when the user wants to benchmark on SST-2, IMDB, Yelp, Sentiment140, RTE, QQP, or asks about evaluating this task. Reports accuracy.
- ▌ Space Ris Satellite Eval · qhjqhj00Evaluates a multi-agent DRL framework (MAPPO) combined with whale optimization for maximizing satellite coverage and data rates in 6G sub-THz networks using reconfigurable intelligent surfaces (RIS). Use when the user wants to benchmark on Simulated LEO Satellite-RIS Network Environment, or asks about evaluating this task. Reports average data rate.
- ▌ Specificityatsensitivity · qhjqhj00Compute the SpecificityAtSensitivity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpecificityAtSensitivity, or asks how to score with SpecificityAtSensitivity.
- ▌ Speechinstructbench Eval · qhjqhj00Evaluates speech instruction-following capabilities across closed-ended, open-ended, and adjustment tasks under varying acoustic conditions (background noise, accents, disfluencies) in English and Chinese. Use when the user wants to benchmark on SpeechInstructBench, or asks about evaluating this task. Reports instruction-level accuracy (I).
- ▌ Ssrs Retrosynthesis Eval · qhjqhj00Evaluates single-step retrosynthesis capability by predicting reactant molecules from a given target product, testing both in-distribution chemical knowledge and out-of-distribution generalization. Use when the user wants to benchmark on USPTO-50K-test, URSA-expert-2026, or asks about evaluating this task. Reports Unique.
- ▌ Structured3d Layout Eval · qhjqhj00Evaluates a model's ability to predict architectural elements (walls, doors, windows) and room layouts within indoor 3D scenes. It tests the model's capacity for structured scene understanding and spatial reasoning by comparing predicted layouts against ground-truth annotations. Use when the user wants to benchmark on Structured3D, or asks about evaluating this task. Reports F1.
- ▌ Synthetic Drug Data Eval · qhjqhj00Evaluates the distributional fidelity of generated pharmacokinetic and drug-target interaction properties against real data, and measures the utility of the synthetic data for downstream regression tasks. Use when the user wants to benchmark on TDCommons/BindingDB PK & DTI Collection, or asks about evaluating this task. Reports Hellinger Distance (HD).
- ▌ T2i Deanonymization Eval · qhjqhj00Evaluates the ability to deanonymize text-to-image models by identifying which model generated a given image, exploiting model-specific visual signatures in embedding space. Use when the user wants to benchmark on T2I Leaderboard Prompts, or asks about evaluating this task. Reports Top-1 accuracy.
- ▌ Table Row Detection Eval · qhjqhj00Evaluates graph-based machine learning models for sequence labeling (BIESO) and table row detection on handwritten historical register books. The protocol tests the models' ability to segment table rows and label cell boundaries using pre-extracted textline and column features rather than raw images. Use when the user wants to benchmark on Dataset1, Dataset2, or asks about evaluating this task. Reports F1 score.
- ▌ Tanks Temples Truck Eval · qhjqhj00Evaluates novel view synthesis quality and 3D geometric reconstruction accuracy of NeRF variants on a single outdoor scene. Use when the user wants to benchmark on Tanks & Temples (Truck scene), or asks about evaluating this task. Reports CD.
- ▌ Tcga Histopathology Eval · qhjqhj00Evaluates the representation quality and generalization of self-supervised histopathology models across diverse patch-level diagnostic tasks and weakly supervised slide-level tasks using linear probing and fine-tuning on TCGA whole slide images. Use when the user wants to benchmark on TCGA Histopathology, or asks about evaluating this task. Reports average AUC.
- ▌ Text Classification Eval · qhjqhj00Evaluates text classification performance across multiple sentiment, subjectivity, question classification, and topic categorization tasks. It probes the model's ability to capture contextual and syntactic features from sequential text using 2D matrix representations and spatial pooling. Use when the user wants to benchmark on MR, SST-1, SST-2, Subj, TREC, 20Newsgroups, or asks about evaluating this task. Reports accuracy.
- ▌ Tfbs Classification Eval · qhjqhj00Evaluates a model's ability to classify short DNA sequences as transcription factor binding sites or not, capturing its capacity to learn regulatory sequence patterns from genomic data. Use when the user wants to benchmark on TFBS classification, or asks about evaluating this task. Reports AUC.
- ▌ Themisio Io Sharing Eval · qhjqhj00Evaluates a policy-driven I/O sharing framework for burst buffers by measuring how effectively it allocates bandwidth, maintains fairness, and reduces interference across concurrent workloads. It probes the system's ability to enforce primitive and composite sharing policies under varying load conditions and compares performance against baseline schedulers. Use when the user wants to benchmark on ThemisIO Benchmark & Application Suite, or asks about evaluating this task. Reports sustained I/O throughput.
- ▌ Threat Intelligence Eval · qhjqhj00This benchmark evaluates an AI system's ability to extract actionable insights from threat intelligence reports and perform security reasoning. It probes multi-document comprehension, attack chain reconstruction, and MITRE ATT&CK framework mapping capabilities. Use when the user wants to benchmark on CyberSOCEval Threat Intelligence Reasoning, or asks about evaluating this task. Reports accuracy.
- ▌ Tongue Segmentation Eval · qhjqhj00Evaluates zero-shot and cross-dataset generalization of a tongue segmentation model adapted from SAM. It probes the model's ability to segment tongue regions in medical images without task-specific fine-tuning on the target datasets. Use when the user wants to benchmark on TongueSet1, BioHit (TongueSet2), Webset (TongueSet3), or asks about evaluating this task. Reports mIoU.
- ▌ Toto TS Forecasting Eval · qhjqhj00Evaluates zero-shot and fine-tuned time series forecasting capabilities on real-world observability telemetry and general-purpose benchmarks. Probes model robustness to high-dimensional, nonstationary, multivariate series with skewed distributions and varying temporal intervals. Use when the user wants to benchmark on Boom, Boomlet, GIFT-Eval, LSF, or asks about evaluating this task. Reports CRPS.
- ▌ Trec Session Tracks Eval · qhjqhj00Evaluates the effectiveness of a learning-to-rank personalization approach for web search sessions by measuring relevance prediction quality across multiple years of session track data. It probes the model's ability to leverage historical query sequences, document rankings, and user click behavior to improve session-level relevance ranking. Use when the user wants to benchmark on TREC 2011-2014 Session Tracks, or asks about evaluating this task. Reports nDCG@k.
- ▌ Trec2024 RAG Nugget Eval · qhjqhj00Evaluates the factual accuracy and content grounding of RAG-generated answers by checking for the presence of key factual claims (nuggets) extracted from source documents. It also measures answer length to assess the trade-off between conciseness and completeness in system outputs. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports V_strict.
- ▌ Umlip High Temp Mof Eval · qhjqhj00Evaluates the accuracy of universal machine-learned interatomic potentials (uMLIPs) in predicting energy, forces, and stress tensors during high-temperature molecular dynamics simulations of metal-organic frameworks (MOFs), including their stability and thermal decomposition behavior. Use when the user wants to benchmark on High-Temperature MOF AIMD Benchmark, or asks about evaluating this task. Reports energy MAE.
- ▌ Unet Biomedical Seg Eval · qhjqhj00Evaluates pixel-level biomedical image segmentation capability using convolutional networks. Probes the model's ability to precisely delineate cellular structures and membranes in electron and light microscopy images with limited training data. Use when the user wants to benchmark on EM segmentation challenge (ISBI 2012), PhC-U373, DIC-HeLa, or asks about evaluating this task. Reports warping error, IOU.
- ▌ Us Grid Forecasting Eval · qhjqhj00Evaluates the ability of deep learning architectures (SSMs, Transformers, RNNs) to forecast hourly electricity load across major US power grids. It probes how well models capture temporal patterns, handle varying prediction horizons, and integrate exogenous weather covariates for accurate grid-scale forecasting. Use when the user wants to benchmark on US ISO Hourly Load Data (EIA-930), or asks about evaluating this task. Reports MSE (%).
- ▌ Velocity Dealiasing Eval · qhjqhj00Evaluates a U-Net model's ability to predict velocity fold numbers and produce dealiased radar velocity fields from folded inputs. It measures both classification accuracy for fold detection and reconstruction fidelity via velocity error metrics. Use when the user wants to benchmark on WSR-88D Level-II/III Radar Data, or asks about evaluating this task. Reports velocity RMSE.
- ▌ Verafi Financial QA Eval · qhjqhj00Probes an agentic RAG system's ability to retrieve relevant SEC filings and generate factually correct, complete financial answers. It specifically tests the impact of neurosymbolic policy validation on suppressing hallucinations and mathematical errors in high-stakes financial domains. Use when the user wants to benchmark on FinanceBench-style Financial QA Dataset, or asks about evaluating this task. Reports Factual Correctness.
- ▌ Video Thinking Test Eval · qhjqhj00Evaluates video large language models on their ability to understand complex visual narratives and answer questions correctly. It specifically probes robustness by testing model performance on naturally adversarial or misleading variations of the same video question. Use when the user wants to benchmark on Video Thinking Test, or asks about evaluating this task. Reports Correctness score (accuracy).
- ▌ Vincicoder Code Gen Eval · qhjqhj00Probes a model's ability to generate executable, visually faithful code (HTML, SVG, LaTeX, SMILES) from input images across diverse domains. It evaluates both syntactic correctness via execution rate and perceptual alignment with target images using coarse-to-fine visual similarity metrics. Use when the user wants to benchmark on ChartMimic, Design2Code, UniSVG, Image2Struct, Cosyn-400k, or asks about evaluating this task. Reports UniSVG Final Score.
- ▌ Vision Language Ood Eval · qhjqhj00Probes the ability of vision-language models to distinguish in-distribution from out-of-distribution samples under semantic, covariate, and real-world distribution shifts. It evaluates both zero-shot and few-shot prompt learning approaches across multiple benchmarks to assess robustness and ranking consistency. Use when the user wants to benchmark on ImageNet-X, ImageNet-FS-X, Wilds-FS-X, or asks about evaluating this task. Reports AUROC.
- ▌ Visual Wetlandbirds Eval · qhjqhj00This benchmark evaluates deep learning models on fine-grained bird species classification and spatio-temporal behavior recognition in ecological video footage. It probes the model's ability to localize birds, identify their species, and classify their actions across video frames in real-world wetland environments. Use when the user wants to benchmark on Visual WetlandBirds Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Vl Compositionality Eval · qhjqhj00Evaluates vision-language models on compositional reasoning capabilities, specifically testing their ability to correctly bind attributes, understand semantic relations, and parse word order in image-text pairs. It also measures systematic generalization to unseen concept combinations and zero-shot classification and retrieval performance. Use when the user wants to benchmark on ARO, CREPE, SVO, VL-Checklist, or asks about evaluating this task. Reports accuracy.
- ▌ Wals Metalinguistic Eval · qhjqhj00Probes large language models' ability to recall and identify structural and grammatical properties of languages across diverse linguistic domains. It measures whether models have internalized typological facts from the World Atlas of Language Structures (WALS) by answering multiple-choice questions about specific language features. Use when the user wants to benchmark on WALS, WALS-100, or asks about evaluating this task. Reports accuracy.
- ▌ Weak Annotation Har Eval · qhjqhj00Evaluates inertial-based activity recognition models trained on weakly-supervised labels generated via vision foundation model clustering, benchmarked against fully-supervised and few-shot baselines. Use when the user wants to benchmark on WEAR, Wetlab, ActionSense, or asks about evaluating this task. Reports Acc.
- ▌ Web Agent Benchmark Eval · qhjqhj00Evaluates long-horizon web navigation and information-seeking capabilities. Probes the agent's ability to formulate search queries, browse multiple web pages, synthesize information from diverse sources, and answer complex multi-step questions. Use when the user wants to benchmark on BrowseComp-en, BrowseComp-zh, GAIA, WebWalkerQA, FRAMES, XBench-DeepSearch, HLE, or asks about evaluating this task. Reports Avg@4 Accuracy.
- ▌ Wind Power Ensemble Eval · qhjqhj00Evaluates the calibration, sharpness, and accuracy of probabilistic wind power forecasts under different ensemble post-processing strategies (raw, weather-only, power-only, and joint weather-power post-processing). It probes whether correcting biases at the weather stage alone is sufficient, or if direct post-processing of the final power ensemble is required to handle non-linear power curve biases. Use when the user wants to benchmark on Benchmark Data, Swedish Data Set, or asks about evaluating this task. Reports CRPS.
- ▌ Wmt17 Nmt Benchmark Eval · qhjqhj00Evaluates the training efficiency, inference speed, and translation quality of neural machine translation systems on standard WMT17 benchmarks. It probes the trade-offs between model architecture, hardware acceleration (FP16/INT8), and batching strategies. Use when the user wants to benchmark on WMT17 English-German, WMT17 Russian-English, or asks about evaluating this task. Reports BLEU.
- ▌ Workload Allocation Eval · qhjqhj00Evaluates the latency performance of AI workload allocation strategies across hierarchical cloud/edge/device computing environments for latency-sensitive medical ICU applications. It measures how effectively dynamic routing minimizes end-to-end response time when processing and transmission delays are factored in. Use when the user wants to benchmark on Edge AIBench ICU Applications (MIMIC-III derived), or asks about evaluating this task. Reports response time.
- ▌ World Music Corpora Eval · qhjqhj00Evaluates audio foundation models' cross-cultural generalization across diverse musical traditions (Western, Greek, Turkish, Indian) using multi-label tagging and few-shot learning. Probes whether pre-trained representations capture cultural musical knowledge without extensive adaptation. Use when the user wants to benchmark on Turkish-makam, Hindustani, Carnatic, MagnaTagATune, FMA-medium, Lyra, or asks about evaluating this task. Reports ROC-AUC.
- ▌ 3d Spatial Reasoning Eval · qhjqhj00Evaluates a model's ability to perform 3D visual grounding and situated question answering by reasoning over object coordinates and spatial relations in 3D scenes. It probes whether the model can accurately locate objects based on natural language instructions and answer spatial questions about scene layouts without linguistic interference. Use when the user wants to benchmark on ScanRefer, Multi3DRef, SQA3D, or asks about evaluating this task. Reports accuracy.
- ▌ 6dof Camera Tracking Eval · qhjqhj00Evaluates the tracking accuracy of a 6-DoF autonomous camera algorithm in a simulated surgical environment. It also measures how different camera control strategies impact human rater accuracy when assessing surgical skill from video. Use when the user wants to benchmark on da Vinci wire chaser simulation, or asks about evaluating this task. Reports assessment_error.
- ▌ Acappella Separation Eval · qhjqhj00This benchmark evaluates audio-visual singing voice separation models by measuring how accurately they isolate target singing voices from mixed audio accompanied by video. It probes the model's ability to leverage visual motion cues (face landmarks) to separate overlapping or low-volume singing voices across different languages and volume conditions. Use when the user wants to benchmark on Acappella, or asks about evaluating this task. Reports SDR.
- ▌ Activity Recognition Eval · qhjqhj00Evaluates a model's ability to recognize human activities in real-time by simultaneously learning from skeletal pose data and object attributes. It probes the integration of multi-modal cues (color, shape, distance, or object probabilities) for accurate and efficient activity classification in robotics scenarios. Use when the user wants to benchmark on Cornell Activity Dataset (CAD-60), MSR Daily Activity 3D Dataset, or asks about evaluating this task. Reports accuracy.
- ▌ Adaptive LLM Testing Eval · qhjqhj00This evaluation probes the effectiveness of diversity-based adaptive test selection strategies for black-box LLM applications. It measures how quickly and reliably different prioritization methods detect failures in prompt templates compared to random baselines, while also assessing the diversity of generated outputs. Use when the user wants to benchmark on BBH & P3 Prompt Templates, or asks about evaluating this task. Reports APFD.
- ▌ AI Paper Error Audit Eval · qhjqhj00Evaluates an LLM-based auditing system's ability to detect, categorize, and quantify objective mistakes in published AI research papers. It measures the system's precision against human verification and its recall against injected ground-truth errors across mathematical, textual, tabular, and cross-reference categories. Use when the user wants to benchmark on Published AI Papers (ICLR, NeurIPS, TMLR), or asks about evaluating this task. Reports precision.
- ▌ Alzheimer Mri 4class Eval · qhjqhj00Evaluates multi-class classification performance on Alzheimer's disease MRI scans to assess a model's ability to distinguish between different stages of dementia and healthy controls under resource-constrained hardware conditions. Use when the user wants to benchmark on Alzheimer MRI 4 Classes Dataset, or asks about evaluating this task. Reports Accuracy.