all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 24 of 76

  1. ▌
    Ms Marco V2 Ranking Eval · qhjqhj00
    Evaluates passage and document ranking systems on a large-scale, document-native corpus. It probes a model's ability to retrieve relevant content from millions of documents using sparse, crowd-sourced relevance judgments, while handling realistic corpus drift and query-independent passage extraction. Use when the user wants to benchmark on MS MARCO v2, or asks about evaluating this task. Reports NDCG@10.
    3 repo stars
  2. ▌
    Mtzig Cross Entropy Loss · qhjqhj00
    Compute mtzig/cross_entropy_loss via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of mtzig/cross_entropy_loss.
    3 repo stars
  3. ▌
    Multi Agent Latency Eval · qhjqhj00
    Evaluates the end-to-end latency, success rate, and cost-efficiency of a multi-agent LLM tutoring system under varying concurrency levels across different cloud inference throughput tiers. It isolates the impact of shared vs. priority vs. provisioned inference pools on response time variance and system reliability. Use when the user wants to benchmark on ITAS Student Query Corpus, or asks about evaluating this task. Reports end-to-end latency.
    3 repo stars
  4. ▌
    Multilingual Safety Eval · qhjqhj00
    Evaluates a parameter-efficient multilingual safety guardrail's ability to classify content as safe or unsafe across high-resource and low-resource languages. It probes cross-lingual generalization and robustness against diverse harm categories using cluster-guided transfer. Use when the user wants to benchmark on Aegis-Content-Safety-2.0-Test (Aegis-CS2), HarmBench, Redteam2k, JBB-Behaviors, StrongReject, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  5. ▌
    Multimodal Rec Benchmark · qhjqhj00
    Evaluates the performance of multimodal deep learning recommender systems across five Amazon product categories. It probes both standard recommendation accuracy and beyond-accuracy dimensions such as novelty, diversity, popularity bias, and catalog coverage. Use when the user wants to benchmark on Amazon (Office, Toys, Beauty, Sports, Clothing), or asks about evaluating this task. Reports Recall@k.
    3 repo stars
  6. ▌
    Multimodal Tool Use Eval · qhjqhj00
    Evaluates the ability of agentic multimodal models to perform visual perception, document understanding, and mathematical reasoning. It specifically probes whether models can strategically decide when to invoke external tools (e.g., image cropping, web search, Python code execution) versus answering directly, balancing task accuracy with tool efficiency. Use when the user wants to benchmark on V-Bench, HRBench-4K/8K, TreeBench, MME-RealWorld, SEEDBench2-Plus, CharXiv, MathVista_mini, MathVerse_mini, WeMath, DynaMath, LogicVista, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  7. ▌
    Multiple Choice Vqa Eval · qhjqhj00
    This evaluation probes the true multimodal reasoning capability of vision-language models on multiple-choice question answering tasks. It specifically measures whether models rely on actual question understanding or exploit visual relevance imbalances between correct answers and distractors. Performance is assessed under both standard (vision, question, options) and question-omitted (vision, options) settings to detect easy-option bias. Use when the user wants to benchmark on NExT-QA, MMStar, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  8. ▌
    Music Audio Tagging Eval · qhjqhj00
    This evaluation probes a model's ability to perform large-scale music audio tagging by predicting a fixed set of semantic labels (e.g., genre, mood, instrumentation) from raw 30-second audio clips. It measures how well architectures generalize across varying dataset sizes and label granularities. Use when the user wants to benchmark on MagnaTagATune (MTT), Million Song Dataset (MSD), Private Dataset, or asks about evaluating this task. Reports top-50 tag prediction.
    3 repo stars
  9. ▌
    Nemotron Nano V2 Vl Eval · qhjqhj00
    Evaluates a 12B vision-language model's capabilities across multimodal understanding, long-context reasoning, document/OCR processing, video comprehension, and pure text reasoning. It probes the model's ability to handle diverse visual inputs, follow instructions, and perform complex STEM and code reasoning under varying decoding and reasoning budget constraints. Use when the user wants to benchmark on MMBench V1.1, MMMU, OCRBench, DocVQA, LongVideoBench, MATH-500, GPQA-Diamond, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Nerf View Synthesis Eval · qhjqhj00
    Evaluates the ability of a neural radiance field to synthesize photorealistic novel views of 3D scenes from a sparse set of input images. It probes geometric reconstruction fidelity, appearance modeling (including non-Lambertian materials), and multi-view consistency across synthetic and real-world captures. Use when the user wants to benchmark on Diffuse Synthetic 360° (DeepVoxels), Realistic Synthetic 360°, Real Forward-Facing, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  11. ▌
    News Bias Detection Eval · qhjqhj00
    Evaluates transformer models' ability to classify news articles as biased or unbiased, comparing standard fine-tuning against domain-adapted training. It further probes model decision-making by analyzing word-level SHAP attribution magnitudes and lexical feature importance across true and false predictions. Use when the user wants to benchmark on BABE, or asks about evaluating this task. Reports Binary F1.
    3 repo stars
  12. ▌
    Nl Code Pair Mining Eval · qhjqhj00
    Evaluates a machine learning model's ability to automatically extract high-quality, aligned natural language intent and code snippet pairs from Stack Overflow posts. It probes the system's ranking capability, precision, and recall across different programming languages and feature combinations. Use when the user wants to benchmark on Stack Overflow NL-Code Pairs, or asks about evaluating this task. Reports AUC.
    3 repo stars
  13. ▌
    Nl Object Retrieval Eval · qhjqhj00
    Evaluates a model's ability to ground natural language queries to specific regions within images by scoring candidate bounding boxes. It probes spatial reasoning, contextual understanding, and cross-modal alignment between text and visual features. Use when the user wants to benchmark on ReferIt, Kitchen, or asks about evaluating this task. Reports P@1.
    3 repo stars
  14. ▌
    Nllb Clip Retrieval Eval · qhjqhj00
    Probes multilingual image-text retrieval capability across low-resource languages. It evaluates how effectively the model aligns visual and textual representations when trained with limited data and frozen encoders. Use when the user wants to benchmark on XTD200, Flickr30k-200, or asks about evaluating this task. Reports R@10.
    3 repo stars
  15. ▌
    Nllp 2024 Legal Nli Eval · qhjqhj00
    Probes an LLM's ability to perform Natural Language Inference (NLI) within the legal domain. Specifically, it tests whether the model can correctly classify the logical relationship (entailment, neutral, or contradiction) between a formal legal case summary and an informal social media review. Use when the user wants to benchmark on NLLP 2024 Legal NLI, or asks about evaluating this task. Reports F1.
    3 repo stars
  16. ▌
    Node Classification Eval · qhjqhj00
    This evaluation probes a model's ability to perform node classification on graphs across a spectrum of homophily regimes, from strongly heterophilic to strongly homophilic. It specifically tests the framework's adaptive capability to switch between a combinatorial predictor and a neural refinement stage based on validation performance. Use when the user wants to benchmark on Texas, Cornell, Actor, CiteSeer, Cora, Pubmed, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  17. ▌
    Optimam Mammography Eval · qhjqhj00
    Binary classification of high-resolution mammography images to detect malignant breast tissue. It probes the model's ability to distinguish between malignant and non-malignant cases using both localized patches and full-resolution inputs. Use when the user wants to benchmark on OPTIMAM, or asks about evaluating this task. Reports AUC.
    3 repo stars
  18. ▌
    Ota Firmware Update Eval · qhjqhj00
    Evaluates the energy efficiency and update latency of Over-The-Air (OTA) firmware update strategies on flash-based, batteryless IoT devices under simulated energy-harvesting conditions. Use when the user wants to benchmark on OTA Firmware Update Benchmarks, or asks about evaluating this task. Reports Total Update Energy Consumption.
    3 repo stars
  19. ▌
    Padt Unified Vision Eval · qhjqhj00
    Evaluates a multimodal large language model's ability to perform visual grounding, segmentation, open-vocabulary detection, and referring image captioning by predicting structured visual outputs directly from interleaved visual reference tokens and text. Use when the user wants to benchmark on RefCOCO/+/g, COCO 2017, RIC, or asks about evaluating this task. Reports IoU@0.5 accuracy.
    3 repo stars
  20. ▌
    Paired T Test Evaluation · qhjqhj00
    Tests whether a peer prediction mechanism's scoring function is sensitive to report quality by verifying that replacing high-quality reports with degraded or LLM-generated low-quality reports leads to a statistically significant decrease in expected scores. Use when the user has predictions and gold and needs to compute paired difference t-test (p-value).
    3 repo stars
  21. ▌
    Parrot Multilingual Eval · qhjqhj00
    Evaluates the multilingual visual-language understanding capabilities of multimodal large language models (MLLMs) across six languages (English, Chinese, Portuguese, Arabic, Turkish, Russian). It probes how well models align visual features with non-English textual instructions and handle cross-lingual multimodal tasks without relying on naive translation. Use when the user wants to benchmark on MMMB, MMBench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  22. ▌
    Pencil Puzzle Bench Eval · qhjqhj00
    Evaluates multi-step verifiable reasoning and agentic iteration on constraint-satisfaction puzzles. It probes a model's ability to plan, execute moves, check constraints step-by-step, and course-correct over long contexts. Use when the user wants to benchmark on Pencil Puzzle Bench, or asks about evaluating this task. Reports success rate.
    3 repo stars
  23. ▌
    Phonebit Inference Bench · qhjqhj00
    Evaluates the inference speed, energy efficiency, and classification accuracy of a GPU-accelerated binary neural network engine on mobile devices against standard frameworks. Use when the user has predictions and gold and needs to compute runtime.
    3 repo stars
  24. ▌
    Platinum Benchmarks Eval · qhjqhj00
    Evaluates LLM reliability on curated, low-noise subsets of standard benchmarks (VQA v2.0, SQuAD 2.0, HotPotQA, DROP, BIG-bench) by removing ambiguous examples and re-labeling to minimize ground-truth errors, revealing true model failures on elementary reasoning tasks. Use when the user wants to benchmark on VQA v2.0, SQuAD 2.0, HotPotQA, DROP, BIG-bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  25. ▌
    Polarmem Multimodal Eval · qhjqhj00
    Evaluates training-free multimodal agents on retrieval-augmented generation, general reasoning, and hallucination robustness by testing a polarized latent graph memory that injects logical constraints at inference time. Use when the user wants to benchmark on MRAMG-Bench, MRAG-Bench, Visual-RAG, MMMU, MMStar, HallusionBench, or asks about evaluating this task. Reports performance.
    3 repo stars
  26. ▌
    Principle Alignment Eval · qhjqhj00
    Probes an LLM's ability to align generated responses with a set of natural language constitutional principles without parameter fine-tuning. It measures both overall conformance quality and the reduction of critical principle violations through an inference-time self-correction pipeline. Use when the user wants to benchmark on SafeRLHF, HH-RLHF, or asks about evaluating this task. Reports 5-Point Likert Score Ranking.
    3 repo stars
  27. ▌
    Prior Loss Seg Benchmark · qhjqhj00
    Evaluates the effectiveness of various prior-based loss functions (low-level boundary/distance and high-level shape/size constraints) for medical image segmentation across diverse anatomical structures and imaging modalities. Use when the user wants to benchmark on WMH, ISLES, Atrium, Colon, Spleen, Hippocampus, Prostate, ACDC, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  28. ▌
    Prism Hallucination Eval · qhjqhj00
    Probes LLM hallucinations across four dimensions (knowledge missing, knowledge errors, reasoning errors, and instruction-following errors) by isolating error sources through controlled, task-specific queries. It evaluates model reliability and guides optimization by measuring error rates across memory, instruction, and reasoning generation stages. Use when the user wants to benchmark on PRISM, or asks about evaluating this task. Reports H-Score.
    3 repo stars
  29. ▌
    Procgen Competition Eval · qhjqhj00
    Evaluates reinforcement learning agents on their ability to learn efficiently and generalize to unseen, procedurally generated environments. It measures how well algorithms adapt to novel level distributions under strict computational and timestep constraints. Use when the user wants to benchmark on Procgen Benchmark, or asks about evaluating this task. Reports mean normalized return.
    3 repo stars
  30. ▌
    Pyroclast Benchmark Eval · qhjqhj00
    Evaluates the numerical correctness, stability, and parallel performance (CPU/GPU strong scaling and distributed weak scaling) of a matrix-free finite difference geodynamic solver. Use when the user wants to benchmark on Pyroclast Stokes & Advection Benchmarks, or asks about evaluating this task. Reports parallel speedup.
    3 repo stars
  31. ▌
    Quadsentinel Safety Eval · qhjqhj00
    This evaluation probes the ability of multi-agent guardrail systems to enforce machine-checkable safety policies over agent trajectories in real-time. It measures how effectively a system detects and blocks unsafe actions while minimizing false positives on enterprise web-agent and malicious behavior benchmarks. Use when the user wants to benchmark on ST-WebAgentBench, AgentHarm, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  32. ▌
    R1 Code Interpreter Eval · qhjqhj00
    Probes LLMs' ability to autonomously generate and execute code for complex reasoning and planning tasks across logic, spatial, order, optimization, search, and math domains. It evaluates how well models can iteratively explore, optimize, and self-check solutions in multi-turn code execution environments. Use when the user wants to benchmark on SymBench, Big-Bench-Hard, Reasoning-Gym, or asks about evaluating this task. Reports exact match or constraint check.
    3 repo stars
  33. ▌
    Radgpt Tumor Report Eval · qhjqhj00
    Evaluates an AI pipeline's ability to generate clinically accurate radiology reports from 3D CT scans by detecting tumors, measuring their size, localizing them within organ sub-segments, and staging cancers. It also assesses the textual similarity and diagnostic utility of generated reports compared to ground-truth clinical notes. Use when the user wants to benchmark on AbdomenAtlas 3.0, or asks about evaluating this task. Reports Tumor Detection Sensitivity & Specificity.
    3 repo stars
  34. ▌
    Reconstruction Pose Eval · qhjqhj00
    Evaluates how well different 3D reconstruction methods perform in a downstream object pose estimation task, rather than measuring standalone geometric reconstruction accuracy. It compares pose estimation results using reconstructed 3D models against those using ground-truth CAD models. Use when the user wants to benchmark on YCB-V, or asks about evaluating this task. Reports accuracy of the estimated poses.
    3 repo stars
  35. ▌
    Refugee Law Outcome Eval · qhjqhj00
    Evaluates NLP models' ability to perform legal information extraction (NER) and predict refugee claim decision outcomes from Canadian legal documents. It probes the model's capacity to handle domain-specific terminology, extract structured entities from unstructured text, and classify case outcomes based on judicial reasoning. Use when the user wants to benchmark on Canadian Refugee Status Determination (RSD) Cases, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  36. ▌
    Relative Decrease In Mae · qhjqhj00
    Evaluates the effectiveness of machine-learned preconditioners for linear solvers in semi-implicit shallow-water models by measuring error reduction in the first solver iteration and overall convergence rates during free-running simulations. Use when the user has predictions and gold and needs to compute Relative decrease in MAE.
    3 repo stars
  37. ▌
    Resel Scientific Ie Eval · qhjqhj00
    Evaluates a model's ability to perform N-ary relation extraction from scientific documents by first retrieving relevant text/table components and then selecting the correct entities within those components. Use when the user wants to benchmark on SciREX, PubMed, NLP-TDMS (Full), or asks about evaluating this task. Reports Accuracy (Acc).
    3 repo stars
  38. ▌
    Roboverse Imitation Eval · qhjqhj00
    Evaluates robot manipulation policies on a unified set of contact-rich pick-and-place and articulation tasks across multiple simulators. It probes both specialist and generalist vision-language-action models on success rates under standard and progressively challenging generalization levels. Use when the user wants to benchmark on ROBOVERSE Imitation Learning Benchmark, or asks about evaluating this task. Reports success rate.
    3 repo stars
  39. ▌
    Robustbench Cifar10 Eval · qhjqhj00
    Evaluates adversarial robustness of image classifiers under multi-norm threat models, testing whether pre-screening diagnostics (FOSC, RDI) reliably predict full attack performance and expose worst-case vulnerabilities masked by single-norm evaluations. Use when the user wants to benchmark on RobustBench CIFAR-10, or asks about evaluating this task. Reports robust_accuracy.
    3 repo stars
  40. ▌
    Roughness Index Distance · qhjqhj00
    Evaluates surface irregularities and topological consistency between predicted and ground-truth 3D medical segmentation masks. It quantifies local surface roughness, relative roughness differences, and average surface distance to detect spikes, holes, and smoothing artifacts. Use when the user has predictions and gold and needs to compute Roughness Index (RI).
    3 repo stars
  41. ▌
    Rowen Hallucination Eval · qhjqhj00
    This evaluation probes an LLM's ability to generate factually correct responses and mitigate hallucinations under adaptive retrieval-augmented generation. It measures factual accuracy via LLM-based scoring on open-ended questions and exact-match accuracy on multi-step reasoning yes/no questions. Use when the user wants to benchmark on TruthfulQA, StrategyQA, or asks about evaluating this task. Reports FactScore.
    3 repo stars
  42. ▌
    Rscc Change Caption Eval · qhjqhj00
    Evaluates vision-language models' ability to generate detailed, semantically accurate captions describing changes between bi-temporal remote sensing image pairs, particularly in disaster scenarios. It probes spatiotemporal reasoning, fine-grained environmental change detection, and long-text generation quality. Use when the user wants to benchmark on RSCC, or asks about evaluating this task. Reports ST5-SCS.
    3 repo stars
  43. ▌
    Saga 3d Mesh Attack Eval · qhjqhj00
    Evaluates the effectiveness of a spectral geometric adversarial attack on 3D mesh autoencoders by measuring how well perturbed meshes deceive a downstream classifier and evade detection. Use when the user wants to benchmark on CoMA, SMAL, or asks about evaluating this task. Reports Targeted classification accuracy.
    3 repo stars
  44. ▌
    Saplma Truthfulness Eval · qhjqhj00
    Evaluates whether an LLM's internal hidden layer activations can predict the veracity of a given statement. It probes the model's implicit knowledge of truthfulness by training a classifier on neural activations rather than relying on explicit prompting or output probabilities. Use when the user wants to benchmark on True-False Dataset, LLM-Generated Statements, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  45. ▌
    Scandeval Benchmark Eval · qhjqhj00
    Evaluates the performance of monolingual and multilingual language models across five Scandinavian languages (Danish, Norwegian, Swedish, Icelandic, Faroese) on question answering, linguistic acceptability, and named entity recognition. It also probes cross-lingual transfer capabilities between these languages by measuring performance variance across language groups. Use when the user wants to benchmark on ScandiQA, ScaLA, MIM-GOLD-NER, WikiANN, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  46. ▌
    Schema Guided Dstc8 Eval · qhjqhj00
    Evaluates zero-shot dialogue state tracking across single and multi-domain conversations. It measures the model's ability to predict intents, extract slot values, and maintain accurate dialogue states over long contexts without prior exposure to unseen service domains. Use when the user wants to benchmark on Schema-Guided Dialogue (DSTC8 Track 4), or asks about evaluating this task. Reports Joint Goal Accuracy.
    3 repo stars
  47. ▌
    Scientific Ideation Eval · qhjqhj00
    Evaluates the ability of LLMs to generate novel, feasible, and effective scientific research ideas given a research question. Probes open-ended scientific reasoning and ideation quality under compute-matched inference budgets. Use when the user wants to benchmark on ICLR 2024 & NeurIPS 2025, or asks about evaluating this task. Reports Absolute Novelty.
    3 repo stars
  48. ▌
    Scmamba Integration Eval · qhjqhj00
    Evaluates computational methods for single-cell multi-omics integration by measuring their ability to preserve biological variation, align different omics layers at the cell and single-cell levels, and enable accurate cell type annotation. Use when the user wants to benchmark on SHARE-seq BMMC, SNARE-seq, 10X Genomics Multiome, Human fetal atlas, CITE-seq BMMC S1, CITE-seq BMMC S4, Human brain multi-omics, Human brain 3k, or asks about evaluating this task. Reports biological variation conservation score, omics alignment score, overall integration score.
    3 repo stars
  49. ▌
    Secure Inference Latency · qhjqhj00
    Measures the online and offline computation latency and communication bandwidth for cryptographic primitives and neural network operations under secure two-party computation. It evaluates how efficiently packed homomorphic encryption and garbled circuits handle matrix-vector products, convolutions, and activation functions without revealing inputs or model parameters. Use when the user has predictions and gold and needs to compute t_online.
    3 repo stars
  50. ▌
    Semeval 2023 Task12 Eval · qhjqhj00
    Sentiment classification across twelve low-resource African languages and Creoles, evaluating model robustness to code-switching and varying degrees of lexical similarity to pretraining data. Use when the user wants to benchmark on SemEval-2023 Task 12, or asks about evaluating this task. Reports macro-F1.
    3 repo stars
  51. ▌
    Semigda Medical Seg Eval · qhjqhj00
    Evaluates semi-supervised medical image segmentation performance under limited labeled data ratios (10% and 30%). It measures segmentation accuracy and boundary precision across multiple medical domains including colonoscopy, dermoscopy, pathology, and ultrasound. Use when the user wants to benchmark on Colonoscopy (CVC-ClinicDB, Kvasir, CVC-300), ISIC-2018, BCSS, BUSI, or asks about evaluating this task. Reports Dice coefficient (Dice).
    3 repo stars
  52. ▌
    Sensitivityatspecificity · qhjqhj00
    Compute the SensitivityAtSpecificity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SensitivityAtSpecificity, or asks how to score with SensitivityAtSpecificity.
    3 repo stars
  53. ▌
    Shalakasatheesh Squad V2 · qhjqhj00
    Compute shalakasatheesh/squad_v2 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of shalakasatheesh/squad_v2.
    3 repo stars
  54. ▌
    Shirt Pose Tracking Eval · qhjqhj00
    Evaluates the accuracy and robustness of a neural network-integrated Unscented Kalman Filter for monocular pose tracking of tumbling noncooperative spacecraft. It probes the system's ability to maintain steady-state position and orientation accuracy under domain gaps between synthetic training data and real hardware-in-the-loop test images. Use when the user wants to benchmark on SHIRT, or asks about evaluating this task. Reports e_pose.
    3 repo stars
  55. ▌
    Shpi Recommendation Eval · qhjqhj00
    Evaluates offline reinforcement learning methods for session-based recommendation systems in optimizing long-term user retention versus short-term clicks. It tests the ability of algorithms to learn from fixed logging policies and generalize to online rollouts across synthetic, simulated, and real-world recommendation environments. Use when the user wants to benchmark on Synthetic recommendation problem, RecoGym, HIV treatment simulator, Private dataset X, or asks about evaluating this task. Reports undiscounted test performance on true environment rewards.
    3 repo stars
  56. ▌
    Simmotion Retrieval Eval · qhjqhj00
    This benchmark evaluates a model's ability to retrieve videos based on semantic motion similarity, disentangling dynamic behavior from static appearance, camera viewpoint, and scene context. It tests robustness to appearance variations in controlled synthetic settings and unsynchronized, in-the-wild video pairs. Use when the user wants to benchmark on SimMotion-Synthetic, SimMotion-Real-1K, Jester, or asks about evaluating this task. Reports Retrieval accuracy.
    3 repo stars
  57. ▌
    Slamming Additional Eval · qhjqhj00
    Evaluates the speech language model's performance across multiple benchmarks including linguistic acceptability, story completion, audio generation quality, and cross-domain text generation perplexity. Use when the user wants to benchmark on sBLIMP, StoryCloze, People Speech, or asks about evaluating this task. Reports MOSnet.
    3 repo stars
  58. ▌
    Soi Id Ood Accuracy Eval · qhjqhj00
    Evaluates pretrained language models' in-distribution (ID) and out-of-distribution (OOD) classification accuracy under single-setting and multi-setting fine-tuning configurations. It probes how training dynamics and subset selection affect robustness and generalization across languages, sources, and tasks. Use when the user wants to benchmark on SST-2, IMDB, Yelp, Sentiment140, RTE, QQP, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  59. ▌
    Space Ris Satellite Eval · qhjqhj00
    Evaluates a multi-agent DRL framework (MAPPO) combined with whale optimization for maximizing satellite coverage and data rates in 6G sub-THz networks using reconfigurable intelligent surfaces (RIS). Use when the user wants to benchmark on Simulated LEO Satellite-RIS Network Environment, or asks about evaluating this task. Reports average data rate.
    3 repo stars
  60. ▌
    Specificityatsensitivity · qhjqhj00
    Compute the SpecificityAtSensitivity metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpecificityAtSensitivity, or asks how to score with SpecificityAtSensitivity.
    3 repo stars
  61. ▌
    Speechinstructbench Eval · qhjqhj00
    Evaluates speech instruction-following capabilities across closed-ended, open-ended, and adjustment tasks under varying acoustic conditions (background noise, accents, disfluencies) in English and Chinese. Use when the user wants to benchmark on SpeechInstructBench, or asks about evaluating this task. Reports instruction-level accuracy (I).
    3 repo stars
  62. ▌
    Ssrs Retrosynthesis Eval · qhjqhj00
    Evaluates single-step retrosynthesis capability by predicting reactant molecules from a given target product, testing both in-distribution chemical knowledge and out-of-distribution generalization. Use when the user wants to benchmark on USPTO-50K-test, URSA-expert-2026, or asks about evaluating this task. Reports Unique.
    3 repo stars
  63. ▌
    Structured3d Layout Eval · qhjqhj00
    Evaluates a model's ability to predict architectural elements (walls, doors, windows) and room layouts within indoor 3D scenes. It tests the model's capacity for structured scene understanding and spatial reasoning by comparing predicted layouts against ground-truth annotations. Use when the user wants to benchmark on Structured3D, or asks about evaluating this task. Reports F1.
    3 repo stars
  64. ▌
    Synthetic Drug Data Eval · qhjqhj00
    Evaluates the distributional fidelity of generated pharmacokinetic and drug-target interaction properties against real data, and measures the utility of the synthetic data for downstream regression tasks. Use when the user wants to benchmark on TDCommons/BindingDB PK & DTI Collection, or asks about evaluating this task. Reports Hellinger Distance (HD).
    3 repo stars
  65. ▌
    T2i Deanonymization Eval · qhjqhj00
    Evaluates the ability to deanonymize text-to-image models by identifying which model generated a given image, exploiting model-specific visual signatures in embedding space. Use when the user wants to benchmark on T2I Leaderboard Prompts, or asks about evaluating this task. Reports Top-1 accuracy.
    3 repo stars
  66. ▌
    Table Row Detection Eval · qhjqhj00
    Evaluates graph-based machine learning models for sequence labeling (BIESO) and table row detection on handwritten historical register books. The protocol tests the models' ability to segment table rows and label cell boundaries using pre-extracted textline and column features rather than raw images. Use when the user wants to benchmark on Dataset1, Dataset2, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  67. ▌
    Tanks Temples Truck Eval · qhjqhj00
    Evaluates novel view synthesis quality and 3D geometric reconstruction accuracy of NeRF variants on a single outdoor scene. Use when the user wants to benchmark on Tanks & Temples (Truck scene), or asks about evaluating this task. Reports CD.
    3 repo stars
  68. ▌
    Tcga Histopathology Eval · qhjqhj00
    Evaluates the representation quality and generalization of self-supervised histopathology models across diverse patch-level diagnostic tasks and weakly supervised slide-level tasks using linear probing and fine-tuning on TCGA whole slide images. Use when the user wants to benchmark on TCGA Histopathology, or asks about evaluating this task. Reports average AUC.
    3 repo stars
  69. ▌
    Text Classification Eval · qhjqhj00
    Evaluates text classification performance across multiple sentiment, subjectivity, question classification, and topic categorization tasks. It probes the model's ability to capture contextual and syntactic features from sequential text using 2D matrix representations and spatial pooling. Use when the user wants to benchmark on MR, SST-1, SST-2, Subj, TREC, 20Newsgroups, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  70. ▌
    Tfbs Classification Eval · qhjqhj00
    Evaluates a model's ability to classify short DNA sequences as transcription factor binding sites or not, capturing its capacity to learn regulatory sequence patterns from genomic data. Use when the user wants to benchmark on TFBS classification, or asks about evaluating this task. Reports AUC.
    3 repo stars
  71. ▌
    Themisio Io Sharing Eval · qhjqhj00
    Evaluates a policy-driven I/O sharing framework for burst buffers by measuring how effectively it allocates bandwidth, maintains fairness, and reduces interference across concurrent workloads. It probes the system's ability to enforce primitive and composite sharing policies under varying load conditions and compares performance against baseline schedulers. Use when the user wants to benchmark on ThemisIO Benchmark & Application Suite, or asks about evaluating this task. Reports sustained I/O throughput.
    3 repo stars
  72. ▌
    Threat Intelligence Eval · qhjqhj00
    This benchmark evaluates an AI system's ability to extract actionable insights from threat intelligence reports and perform security reasoning. It probes multi-document comprehension, attack chain reconstruction, and MITRE ATT&CK framework mapping capabilities. Use when the user wants to benchmark on CyberSOCEval Threat Intelligence Reasoning, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  73. ▌
    Tongue Segmentation Eval · qhjqhj00
    Evaluates zero-shot and cross-dataset generalization of a tongue segmentation model adapted from SAM. It probes the model's ability to segment tongue regions in medical images without task-specific fine-tuning on the target datasets. Use when the user wants to benchmark on TongueSet1, BioHit (TongueSet2), Webset (TongueSet3), or asks about evaluating this task. Reports mIoU.
    3 repo stars
  74. ▌
    Toto TS Forecasting Eval · qhjqhj00
    Evaluates zero-shot and fine-tuned time series forecasting capabilities on real-world observability telemetry and general-purpose benchmarks. Probes model robustness to high-dimensional, nonstationary, multivariate series with skewed distributions and varying temporal intervals. Use when the user wants to benchmark on Boom, Boomlet, GIFT-Eval, LSF, or asks about evaluating this task. Reports CRPS.
    3 repo stars
  75. ▌
    Trec Session Tracks Eval · qhjqhj00
    Evaluates the effectiveness of a learning-to-rank personalization approach for web search sessions by measuring relevance prediction quality across multiple years of session track data. It probes the model's ability to leverage historical query sequences, document rankings, and user click behavior to improve session-level relevance ranking. Use when the user wants to benchmark on TREC 2011-2014 Session Tracks, or asks about evaluating this task. Reports nDCG@k.
    3 repo stars
  76. ▌
    Trec2024 RAG Nugget Eval · qhjqhj00
    Evaluates the factual accuracy and content grounding of RAG-generated answers by checking for the presence of key factual claims (nuggets) extracted from source documents. It also measures answer length to assess the trade-off between conciseness and completeness in system outputs. Use when the user wants to benchmark on TREC 2024 RAG Track, or asks about evaluating this task. Reports V_strict.
    3 repo stars
  77. ▌
    Umlip High Temp Mof Eval · qhjqhj00
    Evaluates the accuracy of universal machine-learned interatomic potentials (uMLIPs) in predicting energy, forces, and stress tensors during high-temperature molecular dynamics simulations of metal-organic frameworks (MOFs), including their stability and thermal decomposition behavior. Use when the user wants to benchmark on High-Temperature MOF AIMD Benchmark, or asks about evaluating this task. Reports energy MAE.
    3 repo stars
  78. ▌
    Unet Biomedical Seg Eval · qhjqhj00
    Evaluates pixel-level biomedical image segmentation capability using convolutional networks. Probes the model's ability to precisely delineate cellular structures and membranes in electron and light microscopy images with limited training data. Use when the user wants to benchmark on EM segmentation challenge (ISBI 2012), PhC-U373, DIC-HeLa, or asks about evaluating this task. Reports warping error, IOU.
    3 repo stars
  79. ▌
    Us Grid Forecasting Eval · qhjqhj00
    Evaluates the ability of deep learning architectures (SSMs, Transformers, RNNs) to forecast hourly electricity load across major US power grids. It probes how well models capture temporal patterns, handle varying prediction horizons, and integrate exogenous weather covariates for accurate grid-scale forecasting. Use when the user wants to benchmark on US ISO Hourly Load Data (EIA-930), or asks about evaluating this task. Reports MSE (%).
    3 repo stars
  80. ▌
    Velocity Dealiasing Eval · qhjqhj00
    Evaluates a U-Net model's ability to predict velocity fold numbers and produce dealiased radar velocity fields from folded inputs. It measures both classification accuracy for fold detection and reconstruction fidelity via velocity error metrics. Use when the user wants to benchmark on WSR-88D Level-II/III Radar Data, or asks about evaluating this task. Reports velocity RMSE.
    3 repo stars
  81. ▌
    Verafi Financial QA Eval · qhjqhj00
    Probes an agentic RAG system's ability to retrieve relevant SEC filings and generate factually correct, complete financial answers. It specifically tests the impact of neurosymbolic policy validation on suppressing hallucinations and mathematical errors in high-stakes financial domains. Use when the user wants to benchmark on FinanceBench-style Financial QA Dataset, or asks about evaluating this task. Reports Factual Correctness.
    3 repo stars
  82. ▌
    Video Thinking Test Eval · qhjqhj00
    Evaluates video large language models on their ability to understand complex visual narratives and answer questions correctly. It specifically probes robustness by testing model performance on naturally adversarial or misleading variations of the same video question. Use when the user wants to benchmark on Video Thinking Test, or asks about evaluating this task. Reports Correctness score (accuracy).
    3 repo stars
  83. ▌
    Vincicoder Code Gen Eval · qhjqhj00
    Probes a model's ability to generate executable, visually faithful code (HTML, SVG, LaTeX, SMILES) from input images across diverse domains. It evaluates both syntactic correctness via execution rate and perceptual alignment with target images using coarse-to-fine visual similarity metrics. Use when the user wants to benchmark on ChartMimic, Design2Code, UniSVG, Image2Struct, Cosyn-400k, or asks about evaluating this task. Reports UniSVG Final Score.
    3 repo stars
  84. ▌
    Vision Language Ood Eval · qhjqhj00
    Probes the ability of vision-language models to distinguish in-distribution from out-of-distribution samples under semantic, covariate, and real-world distribution shifts. It evaluates both zero-shot and few-shot prompt learning approaches across multiple benchmarks to assess robustness and ranking consistency. Use when the user wants to benchmark on ImageNet-X, ImageNet-FS-X, Wilds-FS-X, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  85. ▌
    Visual Wetlandbirds Eval · qhjqhj00
    This benchmark evaluates deep learning models on fine-grained bird species classification and spatio-temporal behavior recognition in ecological video footage. It probes the model's ability to localize birds, identify their species, and classify their actions across video frames in real-world wetland environments. Use when the user wants to benchmark on Visual WetlandBirds Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  86. ▌
    Vl Compositionality Eval · qhjqhj00
    Evaluates vision-language models on compositional reasoning capabilities, specifically testing their ability to correctly bind attributes, understand semantic relations, and parse word order in image-text pairs. It also measures systematic generalization to unseen concept combinations and zero-shot classification and retrieval performance. Use when the user wants to benchmark on ARO, CREPE, SVO, VL-Checklist, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  87. ▌
    Wals Metalinguistic Eval · qhjqhj00
    Probes large language models' ability to recall and identify structural and grammatical properties of languages across diverse linguistic domains. It measures whether models have internalized typological facts from the World Atlas of Language Structures (WALS) by answering multiple-choice questions about specific language features. Use when the user wants to benchmark on WALS, WALS-100, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  88. ▌
    Weak Annotation Har Eval · qhjqhj00
    Evaluates inertial-based activity recognition models trained on weakly-supervised labels generated via vision foundation model clustering, benchmarked against fully-supervised and few-shot baselines. Use when the user wants to benchmark on WEAR, Wetlab, ActionSense, or asks about evaluating this task. Reports Acc.
    3 repo stars
  89. ▌
    Web Agent Benchmark Eval · qhjqhj00
    Evaluates long-horizon web navigation and information-seeking capabilities. Probes the agent's ability to formulate search queries, browse multiple web pages, synthesize information from diverse sources, and answer complex multi-step questions. Use when the user wants to benchmark on BrowseComp-en, BrowseComp-zh, GAIA, WebWalkerQA, FRAMES, XBench-DeepSearch, HLE, or asks about evaluating this task. Reports Avg@4 Accuracy.
    3 repo stars
  90. ▌
    Wind Power Ensemble Eval · qhjqhj00
    Evaluates the calibration, sharpness, and accuracy of probabilistic wind power forecasts under different ensemble post-processing strategies (raw, weather-only, power-only, and joint weather-power post-processing). It probes whether correcting biases at the weather stage alone is sufficient, or if direct post-processing of the final power ensemble is required to handle non-linear power curve biases. Use when the user wants to benchmark on Benchmark Data, Swedish Data Set, or asks about evaluating this task. Reports CRPS.
    3 repo stars
  91. ▌
    Wmt17 Nmt Benchmark Eval · qhjqhj00
    Evaluates the training efficiency, inference speed, and translation quality of neural machine translation systems on standard WMT17 benchmarks. It probes the trade-offs between model architecture, hardware acceleration (FP16/INT8), and batching strategies. Use when the user wants to benchmark on WMT17 English-German, WMT17 Russian-English, or asks about evaluating this task. Reports BLEU.
    3 repo stars
  92. ▌
    Workload Allocation Eval · qhjqhj00
    Evaluates the latency performance of AI workload allocation strategies across hierarchical cloud/edge/device computing environments for latency-sensitive medical ICU applications. It measures how effectively dynamic routing minimizes end-to-end response time when processing and transmission delays are factored in. Use when the user wants to benchmark on Edge AIBench ICU Applications (MIMIC-III derived), or asks about evaluating this task. Reports response time.
    3 repo stars
  93. ▌
    World Music Corpora Eval · qhjqhj00
    Evaluates audio foundation models' cross-cultural generalization across diverse musical traditions (Western, Greek, Turkish, Indian) using multi-label tagging and few-shot learning. Probes whether pre-trained representations capture cultural musical knowledge without extensive adaptation. Use when the user wants to benchmark on Turkish-makam, Hindustani, Carnatic, MagnaTagATune, FMA-medium, Lyra, or asks about evaluating this task. Reports ROC-AUC.
    3 repo stars
  94. ▌
    3d Spatial Reasoning Eval · qhjqhj00
    Evaluates a model's ability to perform 3D visual grounding and situated question answering by reasoning over object coordinates and spatial relations in 3D scenes. It probes whether the model can accurately locate objects based on natural language instructions and answer spatial questions about scene layouts without linguistic interference. Use when the user wants to benchmark on ScanRefer, Multi3DRef, SQA3D, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  95. ▌
    6dof Camera Tracking Eval · qhjqhj00
    Evaluates the tracking accuracy of a 6-DoF autonomous camera algorithm in a simulated surgical environment. It also measures how different camera control strategies impact human rater accuracy when assessing surgical skill from video. Use when the user wants to benchmark on da Vinci wire chaser simulation, or asks about evaluating this task. Reports assessment_error.
    3 repo stars
  96. ▌
    Acappella Separation Eval · qhjqhj00
    This benchmark evaluates audio-visual singing voice separation models by measuring how accurately they isolate target singing voices from mixed audio accompanied by video. It probes the model's ability to leverage visual motion cues (face landmarks) to separate overlapping or low-volume singing voices across different languages and volume conditions. Use when the user wants to benchmark on Acappella, or asks about evaluating this task. Reports SDR.
    3 repo stars
  97. ▌
    Activity Recognition Eval · qhjqhj00
    Evaluates a model's ability to recognize human activities in real-time by simultaneously learning from skeletal pose data and object attributes. It probes the integration of multi-modal cues (color, shape, distance, or object probabilities) for accurate and efficient activity classification in robotics scenarios. Use when the user wants to benchmark on Cornell Activity Dataset (CAD-60), MSR Daily Activity 3D Dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  98. ▌
    Adaptive LLM Testing Eval · qhjqhj00
    This evaluation probes the effectiveness of diversity-based adaptive test selection strategies for black-box LLM applications. It measures how quickly and reliably different prioritization methods detect failures in prompt templates compared to random baselines, while also assessing the diversity of generated outputs. Use when the user wants to benchmark on BBH & P3 Prompt Templates, or asks about evaluating this task. Reports APFD.
    3 repo stars
  99. ▌
    AI Paper Error Audit Eval · qhjqhj00
    Evaluates an LLM-based auditing system's ability to detect, categorize, and quantify objective mistakes in published AI research papers. It measures the system's precision against human verification and its recall against injected ground-truth errors across mathematical, textual, tabular, and cross-reference categories. Use when the user wants to benchmark on Published AI Papers (ICLR, NeurIPS, TMLR), or asks about evaluating this task. Reports precision.
    3 repo stars
  100. ▌
    Alzheimer Mri 4class Eval · qhjqhj00
    Evaluates multi-class classification performance on Alzheimer's disease MRI scans to assess a model's ability to distinguish between different stages of dementia and healthy controls under resource-constrained hardware conditions. Use when the user wants to benchmark on Alzheimer MRI 4 Classes Dataset, or asks about evaluating this task. Reports Accuracy.
    3 repo stars