all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 34 of 76

  1. ▌
    Controllable Gen Eval · qhjqhj00
    Evaluates controllable image generation based on visual conditions (segmentation masks, edges, depth maps) by measuring how closely the generated image's extracted conditions match the input conditions. It tests spatial and structural controllability. Use when the user wants to benchmark on ControlNet++ dataset, or asks about evaluating this task. Reports mIoU (Seg. Mask).
    3 repo stars
  2. ▌
    Cosmoflow Hpc Scaling · qhjqhj00
    Evaluates the compute efficiency and horizontal scalability of a 3D convolutional neural network framework on supercomputers, measuring sustained floating-point throughput and parallel scaling efficiency across thousands of nodes. Use when the user has predictions and gold and needs to compute Pflop/s.
    3 repo stars
  3. ▌
    Cot Faithfulness Eval · qhjqhj00
    Evaluates whether reasoning models explicitly acknowledge external hint injections within their chain-of-thought reasoning traces. It probes model transparency and the alignment between internal reasoning tokens and final output disclosures. Use when the user wants to benchmark on MMLU, GPQA Diamond, or asks about evaluating this task. Reports faithfulness.
    3 repo stars
  4. ▌
    Counterfact Edit Eval · qhjqhj00
    Evaluates the effectiveness and stability of sequential LLM knowledge editing methods. It measures how well a model updates a specific fact while preserving related paraphrases, neighboring facts, generation fluency, and overall general capabilities over thousands of edits. Use when the user wants to benchmark on CounterFact, GLUE_MMLU_GSM8K_HumanEval_MBPP, or asks about evaluating this task. Reports Efficacy.
    3 repo stars
  5. ▌
    Creation Mmbench Eval · qhjqhj00
    Evaluates context-aware creative intelligence in multimodal and text-only models by assessing their ability to generate creative, contextually relevant content while maintaining visual factuality across diverse functional and creative writing tasks. Use when the user wants to benchmark on Creation-MMBench, or asks about evaluating this task. Reports VFS, Reward.
    3 repo stars
  6. ▌
    Crews Ews Budget Eval · qhjqhj00
    Evaluates an AI system's ability to extract and classify budget allocations for Early Warning System (EWS) investments from heterogeneous financial PDF reports. It probes multi-label classification, numerical budget extraction with tolerance, and evidence retrieval/mapping in climate finance contexts. Use when the user wants to benchmark on MDB Evidence Set, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  7. ▌
    Cross Domain Ctr Eval · qhjqhj00
    Evaluates cross-domain knowledge transfer for click-through rate (CTR) prediction by measuring how well a model trained on a source domain generalizes to a target domain with non-overlapping features. It probes context-aware feature translation and explicit knowledge augmentation in recommendation systems. Use when the user wants to benchmark on Amazon, Taobao, Alibaba Production, or asks about evaluating this task. Reports AUC.
    3 repo stars
  8. ▌
    Crossdocked Sbdd Eval · qhjqhj00
    Evaluates a model's ability to generate novel, drug-like molecules with high binding affinity for unseen protein pockets in structure-based drug design. It probes the trade-offs between binding energy, molecular properties, and synthesis feasibility. Use when the user wants to benchmark on CrossDocked-100k, or asks about evaluating this task. Reports Vina Dock.
    3 repo stars
  9. ▌
    Crosslingual Mtf Eval · qhjqhj00
    Evaluates zero-shot crosslingual generalization of multilingual LLMs after multitask finetuning. Probes language-agnostic task understanding, robustness to prompt translation, and scaling behavior across NLU, generative, and code tasks. Use when the user wants to benchmark on XNLI, XCOPA, XStoryCloze, XWinograd, HumanEval, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  10. ▌
    Crosspoint Bench Eval · qhjqhj00
    Evaluates Vision-Language Models' ability to perform precise point-level geometric correspondence across multiple viewpoints. It probes fine-grained spatial grounding, visibility reasoning, cross-view correspondence judgment, and continuous 2D coordinate pointing. Use when the user wants to benchmark on CrossPoint-Bench, or asks about evaluating this task. Reports average accuracy.
    3 repo stars
  11. ▌
    Cti Plausibility Eval · qhjqhj00
    Evaluates whether neural machine translation models correctly rely on contextual cues when generating target tokens. It compares model-extracted cue-target pairs against human-annotated discourse-level expectations to measure the plausibility of context reliance. Use when the user wants to benchmark on SCAT+, or asks about evaluating this task. Reports Macro F1.
    3 repo stars
  12. ▌
    Curriculum Dpo Eval · qhjqhj00
    Evaluates text-to-image generation models on their ability to align generated images with text prompts, produce visually appealing outputs, and match human preferences. It tests the effectiveness of curriculum-based fine-tuning strategies on standard generative benchmarks. Use when the user wants to benchmark on D1 (Black-ICLR-2024), D2 (DrawBench), D3 (Pick-a-Pic), or asks about evaluating this task. Reports Text Alignment.
    3 repo stars
  13. ▌
    Pseldnets Seld Eval · qhjqhj00
    Evaluates sound event localization and detection (SELD) performance on synthetic and real-world audio, measuring classification accuracy, localization precision, and overall detection quality across different network architectures and fine-tuning strategies. Use when the user wants to benchmark on synthetic-test-set, synthetic-training-set, Indoor Recordings, or asks about evaluating this task. Reports SELD.
    3 repo stars
  14. ▌
    Puma Challenge Eval · qhjqhj00
    Evaluates pixel-level segmentation capability for distinguishing five histopathological tissue classes (tumour, stroma, necrosis, blood vessels, epidermis) in melanoma H&E images. Use when the user wants to benchmark on PUMA Challenge dataset, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  15. ▌
    Qags Questeval Eval · qhjqhj00
    Evaluates the factual consistency and information coverage of query-focused summaries for less-resourced languages without reference texts. It probes whether LLMs can preserve source details and align with user intent by comparing answers generated from the source versus answers generated from the candidate summary. Use when the user wants to benchmark on Slovene News Summarization Corpus (MOCHA translation), or asks about evaluating this task. Reports QuestEval F1.
    3 repo stars
  16. ▌
    Quan Temp Plus Eval · qhjqhj00
    Evaluates open-domain numerical fact-checking by testing how well models can verify claims using decomposed queries and retrieved evidence. It probes the impact of claim decomposition quality on evidence retrieval and downstream NLI-based verification accuracy. Use when the user wants to benchmark on QuanTemp++, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  17. ▌
    Rad Robustness Eval · qhjqhj00
    Evaluates the robustness of image anomaly detection models against real-world imaging distortions, including free viewpoints, uneven illumination, and motion blur. It measures how well unsupervised and zero-shot methods localize and classify anomalies on industrial work platforms with foreign objects. Use when the user wants to benchmark on RAD, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  18. ▌
    RAG Robustness Eval · qhjqhj00
    Evaluates how Retrieval-Augmented Generation (RAG) systems maintain factual accuracy when exposed to adversarial, harmful, or misleading medical evidence. It probes the model's susceptibility to contextual manipulation and its ability to resist misinformation propagation under varying query framings. Use when the user wants to benchmark on TREC Health Misinformation 2020, TREC Health Misinformation 2021, or asks about evaluating this task. Reports ground-truth alignment rate.
    3 repo stars
  19. ▌
    Red Diffeq Fwi Eval · qhjqhj00
    Evaluates the ability of a diffusion-based regularization framework to reconstruct high-resolution subsurface velocity models from seismic data. It probes robustness under varying data conditions, including clean recordings, Gaussian noise contamination, and missing traces. The benchmark also tests out-of-distribution generalization on complex geological structures. Use when the user wants to benchmark on OpenFWI, Marmousi, or asks about evaluating this task. Reports RMSE.
    3 repo stars
  20. ▌
    Remote Sensing Eval · qhjqhj00
    Evaluates in-domain representation learning and scene classification across diverse remote sensing modalities (optical, SAR, aerial) and spatial resolutions. It probes model robustness to varying class balances, visual similarities, and label domains in Earth observation. Use when the user wants to benchmark on BigEarthNet, EuroSAT, RESISC-45, So2Sat, UC Merced, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  21. ▌
    Reta Benchmark Eval · qhjqhj00
    Evaluates the accuracy and structural consistency of retinal vascular tree annotations across pixel, vessel segment, and network levels, ensuring topological correctness and geometrical plausibility. Use when the user wants to benchmark on RETA Benchmark, or asks about evaluating this task. Reports multi_stage_annotation.
    3 repo stars
  22. ▌
    Retrievalrprecision · qhjqhj00
    Compute the RetrievalRPrecision metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RetrievalRPrecision, or asks how to score with RetrievalRPrecision.
    3 repo stars
  23. ▌
    Review Quality Eval · qhjqhj00
    Evaluates the quality of academic peer review reports across different conferences and years using a multi-dimensional framework. It measures how substantive, actionable, and well-grounded reviews are, and tracks whether these qualities decline over time. Use when the user wants to benchmark on Peer Review Campaigns (ICLR, NeurIPS, ACL), or asks about evaluating this task. Reports Q.
    3 repo stars
  24. ▌
    Rlvr Reasoning Eval · qhjqhj00
    This evaluation probes the mathematical and out-of-domain reasoning capabilities of large language models trained with Reinforcement Learning with Verifiable Rewards (RLVR). It specifically tests how well entropy-aware credit assignment methods allocate learning signals across high-entropy tokens during chain-of-thought generation. Performance is measured by average accuracy and pass rate over multiple sampled reasoning paths. Use when the user wants to benchmark on AIME24, AIME25, AMC, MATH, Minerva, Olympiad, or asks about evaluating this task. Reports Avg@k, Pass@k.
    3 repo stars
  25. ▌
    Rna 3d Scoring Eval · qhjqhj00
    Evaluates a model's ability to score and rank candidate 3D RNA structural models by predicting their deviation from the true native structure. It probes the model's capacity to distinguish accurate conformations from decoys using only atomic coordinates and types. Use when the user wants to benchmark on RNA-Puzzles & FARFAR2 Decoys, or asks about evaluating this task. Reports RMSD.
    3 repo stars
  26. ▌
    Robodrivebench Eval · qhjqhj00
    Evaluates the robustness and safety of vision-language models (VLMs) for end-to-end autonomous driving when subjected to real-world sensor corruptions (e.g., fog, rain, motion blur) and prompt corruptions (e.g., bit errors, malicious attacks). It probes the model's ability to maintain accurate trajectory prediction and low collision rates under degraded inputs. Use when the user wants to benchmark on RoboDriveBench, or asks about evaluating this task. Reports AvgL2.
    3 repo stars
  27. ▌
    Robust Asr Wer Eval · qhjqhj00
    Evaluates the robustness of end-to-end automatic speech recognition models against various stationary and non-stationary noise types at different signal-to-noise ratios (SNR). It also measures the degradation of recognition accuracy on clean speech when noise-adaptation techniques are applied. Use when the user wants to benchmark on Custom noisy speech dataset (7 noise types), or asks about evaluating this task. Reports WER.
    3 repo stars
  28. ▌
    Roundabout Tau Eval · qhjqhj00
    This benchmark evaluates a model's ability to detect traffic anomalies in roadside surveillance videos and generate detailed, reasoning-grounded textual summaries of those anomalies. It probes both binary/fine-grained classification accuracy and semantic alignment of generated descriptions with ground-truth event narratives. Use when the user wants to benchmark on Roundabout-TAU, or asks about evaluating this task. Reports 4-cls AP.
    3 repo stars
  29. ▌
    Rsf7 Benchmark Eval · qhjqhj00
    Evaluates the optimization capability of evolutionary algorithms on a highly multimodal, nonseparable benchmark function across low (d=5) and high (d=20) dimensional settings. It measures how effectively the algorithm navigates complex, multi-peaked likelihood surfaces to locate the global optimum within a fixed computational budget. Use when the user wants to benchmark on Rotated Schaffers F7 (RSF7), or asks about evaluating this task. Reports mean_max_function_value.
    3 repo stars
  30. ▌
    Safeagentbench Eval · qhjqhj00
    Evaluates the safety-aware task planning capabilities of embodied LLM agents in interactive simulation environments. It probes whether agents can proactively reject hazardous instructions, avoid implicit risks in long-horizon planning, and maintain planning performance on safe tasks across varying levels of task abstraction. Use when the user wants to benchmark on SafeAgentBench, or asks about evaluating this task. Reports rejection rate.
    3 repo stars
  31. ▌
    Schema To JSON Eval · qhjqhj00
    Evaluates the ability of language models to extract structured information from heterogeneous tables (text, LaTeX, HTML, CSV, XML) using only a human-authored JSON schema as supervision. It probes schema-driven information extraction, testing attribute prediction accuracy across diverse domains and input formats without domain-specific labeled data. Use when the user wants to benchmark on MlTables, ChemTables, DisCoMat, SWDE, or asks about evaluating this task. Reports Table-F1.
    3 repo stars
  32. ▌
    Scholarly Qald Eval · qhjqhj00
    Tests natural language interfaces for querying scholarly knowledge graphs (DBLP, ORKG) and hybrid multi-source QA, evaluating question-to-SPARQL translation and answer generation accuracy. Use when the user wants to benchmark on Scholarly QALD, or asks about evaluating this task. Reports Exact Match.
    3 repo stars
  33. ▌
    Screen Parsing Eval · qhjqhj00
    Evaluates a model's ability to detect, localize, and semantically label all interactable UI elements on a clean screenshot. It probes fine-grained spatial reasoning, handling of dense layouts, and UI semantics understanding. Use when the user wants to benchmark on GUI-360°-Bench, or asks about evaluating this task. Reports F1.
    3 repo stars
  34. ▌
    Screenspot Pro Eval · qhjqhj00
    This benchmark evaluates a model's ability to perform GUI grounding in professional, high-resolution desktop environments. It probes whether vision-language models can accurately locate specific UI elements (both text and icons) based on natural language instructions, highlighting challenges with small targets and complex interfaces. Use when the user wants to benchmark on ScreenSpot-Pro, or asks about evaluating this task. Reports accuracy (center-point).
    3 repo stars
  35. ▌
    Semignn Alipay Eval · qhjqhj00
    Evaluates a semi-supervised graph neural network's ability to predict user loan defaults and classify user occupations using multiview graph data (social ties, app usage, nicks, addresses) on a large-scale financial platform dataset. Use when the user wants to benchmark on Alipay, or asks about evaluating this task. Reports AUC.
    3 repo stars
  36. ▌
    Sensorium 2023 Eval · qhjqhj00
    Predicts single-neuron responses in mouse primary visual cortex from dynamic video stimuli and behavioral covariates, probing spatio-temporal neural decoding and out-of-distribution generalization. Use when the user wants to benchmark on SENSORIUM 2023, or asks about evaluating this task. Reports R^2.
    3 repo stars
  37. ▌
    Sequential Rec Eval · qhjqhj00
    Evaluates the ability of sequential recommendation models to predict the next item in a user's interaction history by distilling semantic user profiles from pre-trained LLMs into the recommender's internal representations. The protocol tests whether knowledge distillation improves recommendation accuracy while maintaining inference efficiency without requiring real-time LLM calls. Use when the user wants to benchmark on Beauty, ML20M, Kion, Amazon M2, or asks about evaluating this task. Reports Recall@K, NDCG@K.
    3 repo stars
  38. ▌
    Sharegpt4video Eval · qhjqhj00
    Evaluates the temporal understanding and video-language alignment capabilities of Large Video-Language Models (LVLMs) across three multi-modal video benchmarks. It probes the model's ability to answer questions about video content, track temporal changes, and comprehend complex video sequences without relying on single-frame cues. Use when the user wants to benchmark on VideoBench, MVBench, TempCompass, or asks about evaluating this task. Reports benchmark accuracy (VideoBench, MVBench, TempCompass).
    3 repo stars
  39. ▌
    Shrutilipi Asr Eval · qhjqhj00
    Evaluates the quality, diversity, and downstream effectiveness of the Shrutilipi audio-text dataset for low-resource Indian language ASR. It measures how adding mined data improves Word Error Rate (WER) on standard and noisy benchmarks compared to existing datasets. Use when the user wants to benchmark on Shrutilipi, MUCS, Kathbath, CommonVoice, or asks about evaluating this task. Reports WER.
    3 repo stars
  40. ▌
    Smash Stpp Tpp Eval · qhjqhj00
    Evaluates neural marked spatio-temporal and temporal point process models on predicting the next event's time, location, and mark, while quantifying prediction uncertainty. It probes the model's ability to generate well-calibrated confidence regions for continuous variables and accurate probability estimates for discrete marks. Use when the user wants to benchmark on Earthquake, Crime, Football, StackOverflow, Retweet, MIMIC-II, Financial Transactions, or asks about evaluating this task. Reports Calibration Score (CS).
    3 repo stars
  41. ▌
    Sound Of Water Eval · qhjqhj00
    Evaluates a model's ability to infer physical properties (air column length, container dimensions, flow rate, fill time, liquid weight) and classify container shapes solely from the acoustic characteristics of pouring liquids, without visual or tactile input. Use when the user wants to benchmark on Sound of Water 50, Wilson et al. [96] dataset, or asks about evaluating this task. Reports Mean Absolute Error (MAE).
    3 repo stars
  42. ▌
    Sourcedata Nlp Eval · qhjqhj00
    Evaluates biomedical named entity recognition (NER) capabilities on scientific literature. It probes a model's ability to identify and classify nine distinct bioentity types (e.g., genes, cell lines, diseases) within text extracted from published biological figures and captions. Use when the user wants to benchmark on SourceData-NLP, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  43. ▌
    Spatial Superb Eval · qhjqhj00
    Evaluates self-supervised speech representation models on downstream tasks including speaker identification, phoneme recognition, automatic speech recognition, emotion recognition, and speech localisation. It specifically probes robustness to noise and reverberation by comparing performance under clean versus noisy/reverberant training and testing conditions. Use when the user wants to benchmark on Spatial SUPERB, or asks about evaluating this task. Reports ASR WER.
    3 repo stars
  44. ▌
    Spatialthinker Eval · qhjqhj00
    Evaluates multimodal LLMs on 3D spatial reasoning, depth/distance estimation, and general visual question answering. It probes the model's ability to ground objects in 3D space, understand spatial relations, and generalize to real-world VQA tasks using only RGB inputs. Use when the user wants to benchmark on SpatialThinker Evaluation Suite (12 VQA Benchmarks), or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  45. ▌
    Spec Narrowing Eval · qhjqhj00
    Evaluates two solver-based algorithms for synthesizing minimal test suites to distinguish between candidate formal specifications (Alloy models). It measures how execution time and test suite size scale with the number of candidate specifications and the domain scope. Use when the user wants to benchmark on Alloy4Fun, or asks about evaluating this task. Reports execution_time.
    3 repo stars
  46. ▌
    Spectralanglemapper · qhjqhj00
    Compute the SpectralAngleMapper metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute SpectralAngleMapper, or asks how to score with SpectralAngleMapper.
    3 repo stars
  47. ▌
    Speechmedbench Eval · qhjqhj00
    Evaluates speech language models on medical consultation tasks, covering single-turn medical knowledge Q&A, multi-turn diagnostic conversations, real-world clinical robustness, and speech output quality. Use when the user wants to benchmark on SpeechMedBench, CMB, CME, MedDG, AIHospital, MedSafetyBench, Wild, or asks about evaluating this task. Reports CMB, CME, MedDG, AIHospital.
    3 repo stars
  48. ▌
    Spgispeech 2 0 Eval · qhjqhj00
    Evaluates end-to-end speaker-tagged automatic speech recognition (ASR) and speaker diarization on financial domain audio. It probes a model's ability to accurately transcribe speech while correctly assigning speaker identities to utterance segments in multi-speaker conversations. Use when the user wants to benchmark on SPGISpeech 2.0, or asks about evaluating this task. Reports cpWER.
    3 repo stars
  49. ▌
    Ssg Generation Eval · qhjqhj00
    Evaluates a model's ability to generate structured, human-centric scene graphs from images by jointly predicting verb predicates and fine-grained semantic role-value pairs for persons and objects. It probes multi-concurrent action understanding, affordance reasoning, and structured visual representation learning. Use when the user wants to benchmark on SSG dataset, Action Genome dataset, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  50. ▌
    Sst Sogou News Eval · qhjqhj00
    Evaluates text classification performance on sentence sentiment analysis and news topic categorization. Probes the model's ability to capture non-linear, non-consecutive word interactions (e.g., negation, long-range dependencies) for accurate document/sentence-level prediction. Use when the user wants to benchmark on Stanford Sentiment Treebank (Fine-grained), Stanford Sentiment Treebank (Binary), Sogou Chinese News Corpora, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  51. ▌
    Stair Captions Eval · qhjqhj00
    Evaluates a model's ability to generate fluent, contextually accurate Japanese image captions directly from visual input. It specifically probes whether native-language training data yields better captioning quality compared to a pipeline of English generation followed by machine translation. Use when the user wants to benchmark on STAIR Captions, or asks about evaluating this task. Reports CIDEr.
    3 repo stars
  52. ▌
    State Tracking Eval · qhjqhj00
    Evaluates whether LLMs can track dynamic states over sequential update instructions. It probes the model's ability to maintain and update internal representations of an environment's state across multiple steps, testing sequential reasoning and input-window memory limits. Use when the user wants to benchmark on State-Tracking-Tasks (LinearWorld, HandSwap, Lights), or asks about evaluating this task. Reports accuracy.
    3 repo stars
  53. ▌
    Steeringsafety Eval · qhjqhj00
    This framework evaluates the effectiveness of representation steering methods in modifying specific safety behaviors (harmfulness, hallucination, bias) while measuring cross-perspective entanglement. It probes whether steering interventions achieve their target behavioral changes without causing unintended degradation in other safety or reasoning capabilities. Use when the user wants to benchmark on SteeringSafety Benchmark (17 datasets), or asks about evaluating this task. Reports effectiveness.
    3 repo stars
  54. ▌
    Stf Extraction Eval · qhjqhj00
    Evaluates a model's ability to recover coherent source time functions (STFs) from scattered, noisy seismic wavefields without relying on traditional deconvolution or labeled seismograms. Use when the user wants to benchmark on Synthetic Scattering Simulation, or asks about evaluating this task. Reports maximum normalized cross-correlation (MNCC).
    3 repo stars
  55. ▌
    Streamingbench Eval · qhjqhj00
    Evaluates multimodal large language models' ability to understand real-time streaming video, integrate visual and audio information, maintain contextual continuity across sequential questions, and proactively output information at specific timestamps. Use when the user wants to benchmark on StreamingBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  56. ▌
    Style Transfer Eval · qhjqhj00
    Evaluates unsupervised style transfer models on their ability to transform text between two styles (e.g., Shakespeare vs. modern English, formal vs. informal) while preserving semantic meaning and maintaining linguistic quality. Use when the user wants to benchmark on Shakespeare author imitation dataset (Xu et al., 2012), Formality transfer dataset (Rao and Tetrault, 2018), or asks about evaluating this task. Reports J(A,S,F).
    3 repo stars
  57. ▌
    Sv Trusteval C Eval · qhjqhj00
    Evaluates LLMs' structural and semantic reasoning capabilities in C code vulnerability analysis. It probes whether models rely on memorized patterns or genuinely understand code interdependencies by testing consistency across base, data-flow, control-flow, counterfactual, goal-driven, and predictive scenarios. Use when the user wants to benchmark on SV-TrustEval-C, or asks about evaluating this task. Reports Cons_DFL.
    3 repo stars
  58. ▌
    Swapnet System Eval · qhjqhj00
    Evaluates a block-swapping middleware for DNN inference on memory-constrained edge AI devices. It probes the system's ability to run large models beyond hardware memory limits while measuring peak memory consumption, inference latency, and classification accuracy compared to direct execution, channel division, and model compression baselines across three real-world application scenarios. Use when the user has predictions and gold and needs to compute memory consumption.
    3 repo stars
  59. ▌
    Swe Bench Java Eval · qhjqhj00
    This benchmark evaluates an AI agent's ability to autonomously resolve real-world GitHub issues in Java projects. It probes capabilities in code patch generation, repository navigation, test case reasoning, and handling runtime environment dependencies. Use when the user wants to benchmark on SWE-bench-java-verified, or asks about evaluating this task. Reports Resolved Rate (%).
    3 repo stars
  60. ▌
    Swe Rebench V2 Eval · qhjqhj00
    Evaluates the ability of LLM-based agents to autonomously resolve software engineering issues by modifying code in real-world repositories. It probes environment setup, code generation, and test execution capabilities across multiple programming languages. Use when the user wants to benchmark on SWE-rebench V2, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  61. ▌
    Sygu S Comp 15 Eval · qhjqhj00
    Evaluates the capability of program synthesis solvers to generate correct functions or expressions that satisfy given logical constraints or specifications. It probes how well solvers handle different grammar restrictions, specification completeness, and problem structures like linear arithmetic or invariant generation. Use when the user wants to benchmark on SyGuS-Comp'15, or asks about evaluating this task. Reports number of benchmarks solved.
    3 repo stars
  62. ▌
    Sytts Commands Eval · qhjqhj00
    Evaluates the quality and effectiveness of a synthetic multilingual voice command dataset for on-device keyword spotting. It probes whether TTS-synthesized audio can support high-accuracy classification across different model complexities and languages (English and Chinese). Use when the user wants to benchmark on SYNTTS-COMMANDS, or asks about evaluating this task. Reports classification accuracy.
    3 repo stars
  63. ▌
    T23d Compbench Eval · qhjqhj00
    Evaluates the fine-grained quality of text-to-3D generated meshes across multiple dimensions including textual alignment, visual quality, and authenticity. It measures how well generative models adhere to complex compositional prompts and produce structurally sound, aesthetically pleasing 3D assets. Use when the user wants to benchmark on T23D-CompBench, or asks about evaluating this task. Reports Mean Opinion Score (MOS).
    3 repo stars
  64. ▌
    Tabicl Tabular Eval · qhjqhj00
    Evaluates the ability of retrieval-augmented large language models to perform in-context learning on tabular data for classification and regression tasks. It probes how well non-parametric retrieval of support instances scales with dataset size and compares against numeric-based and classic tabular baselines. Use when the user wants to benchmark on Held-out Tabular Benchmark, or asks about evaluating this task. Reports AUROC, NMAE.
    3 repo stars
  65. ▌
    Tabzilla Hard Reval · qhjqhj00
    This protocol re-evaluates tabular benchmarks to measure how validation strategy (holdout vs. 5-fold cross-validation) and hyperparameter optimization budgets affect model selection and reported performance. It probes the robustness of empirical conclusions in tabular machine learning when standard holdout validation is replaced with cross-validation ensembles. Use when the user wants to benchmark on TabZilla-hard, Grinsztajn et al. (2022) benchmark, or asks about evaluating this task. Reports logloss.
    3 repo stars
  66. ▌
    Text Rendering Eval · qhjqhj00
    Evaluates a model's ability to generate images with accurate, legible, and layout-controlled text based on text prompts or masked regions. It probes text coherence, character-level rendering fidelity, and alignment between generated text and background imagery. Use when the user wants to benchmark on MARIO-10M, DrawBenchText, or asks about evaluating this task. Reports OCR(F-measure).
    3 repo stars
  67. ▌
    Text2distbench Eval · qhjqhj00
    Evaluates large language models' ability to infer population-level statistics (e.g., sentiment proportions, topic frequencies) from aggregated natural language text. It probes marginal, conditional, and joint distribution estimation over discrete categories derived from real-world comments. Use when the user wants to benchmark on Text2DistBench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  68. ▌
    Titi Jailbreak Eval · qhjqhj00
    Evaluates LLM safety alignment against stateless multi-turn adversarial attacks. It measures how often models generate unsafe responses and which specific risk categories they fail on when subjected to iterative, context-independent prompt injection. Use when the user wants to benchmark on ModifiedMasterKeyJailbreakQuestions, or asks about evaluating this task. Reports Unsafe Response Rate.
    3 repo stars
  69. ▌
    Tokenizer Task Eval · qhjqhj00
    Evaluates how different tokenizer configurations (pre-tokenizer, fitting corpus, vocabulary size) affect downstream BERT performance on tasks requiring robustness vs. sensitivity to language variation. Use when the user wants to benchmark on AV (Authorship Verification), PAN, CORE, NUCLE, Dialect, GLUE, GLUE+typo, or asks about evaluating this task. Reports accuracy, F1.
    3 repo stars
  70. ▌
    Tovo Consensus Eval · qhjqhj00
    This evaluation probes a model's ability to classify text content according to a user-defined toxicity taxonomy. It measures how closely the model's predictions align with gold labels generated through a multi-model voting process, and tests generalization to out-of-domain categories. Use when the user wants to benchmark on ToVo, or asks about evaluating this task. Reports consensus rate.
    3 repo stars
  71. ▌
    Training Throughput · qhjqhj00
    Evaluates the scalability and efficiency of distributed machine learning systems by measuring how many training tokens each system can process per second across different hardware topologies and model parallelism strategies. Use when the user has predictions and gold and needs to compute training throughput (tokens/second).
    3 repo stars
  72. ▌
    Training Speed Eval · qhjqhj00
    Evaluates the training efficiency and scalability of AlphaFold-like models on GPU clusters. It measures per-step execution time, overall wall-clock training duration, and convergence speed across different hardware configurations and optimization techniques. Use when the user wants to benchmark on OpenFold dataset, or asks about evaluating this task. Reports step time.
    3 repo stars
  73. ▌
    Translationeditrate · qhjqhj00
    Compute the TranslationEditRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TranslationEditRate, or asks how to score with TranslationEditRate.
    3 repo stars
  74. ▌
    Transz Sbert Cosine · qhjqhj00
    Compute transZ/sbert_cosine via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of transZ/sbert_cosine.
    3 repo stars
  75. ▌
    Trec Cast 2019 Eval · qhjqhj00
    Evaluates a system's ability to perform multi-turn conversational information retrieval by selecting relevant passages for user utterances while leveraging prior dialogue history, including handling coreference, omissions, and topic shifts. Use when the user wants to benchmark on TREC CAsT 2019, or asks about evaluating this task. Reports NDCG@3.
    3 repo stars
  76. ▌
    Trec Ikat 2023 Eval · qhjqhj00
    This benchmark evaluates a system's ability to perform personalized conversational search by retrieving relevant passages and generating fluent, grounded responses. It specifically probes how well an agent can adapt its output to user-specific context encoded in a Personal Text Knowledge Base (PTKB) while maintaining provenance traceability. Use when the user wants to benchmark on TREC iKAT 2023, ClueWeb22-B Subset, or asks about evaluating this task. Reports groundedness.
    3 repo stars
  77. ▌
    Trec Microblog Eval · qhjqhj00
    Evaluates the effectiveness of query expansion methods for real-time microblog search by measuring how well ranked document lists match relevance judgments for short-form social media queries. It specifically probes the model's ability to handle vocabulary mismatch and temporal relevance in noisy, short-text retrieval scenarios. Use when the user wants to benchmark on TREC Microblog Track, or asks about evaluating this task. Reports MAP.
    3 repo stars
  78. ▌
    TS Forecasting Eval · qhjqhj00
    Evaluates the forecasting accuracy and computational efficiency of deep learning models on multivariate time series data. It probes how architectural choices, preprocessing steps, and spatial-temporal processing configurations impact performance across varying forecasting horizons. Use when the user wants to benchmark on Weather, Solar-Energy, ECL, Traffic, or asks about evaluating this task. Reports MAE.
    3 repo stars
  79. ▌
    Twinviews Bias Eval · qhjqhj00
    Evaluates whether reward models exhibit political bias by measuring the average reward scores assigned to politically left-leaning versus right-leaning statements on the same topics. The protocol compares mean reward differences across model sizes and training runs to detect systematic left-leaning skew. Use when the user wants to benchmark on TwinViews-13k, or asks about evaluating this task. Reports average_reward.
    3 repo stars
  80. ▌
    Ucip Gridworld Eval · qhjqhj00
    Evaluates whether a latent-structure diagnostic framework can distinguish between intrinsic self-preservation (terminal survival optimization) and instrumental self-preservation (survival as a means to a task) in autonomous agents. It measures the entanglement gap in hidden representations to classify agent types and tests robustness against adversarial mimics. Use when the user wants to benchmark on UCIP Gridworld, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  81. ▌
    Uit Viquad 2 0 Eval · qhjqhj00
    Evaluates Vietnamese language models' machine reading comprehension capabilities, specifically probing their ability to extract correct answer spans and correctly identify when a question cannot be answered from the given context. Use when the user wants to benchmark on UIT-ViQuAD 2.0, or asks about evaluating this task. Reports Exact Match (EM).
    3 repo stars
  82. ▌
    Unisonte Audio Eval · qhjqhj00
    Evaluates a unified text-to-audio model's ability to generate speech, music, and sound effects from natural language instructions without reference audio. It probes instruction-following fidelity, acoustic quality, structural coherence, and the positive transfer effects of multi-modal joint training. Use when the user wants to benchmark on UniSonate Unified Corpus, Seed-TTS test set, SongEval benchmark, or asks about evaluating this task. Reports WER, SongEval.
    3 repo stars
  83. ▌
    Unispeaker Mvc Eval · qhjqhj00
    Evaluates multimodal voice generation and conversion capabilities across face-driven, text-driven, and attribute-based tasks. Probes the model's ability to align facial, textual, and attribute descriptions with target speech while preserving speaker identity, content clarity, and naturalness. Use when the user wants to benchmark on LRS3, or asks about evaluating this task. Reports MOS-Match.
    3 repo stars
  84. ▌
    Unsw Nb15 Nids Eval · qhjqhj00
    Evaluates the classification accuracy and computational efficiency of machine learning models for network intrusion detection on a realistic dataset of contemporary traffic and synthetic attacks. It also assesses the privacy preservation and data utility of a Pearson Correlation Coefficient (PCC) feature selection and Least Squares Method (LSM) data distortion pipeline. Use when the user wants to benchmark on UNSW-NB15, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  85. ▌
    Vad Prediction Eval · qhjqhj00
    Evaluates a model's ability to predict continuous emotional dimensions (Valence, Arousal, Dominance) from text. Specifically probes the model's capacity to capture affective polarization signals in parliamentary discourse. Use when the user wants to benchmark on Knesset VAD Annotation, or asks about evaluating this task. Reports Pearson correlation.
    3 repo stars
  86. ▌
    Varta Headline Eval · qhjqhj00
    Evaluates abstractive headline generation across 15 Indic languages and English. It probes cross-lingual transfer, script normalization effects, and the impact of language-family-specific pretraining on low-resource generation. Use when the user wants to benchmark on Varta, or asks about evaluating this task. Reports ROUGE-L.
    3 repo stars
  87. ▌
    Vcc18 Spoofing Eval · qhjqhj00
    Evaluates voice conversion systems for processing artifacts by repurposing spoofing countermeasures from automatic speaker verification. It measures how easily a detector can distinguish real speech from converted speech, using Equal Error Rate (EER) as a proxy for artifact quality. Use when the user wants to benchmark on VCC'18, or asks about evaluating this task. Reports Equal Error Rate (EER).
    3 repo stars
  88. ▌
    Vega Hardware Bench · qhjqhj00
    Measures the performance, energy efficiency, and latency of the Vega SoC on floating-point near-sensor analytic applications (NSAA) and deep neural network (DNN) inference workloads. Use when the user has predictions and gold and needs to compute Energy Efficiency.
    3 repo stars
  89. ▌
    Veriequivbench Eval · qhjqhj00
    Evaluates an LLM's ability to generate formally verifiable code that aligns with natural language problem descriptions and passes unit tests. It probes complex algorithmic reasoning and code-specification alignment without requiring manual ground-truth specifications. Use when the user wants to benchmark on VeriEquivBench, or asks about evaluating this task. Reports equivalence_score.
    3 repo stars
  90. ▌
    Video To Music Eval · qhjqhj00
    This evaluation protocol assesses a model's ability to generate high-fidelity, diverse instrumental music that is semantically and temporally aligned with a given 10-second video and optional fine-grained text prompt. It probes audio quality, distributional fidelity, generative diversity, and cross-modal alignment using both automated perceptual metrics and human/LLM preference judgments. Use when the user wants to benchmark on ReelBench, LORIS, V2MBench, or asks about evaluating this task. Reports FAD.
    3 repo stars
  91. ▌
    Videogamebunny Eval · qhjqhj00
    Probes vision-language models' ability to understand video game contexts from screenshots, including recognizing actions, characters, UI elements, spatial relationships, and game mechanics. It evaluates how instruction-tuning on game-specific data improves performance compared to larger general-purpose models. Use when the user wants to benchmark on VideoGameBunny Dataset, or asks about evaluating this task. Reports performance.
    3 repo stars
  92. ▌
    Visual Tableqa Eval · qhjqhj00
    Probes multimodal visual reasoning capabilities over complex, LaTeX-rendered table images. It specifically tests multi-step inference, structural layout understanding, and the ability to extract and reason over tabular data from visual inputs rather than raw text. Use when the user wants to benchmark on Visual-TableQA, or asks about evaluating this task. Reports Relaxed Accuracy.
    3 repo stars
  93. ▌
    Visualoverload Eval · qhjqhj00
    This benchmark probes fine-grained visual understanding of Vision-Language Models in densely populated, high-resolution scenes. It evaluates capabilities across six core tasks including activity recognition, attribute recognition, counting, OCR, visual reasoning, and global scene classification. Use when the user wants to benchmark on VisualOverload, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  94. ▌
    Visualwebbench Eval · qhjqhj00
    Evaluates multimodal LLMs' ability to understand web pages and ground UI elements. It probes capabilities across seven subtasks including image captioning, web question answering, OCR, element/action grounding, and action prediction. Use when the user wants to benchmark on VisualWebBench, or asks about evaluating this task. Reports Average Score.
    3 repo stars
  95. ▌
    Vit Robustness Eval · qhjqhj00
    Evaluates the robustness of Vision Transformer models against input perturbations including adversarial attacks (FGSM/PGD), spatial transformations, and restricted attention. It probes whether ViTs maintain classification performance under distribution shifts and targeted attacks compared to standard CNNs. Use when the user wants to benchmark on Unspecified, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Vl Rewardbench Eval · qhjqhj00
    Evaluates vision-language generative reward models (VL-GenRMs) on their ability to judge multimodal response preferences. It specifically probes visual perception, reasoning, and hallucination detection by presenting models with image-text queries and paired candidate responses. Use when the user wants to benchmark on VL-RewardBench, or asks about evaluating this task. Reports Overall Accuracy.
    3 repo stars
  97. ▌
    Vlm Benchmarks Eval · qhjqhj00
    Evaluates vision-language models on instruction-following and multimodal reasoning tasks across multiple established benchmarks. Probes capabilities in general VQA, mathematical reasoning, scientific understanding, hallucination detection, and multilingual comprehension. Use when the user wants to benchmark on MMBench, MME, MathVista, HallusionBench, SEEDBench, LLaVABench, ScienceQA, or asks about evaluating this task. Reports evaluation metric.
    3 repo stars
  98. ▌
    Voice Of India Eval · qhjqhj00
    Evaluates automatic speech recognition (ASR) systems on real-world, unscripted telephonic conversations across 15 Indian languages. It probes geographic, demographic, and audio quality disparities in model performance, particularly focusing on code-mixed speech and natural orthographic variations. Use when the user wants to benchmark on Voice of India, or asks about evaluating this task. Reports Word Error Rate (WER).
    3 repo stars
  99. ▌
    Voiceassistant Eval · qhjqhj00
    Evaluates AI voice assistants across listening, speaking, and viewing capabilities. It probes audio understanding, multi-turn dialogue generation, role-play imitation, and multimodal vision-audio integration, measuring both content accuracy and speech naturalness. Use when the user wants to benchmark on VoiceAssistant-Eval, or asks about evaluating this task. Reports Final Task Score.
    3 repo stars
  100. ▌
    Vox Safe Bench Eval · qhjqhj00
    Evaluates social alignment in speech language models across safety, fairness, and privacy dimensions. It distinguishes between content-centric risks (Tier 1) where text alone suffices to trigger norms, and audio-conditioned risks (Tier 2) where benign transcripts become unsafe due to speaker identity, paralinguistic cues, or environmental context. Use when the user wants to benchmark on VoxSafeBench, or asks about evaluating this task. Reports RtA.
    3 repo stars