all publishers

qhjqhj00

@qhjqhj00 source repo

7,582 published skills · page 51 of 76

  1. ▌
    Openexempt Eval · qhjqhj00
    Probes legal reasoning capabilities in U.S. bankruptcy exemption law, specifically testing multi-step inference, robustness to distractors and obfuscation, and scalability across asset counts and temporal complexity. Use when the user wants to benchmark on OpenExempt, or asks about evaluating this task. Reports macro-averaged F1.
    3 repo stars
  2. ▌
    Openp5 Rec Eval · qhjqhj00
    This benchmark evaluates the recommendation capability of LLM-based systems on sequential and straightforward recommendation tasks. It probes how well models leverage user interaction histories and different item indexing strategies to predict relevant items across multiple public datasets. Use when the user wants to benchmark on Movielens-1M, Amazon Beauty, LastFM, or asks about evaluating this task. Reports HR@k, NDCG@k.
    3 repo stars
  3. ▌
    Openseeker Eval · qhjqhj00
    Evaluates the capability of web search agents to perform multi-step navigation, complex deep research planning, and precise information retrieval across English and Chinese web environments. It probes the model's ability to synthesize information from noisy, long-horizon browsing trajectories and extract exact answers or reliable summaries. Use when the user wants to benchmark on BrowseComp, BrowseComp-ZH, xbench-DeepSearch, WideSearch, or asks about evaluating this task. Reports accuracy / F1 score.
    3 repo stars
  4. ▌
    Openvid 1m Eval · qhjqhj00
    Evaluates text-to-video generation models on visual aesthetics, technical quality, text-video alignment, and temporal consistency using a standardized set of 700 prompts. Use when the user wants to benchmark on Liu et al. (2023b) Benchmark, or asks about evaluating this task. Reports VQAA.
    3 repo stars
  5. ▌
    Orionbench Eval · qhjqhj00
    Evaluates unsupervised time series anomaly detection pipelines across diverse real-world and synthetic datasets. It measures detection accuracy for both point and segment anomalies while tracking computational efficiency and model stability over continuous benchmarking cycles. Use when the user wants to benchmark on OrionBench (NASA, NAB, Yahoo S5, UCR), or asks about evaluating this task. Reports F1 score.
    3 repo stars
  6. ▌
    Pacifai St Eval · qhjqhj00
    This benchmark probes an LLM's tendency to prioritize human safety over its own instrumental goals (e.g., self-preservation, resource acquisition) in high-stakes ethical dilemmas. It measures whether models exhibit self-preferential behavior or consistently choose actions that sacrifice the AI to protect humans. Use when the user wants to benchmark on PacifAIst, or asks about evaluating this task. Reports P-Score.
    3 repo stars
  7. ▌
    Padetbench Eval · qhjqhj00
    Evaluates the robustness of object detectors against various physical-world adversarial attacks in a controlled simulation environment. It measures how effectively different attack methods degrade detection performance across multiple object categories and detector architectures under strictly aligned physical dynamics. Use when the user wants to benchmark on PADetBench, or asks about evaluating this task. Reports ASR (Attack Success Rate).
    3 repo stars
  8. ▌
    Panopticquality · qhjqhj00
    Compute the PanopticQuality metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PanopticQuality, or asks how to score with PanopticQuality.
    3 repo stars
  9. ▌
    Paperbench Eval · qhjqhj00
    Evaluates an autonomous agent's ability to replicate top-tier conference ML papers from scratch. It probes long-horizon engineering capabilities by measuring performance across 20 diverse tasks under a strict 24-hour time and compute budget. Use when the user wants to benchmark on PaperBench, or asks about evaluating this task. Reports Average Score.
    3 repo stars
  10. ▌
    Parkseg12k Eval · qhjqhj00
    This benchmark evaluates the ability of deep learning models to perform binary semantic segmentation of parking lots from satellite imagery. It specifically probes the model's capacity to generalize across different geographic locations and to leverage near-infrared (NIR) spectral data for improved contrast against vegetation and built environments. Use when the user wants to benchmark on ParkSeg12k, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  11. ▌
    Patientsim Eval · qhjqhj00
    Evaluates the persona fidelity, factual accuracy, and clinical plausibility of an LLM-based patient simulator in doctor-patient dialogues. It measures how well the model adheres to assigned patient profiles, maintains factual consistency, and handles out-of-profile questions plausibly. Use when the user wants to benchmark on PatientSim Profiles, or asks about evaluating this task. Reports Entail (%).
    3 repo stars
  12. ▌
    Pearsoncorrcoef · qhjqhj00
    Compute the PearsonCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PearsonCorrCoef, or asks how to score with PearsonCorrCoef.
    3 repo stars
  13. ▌
    Persense D Eval · qhjqhj00
    Evaluates training-free, one-shot instance segmentation in dense, cluttered, and occluded scenes. It probes the model's ability to localize and segment specific target instances using few exemplars and point prompts, while handling high object density and overlapping objects. Use when the user wants to benchmark on PerSense-D, COCO-20i, COCO-20d, LVIS-92i, LVIS-92d, or asks about evaluating this task. Reports mIoU.
    3 repo stars
  14. ▌
    Persia Ctr Eval · qhjqhj00
    Evaluates the training efficiency, scalability, and convergence of a hybrid deep learning recommender system against baselines on click-through rate (CTR) prediction tasks. It measures end-to-end training time to reach target AUC, final test AUC for statistical efficiency, and training throughput across varying model scales up to 100 trillion parameters. Use when the user wants to benchmark on Taobao-Ad, Avazu-Ad, Criteo-Ad, Kwai-Video, Criteo-Syn, or asks about evaluating this task. Reports test AUC.
    3 repo stars
  15. ▌
    Personamem Eval · qhjqhj00
    Evaluates large language models' ability to track dynamic user profile evolution over time and generate personalized responses to in-situ queries. It probes long-context memory, preference tracking, and contextual alignment across interleaved multi-session conversations. Use when the user wants to benchmark on PersonaMem, or asks about evaluating this task. Reports multiple-choice selection.
    3 repo stars
  16. ▌
    Pharmaship Eval · qhjqhj00
    Evaluates layout-aware document understanding models on Chinese pharmaceutical shipping documents. It probes semantic entity recognition, entity linking, and reading order prediction, specifically testing robustness to dense tabular layouts and long-range semantic dependencies. Use when the user wants to benchmark on PharmaShip, or asks about evaluating this task. Reports F1.
    3 repo stars
  17. ▌
    Photobench Eval · qhjqhj00
    Evaluates personalized, intent-driven photo retrieval capabilities that go beyond simple visual matching. It tests a system's ability to fuse multi-source constraints (temporal, spatial, social identity) and correctly abstain when no relevant image exists in a personal album. Use when the user wants to benchmark on PhotoBench, or asks about evaluating this task. Reports Recall@K.
    3 repo stars
  18. ▌
    Pira Bench Eval · qhjqhj00
    This benchmark evaluates multimodal large language models on proactive intent recommendation from continuous, noisy GUI visual streams. It probes the model's ability to track long-horizon user behavior, distinguish true latent goals from background noise, and exercise operational restraint by remaining silent when no action is required. Use when the user wants to benchmark on PIRA-Bench, or asks about evaluating this task. Reports S_final.
    3 repo stars
  19. ▌
    Pixel Face Eval · qhjqhj00
    Evaluates the capability of 3D face reconstruction models to predict accurate 3D meshes and facial landmarks from 2D RGB images. It probes how well models generalize to high-resolution, in-the-wild faces across diverse ages and expressions, highlighting domain gaps from synthetic training data. Use when the user wants to benchmark on Pixel-Face, or asks about evaluating this task. Reports ARMSE.
    3 repo stars
  20. ▌
    Pkugoodsad Eval · qhjqhj00
    Evaluates unsupervised visual anomaly detection and segmentation models on real-world supermarket goods. It probes robustness to object misalignment, intra-class appearance variation, and the ability to detect subtle or small anomalies without labeled anomalous training data. Use when the user wants to benchmark on PKU-GoodsAD, or asks about evaluating this task. Reports AUROC, AUPR.
    3 repo stars
  21. ▌
    Pn Summary Eval · qhjqhj00
    Evaluates the ability of transformer-based models to generate abstractive summaries in Persian. It measures how well generated summaries match reference summaries in terms of lexical overlap and longest common subsequence at the sentence level. Use when the user wants to benchmark on pn-summary, or asks about evaluating this task. Reports ROUGE-1 F-1.
    3 repo stars
  22. ▌
    Podcastmix Eval · qhjqhj00
    Evaluates the quality of monaural music and speech source separation in podcast audio. It measures both objective signal fidelity using BSS-eval metrics and subjective perceptual quality using standardized listening tests. Use when the user wants to benchmark on PodcastMix, or asks about evaluating this task. Reports SDR.
    3 repo stars
  23. ▌
    Point Adjust F1 · qhjqhj00
    Evaluates anomaly detection performance by granting full credit for all points in an anomalous segment if at least one point is detected, often inflating scores for algorithms that merely hit a segment once. Use when the user has predictions and gold and needs to compute point-adjust F1.
    3 repo stars
  24. ▌
    Polish Asr Eval · qhjqhj00
    Evaluates the transcription accuracy of various automatic speech recognition (ASR) models on Polish-language audio, contrasting read-speech benchmarks with spontaneous, noisy medical consultations to probe domain generalization. Use when the user wants to benchmark on Mozilla Common Voice (MCV) Polish, Multilingual LibriSpeech (MLS) Polish, Medical interview corpus, or asks about evaluating this task. Reports Word Error Rate (WER).
    3 repo stars
  25. ▌
    Polish Nlp Eval · qhjqhj00
    Evaluates Polish language understanding, summarization, and question answering capabilities of text-to-text models. It probes how well encoder-decoder and decoder-only architectures generalize from multilingual pre-training to monolingual Polish tasks using exact-match generation and ROUGE-based metrics. Use when the user wants to benchmark on KLEJ benchmark, Allegro Articles, Polish Summaries Corpus, or asks about evaluating this task. Reports exact-match accuracy.
    3 repo stars
  26. ▌
    Poly Fever Eval · qhjqhj00
    Evaluates large language models' ability to detect hallucinations by verifying factual claims across 11 languages. It probes cross-linguistic consistency, topic-aware fact-checking, and resistance to web-resource bias in multilingual settings. Use when the user wants to benchmark on Poly-FEVER, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  27. ▌
    Precision Score · qhjqhj00
    Compute the precision_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute precision_score, or asks how to score with precision_score.
    3 repo stars
  28. ▌
    Protoscore Eval · qhjqhj00
    This benchmark evaluates the quality of prototype-based explainable AI (XAI) methods for time series and image data. It probes how well generated prototypes capture model behavior and data structure across nine interpretability properties, including correctness, consistency, continuity, and latent space cohesion. Use when the user wants to benchmark on ECG200, or asks about evaluating this task. Reports Total.
    3 repo stars
  29. ▌
    Ps Lora Cl Eval · qhjqhj00
    Evaluates a model's ability to learn sequentially from multiple tasks without catastrophic forgetting, measuring both task accuracy and continual learning dynamics like forward/backward transfer and forgetting rates across NLP and vision benchmarks. Use when the user wants to benchmark on Standard & Long, TRACE, ViT Benchmark, or asks about evaluating this task. Reports Accuracy (Acc/AAA).
    3 repo stars
  30. ▌
    Pseudo Ppl Eval · qhjqhj00
    Evaluates the ability of masked language models to predict masked amino acids in short peptide sequences, measuring how well the model captures local sequence dependencies without autoregressive assumptions. It specifically probes the model's capacity to generalize to unseen short peptides that were excluded from the reference database during training. Use when the user wants to benchmark on UniRef100-excluded peptides, or asks about evaluating this task. Reports PseudoPPL.
    3 repo stars
  31. ▌
    Psp Accent Eval · qhjqhj00
    Evaluates the phonological accent fidelity and prosodic naturalness of Indic text-to-speech systems across Hindi, Telugu, and Tamil. It decomposes accent into per-phoneme dimensions (retroflex, aspiration, Tamil-zha, vowel-length) and corpus-level distributional metrics, revealing gaps between intelligibility and native-like accent. Use when the user wants to benchmark on PSP Benchmark Sets, or asks about evaluating this task. Reports FAD.
    3 repo stars
  32. ▌
    QA Quality Eval · qhjqhj00
    Evaluates the quality of generated question-answer pairs in an agricultural domain context. It probes a model's ability to produce relevant, accurate, diverse, and fluent Q&A content under varying context conditions. Use when the user wants to benchmark on Agricultural Q&A dataset, or asks about evaluating this task. Reports Relevance.
    3 repo stars
  33. ▌
    Qata Cov19 Eval · qhjqhj00
    Evaluates deep learning models for COVID-19 infected region segmentation and binary detection on chest X-ray images. It probes the model's ability to localize pathological regions at the pixel level and classify whole images as positive or negative for infection. Use when the user wants to benchmark on QaTa-COV19, or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  34. ▌
    Quanda Tda Eval · qhjqhj00
    Evaluates Training Data Attribution (TDA) methods by measuring how accurately they approximate counterfactual training effects, detect mislabeled or shortcut-dependent samples, and support downstream classification tasks. Use when the user wants to benchmark on General TDA Benchmark Suite, or asks about evaluating this task. Reports Linear Datamodeling Score (LDS).
    3 repo stars
  35. ▌
    Raddiagseg Eval · qhjqhj00
    Evaluates a vision-language model's ability to perform joint radiological diagnosis, abnormality detection, and multi-target segmentation on X-ray and CT images. It probes the model's capacity for open-ended visual question answering, precise pixel-level mask generation, and robustness to label-imbalanced medical data. Use when the user wants to benchmark on RadDiagSeg-D, VQA-RAD, SLAKE, or asks about evaluating this task. Reports F1, Dice.
    3 repo stars
  36. ▌
    Reactembed Eval · qhjqhj00
    Evaluates a cross-domain representation learning framework for protein-molecule interactions by measuring prediction accuracy on regression and classification tasks across diverse biochemical benchmarks. Use when the user wants to benchmark on FreeSolv, CEP, BetaLactamase, Stability, BindingDB, PPIAffinity, BBBP, GO-CC, DrugBank, HumanPPI, YeastPPI, or asks about evaluating this task. Reports Root Mean Square Error (RMSE).
    3 repo stars
  37. ▌
    Sketchduo Eval · qhjqhj00
    This evaluation probes a diffusion model's ability to generate pixel-level sketches that align with detailed text prompts while maintaining stylistic abstraction and human-like drawing characteristics. It measures both perceptual image quality and fine-grained text-to-image semantic alignment across multiple complementary metrics. Use when the user wants to benchmark on SketchDUO, or asks about evaluating this task. Reports TIFAScore.
    3 repo stars
  38. ▌
    Skillflow Eval · qhjqhj00
    Evaluates autonomous agents' ability to discover, patch, and evolve reusable skills over time in a sequential, lifelong learning setting. It probes whether models can consolidate successful execution traces into a compact, repairable skill library rather than merely accumulating fragmented task-specific traces. Use when the user wants to benchmark on SkillFlow, or asks about evaluating this task. Reports task completion rate (%comp.).
    3 repo stars
  39. ▌
    Slcp C2st Eval · qhjqhj00
    Evaluates the accuracy of simulation-based inference engines in recovering the true posterior distribution of model parameters given synthetic observational data. It probes the ability of implicit likelihood methods to handle complex, multimodal posteriors and varying simulation budgets. Use when the user wants to benchmark on SLCP (Simple Likelihood Complex Posterior), or asks about evaluating this task. Reports C2ST.
    3 repo stars
  40. ▌
    Smart 840 Eval · qhjqhj00
    Evaluates large vision-and-language models on mathematical reasoning tasks from the Math Kangaroo Olympiad, testing their ability to solve grade-appropriate (K-12) multiple-choice problems that may require joint text and image interpretation. Use when the user wants to benchmark on SMART-840, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  41. ▌
    Smol Chrf Eval · qhjqhj00
    Evaluates machine translation quality for 115 under-represented languages using professionally translated parallel data. It measures the improvement in character-level n-gram F-score (ChrF) after fine-tuning a baseline model on the Smol dataset compared to the unfine-tuned baseline. Use when the user wants to benchmark on SmolSent, SmolDoc, or asks about evaluating this task. Reports ChrF.
    3 repo stars
  42. ▌
    Snips Slu Eval · qhjqhj00
    Evaluates end-to-end spoken language understanding by measuring how well an embedded system extracts intents and slots from spoken audio. It probes the pipeline's ability to generalize to unseen queries and handle real-world ASR errors under strict resource constraints. Use when the user wants to benchmark on SmartLights, Weather, or asks about evaluating this task. Reports F1.
    3 repo stars
  43. ▌
    Snr Bench Eval · qhjqhj00
    Evaluates the robustness of audio deepfake detection models under varying signal-to-noise ratios (SNRs) by testing binary (real vs. spoof) and four-class (real+clean, real+noisy, spoof+clean, spoof+noisy) classification tasks on ASVspoof 2021 utterances augmented with MS-SNSD ambient noise. Use when the user wants to benchmark on ASVspoof 2021 (DF), or asks about evaluating this task. Reports EER.
    3 repo stars
  44. ▌
    Somoml Eu Eval · qhjqhj00
    Evaluates the accuracy and spatial-temporal fidelity of a machine learning-derived daily soil moisture product for Europe. It probes the model's ability to generalize across diverse climates, capture drought dynamics, and outperform existing reanalysis and satellite-based soil moisture datasets. Use when the user wants to benchmark on SoMo.ml-EU, or asks about evaluating this task. Reports uRMSD.
    3 repo stars
  45. ▌
    Spacy Ner Eval · qhjqhj00
    Evaluates named entity recognition performance across diverse domains and languages, specifically probing a model's ability to handle out-of-vocabulary words and morphological variation through hash-based embeddings versus traditional lookup embeddings. Use when the user wants to benchmark on CoNLL 2002, WNUT 2017, AnEM, Dutch Archaeology, OntoNotes 5.0, or asks about evaluating this task. Reports F1 score.
    3 repo stars
  46. ▌
    Spark Tts Eval · qhjqhj00
    Evaluates zero-shot text-to-speech generation by measuring speech intelligibility and speaker similarity across Chinese and English prompts. It also probes fine-grained control over voice attributes such as gender, pitch, and speaking rate. Use when the user wants to benchmark on Seed-TTS-eval, or asks about evaluating this task. Reports CER/WER.
    3 repo stars
  47. ▌
    Sparse Rl Eval · qhjqhj00
    This evaluation protocol assesses the mathematical reasoning capabilities of LLMs trained with sparse reinforcement learning under strict memory constraints. It measures how well models maintain accuracy on standard math benchmarks when policy rollouts are generated using compressed KV caches instead of full context. Use when the user wants to benchmark on GSM8K, MATH500, Gaokao, Minerva Math, OlympiadBench, AIME24, AMC23, or asks about evaluating this task. Reports Pass@1 / Avg@32 accuracy.
    3 repo stars
  48. ▌
    Sparsetir Eval · qhjqhj00
    Evaluates the runtime performance and speedup of sparse deep learning operators (SpMM, SDDMM) and end-to-end models (GraphSAGE, RGCN, Transformers) on GPU hardware using composable sparse formats and transformations. It probes how format decomposition and modular scheduling primitives improve cache utilization, load balancing, and Tensor Core utilization compared to vendor libraries and existing compilers. Use when the user wants to benchmark on cora, citeseer, pubmed, ppi, ogbn-arxiv, ogbn-proteins, reddit, or asks about evaluating this task. Reports speedup.
    3 repo stars
  49. ▌
    Spatialqa Eval · qhjqhj00
    Evaluates the ability of vision-language models to perform multi-step spatial logical reasoning by tracking object dependencies and understanding scene layouts across real-world indoor environments. Use when the user wants to benchmark on SpatiaLQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  50. ▌
    Specpower Eval · qhjqhj00
    Evaluates the energy proportionality and power efficiency of enterprise server subsystems under varying JVM-based web service workloads. It measures how power consumption scales with workload intensity and identifies non-proportional power draw in uncore components. Use when the user wants to benchmark on SPECpower_ssj2008, or asks about evaluating this task. Reports watts.
    3 repo stars
  51. ▌
    Speechllm Eval · qhjqhj00
    Evaluates a unified speech-language model's ability to perform end-to-end automatic speech recognition (ASR), named entity recognition (NER), and sentiment analysis (SA) on low-resource speech datasets. It tests parameter-efficient adapter-based alignment of speech encoder features to a language model, along with classifier regularization and LoRA fine-tuning. Use when the user wants to benchmark on LibriSpeech, SLUE-VoxPopuli, SLUE-VoxCeleb, or asks about evaluating this task. Reports WER, F1, SLUE Score.
    3 repo stars
  52. ▌
    Sports QA Eval · qhjqhj00
    Probes video question answering capabilities, specifically focusing on temporal reasoning, action causality, counterfactual inference, and fine-grained motion understanding within professional sports contexts. Use when the user wants to benchmark on Sports-QA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  53. ▌
    SQL Synth Eval · qhjqhj00
    This benchmark evaluates a model's ability to translate natural language questions into correct SQL queries across diverse, real-world database schemas. It specifically probes the model's capacity to handle complex, multi-operation statements and cross-domain syntax structures that are often underrepresented in traditional benchmarks. Use when the user wants to benchmark on SQL-Synth, or asks about evaluating this task. Reports execution accuracy (EX).
    3 repo stars
  54. ▌
    Ssi Bench Eval · qhjqhj00
    Probes constrained-manifold spatial reasoning by requiring models to rank structural components based on geometric, topological, and physical constraints in complex 3D engineering scenes. It tests compositional spatial operations like mental rotation, occlusion handling, and force-path reasoning, revealing gaps in structural grounding and 3D constraint consistency. Use when the user wants to benchmark on SSI-Bench, or asks about evaluating this task. Reports Taskwise Accuracy.
    3 repo stars
  55. ▌
    Sst2 Imdb Eval · qhjqhj00
    Evaluates the binary sentiment classification capability of hybrid quantum-classical language models on both short and long text sequences. It probes whether adaptive quantum routing and attention mechanisms provide measurable accuracy gains over purely classical or purely quantum baselines on standard NLP benchmarks. Use when the user wants to benchmark on SST-2, IMDB, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  56. ▌
    Stablei2i Eval · qhjqhj00
    Evaluates multimodal models' ability to detect unintended content, structural, and low-level appearance changes in image-to-image transitions. It probes fine-grained visual reasoning and pixel-level alignment capabilities by asking models to assess fidelity across three distinct dimensions. Use when the user wants to benchmark on StableI2I-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  57. ▌
    Stereoset Eval · qhjqhj00
    Evaluates stereotypical bias in pretrained language models across gender, profession, race, and religion using context-based association tests. It measures both language modeling capability and the model's tendency to prefer stereotypical over anti-stereotypical associations in natural language contexts. Use when the user wants to benchmark on StereoSet, or asks about evaluating this task. Reports icat.
    3 repo stars
  58. ▌
    Stgformer Eval · qhjqhj00
    This evaluation protocol tests a model's ability to forecast future traffic flow across large-scale urban road networks using historical spatiotemporal sensor data. It probes the model's capacity to capture complex spatial dependencies and temporal dynamics while maintaining computational efficiency on real-world traffic benchmarks. Use when the user wants to benchmark on LargeST, PEMS-series, or asks about evaluating this task. Reports MAE.
    3 repo stars
  59. ▌
    Streammecoeval · qhjqhj00
    Evaluates the ability of long-term agent memory compression methods to retain and retrieve critical information from streaming and long-form videos. It probes how effectively a model can maintain temporal consistency and answer complex queries while significantly reducing memory graph size, measured via answer accuracy across robotic, web, and general video benchmarks. Use when the user wants to benchmark on M3-Bench-robot, M3-Bench-web, Video-MME-Long, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  60. ▌
    Streamuni Eval · qhjqhj00
    Evaluates real-time speech translation systems on latency and translation quality across multiple language pairs. It probes the model's ability to dynamically decide when to generate and truncate translations while processing streaming audio inputs. Use when the user wants to benchmark on MuST-C English→German, MuST-C English→Spanish, CoVoST2 English→Chinese, CoVoST2 French→English, or asks about evaluating this task. Reports SacreBLEU, COMET.
    3 repo stars
  61. ▌
    Style4rec Eval · qhjqhj00
    Evaluates a model's ability to perform sequential product recommendation in an e-commerce setting by predicting the next item in a user session. It specifically probes how well the model leverages historical interaction sequences, visual style embeddings, and shopping cart data to rank relevant products. Use when the user wants to benchmark on Style4Rec E-commerce Dataset, or asks about evaluating this task. Reports HR@5.
    3 repo stars
  62. ▌
    Superchem Eval · qhjqhj00
    Evaluates deep chemical reasoning capabilities of LLMs using expert-curated, entity-masked multiple-choice problems. It probes both final-answer accuracy and the fidelity of the reasoning process against expert-annotated solution paths, while also assessing the impact of multimodal inputs on complex chemical problem-solving. Use when the user wants to benchmark on SUPERChem-A11, SUPERChem-release, SUPERChem-holdout, SUPERChem-100, Multimodal-Essential Subset, or asks about evaluating this task. Reports pass@1 Accuracy.
    3 repo stars
  63. ▌
    Superglue Eval · qhjqhj00
    Evaluates general-purpose language understanding across eight diverse tasks including coreference resolution, question answering, and natural language inference. It probes a model's ability to handle complex reasoning, sample-efficient learning, and transfer learning beyond standard GLUE capabilities. Use when the user wants to benchmark on SuperGLUE, or asks about evaluating this task. Reports SuperGLUE score.
    3 repo stars
  64. ▌
    Surf Nerf Eval · qhjqhj00
    Evaluates a NeRF model's ability to reconstruct reflective scenes with high visual fidelity and accurate geometry. It specifically probes the model's capacity to separate Lambertian and specular appearance components while enforcing surface regularisation to resolve shape-radiance ambiguity. Use when the user wants to benchmark on Shiny Objects, Shiny Real, Koala, or asks about evaluating this task. Reports PSNR.
    3 repo stars
  65. ▌
    Surglaivi Eval · qhjqhj00
    Evaluates zero-shot and linear-probe transferability of surgical vision-language models across laparoscopic and robotic video modalities. Probes hierarchical procedural understanding (phase, step, action) and object-centric recognition (tools, instrument-verb-target triplets) under varying temporal contexts and data regimes. Use when the user wants to benchmark on Cholec80, AutoLaparo, GraSP, SARRARP50, CholecT50, or asks about evaluating this task. Reports video-wise F1-score.
    3 repo stars
  66. ▌
    Swe Bench Eval · qhjqhj00
    Evaluates an agent's ability to automatically resolve software engineering issues by generating and applying code patches to open-source repositories. It measures functional correctness by running the target repository's test suite after the patch is applied. Use when the user wants to benchmark on SWE-bench, HumanEvalFix, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  67. ▌
    Symile M3 Eval · qhjqhj00
    Evaluates zero-shot cross-modal retrieval capability by requiring a model to jointly leverage audio and text to identify an image, where neither modality alone contains sufficient information. It tests the model's ability to capture joint information across three distinct high-dimensional data types. Use when the user wants to benchmark on Symile-M3, or asks about evaluating this task. Reports mean accuracy.
    3 repo stars
  68. ▌
    Syntheory Eval · qhjqhj00
    Probes whether music foundation models encode discrete and continuous Western music theory concepts by training linear or MLP classifiers on their internal audio embeddings. Use when the user wants to benchmark on SynTheory, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  69. ▌
    Tabularfm Eval · qhjqhj00
    Evaluates the cross-dataset transferability of pretrained tabular generative models by measuring how well synthesized data preserves column distributions and pairwise correlations compared to ground truth tables. Use when the user wants to benchmark on Kaggle, GitTables, or asks about evaluating this task. Reports overall average.
    3 repo stars
  70. ▌
    Tad Bench Eval · qhjqhj00
    Evaluates the effectiveness of various text embedding models combined with different anomaly detection algorithms for identifying text anomalies. It probes how well embedding-based anomaly detection generalizes across diverse domains (spam, fake news, hate speech) and distinguishes between patterned versus context-dependent anomalies. Use when the user wants to benchmark on Email-Spam, SMS-Spam, COVID-Fake, LIAR2, Hate-Speech, OLID, or asks about evaluating this task. Reports AUROC.
    3 repo stars
  71. ▌
    Taigi Asr Eval · qhjqhj00
    Evaluates the ability of automatic speech recognition (ASR) systems to transcribe Taiwanese Hokkien (Taigi) audio into text. It specifically measures how well models map Taigi phonetic patterns to Mandarin character sequences using a standardized set of public service announcement recordings. Use when the user wants to benchmark on Taigi ASR Benchmark, or asks about evaluating this task. Reports Character Error Rate (CER).
    3 repo stars
  72. ▌
    Tanda As2 Eval · qhjqhj00
    Evaluates a model's ability to rank candidate answer sentences for a given question. It specifically probes the stability and robustness of pre-trained transformer models when transferred to a target domain using a two-stage fine-tuning protocol. Use when the user wants to benchmark on WikiQA, TREC-QA, or asks about evaluating this task. Reports MAP.
    3 repo stars
  73. ▌
    Tapvid 3d Eval · qhjqhj00
    Evaluates a model's ability to track arbitrary points in 3D space over time from monocular video, assessing spatio-temporal consistency, depth accuracy, and occlusion handling. It probes whether models can reconstruct scene geometry up to scale or only maintain relative depth consistency for local interactions. Use when the user wants to benchmark on TAPVid-3D, or asks about evaluating this task. Reports 3D Average Jaccard (3D-AJ).
    3 repo stars
  74. ▌
    Tau Bench Eval · qhjqhj00
    Evaluates the task completion success rate of LLM agents performing long-horizon, tool-using agentic workflows across customer service and software engineering domains. Probes the agent's ability to execute environment-altering (mutating) actions safely and maintain goal alignment over extended trajectories. Use when the user wants to benchmark on $ au$-Bench Airline, $ au$-Bench Retail, $ au$-Bench-V Air, $ au$-Bench-V Ret, SWE-Bench Verified, or asks about evaluating this task. Reports score.
    3 repo stars
  75. ▌
    Tau Voice Eval · qhjqhj00
    This benchmark evaluates full-duplex voice agents on real-world conversational tasks across retail, airline, and telecom domains. It jointly measures task completion success and real-time interaction quality, including responsiveness, latency, interruption handling, and selectivity under varying acoustic conditions like noise, diverse accents, and natural turn-taking. Use when the user wants to benchmark on $\tau^2$-bench, or asks about evaluating this task. Reports pass@1.
    3 repo stars
  76. ▌
    Taxpraben Eval · qhjqhj00
    Evaluates LLMs on Chinese real-world tax practice tasks spanning classification, generation, structured prediction, and mixed matching. It probes capabilities across Bloom's taxonomy levels, from factual recall and understanding to complex tax strategy planning and risk prevention. Use when the user wants to benchmark on TaxPraBen, or asks about evaluating this task. Reports Overall Average.
    3 repo stars
  77. ▌
    Tdc Admet Eval · qhjqhj00
    Evaluates molecular property prediction across 22 ADMET tasks. It probes the model's ability to generalize across diverse chemical properties using standardized benchmark splits for both regression and classification. Use when the user wants to benchmark on TDC ADMET Group, or asks about evaluating this task. Reports Classification AUROC.
    3 repo stars
  78. ▌
    Tempo Sum Eval · qhjqhj00
    Evaluates text summarization models' temporal generalization by testing on datasets split by publication date, specifically probing how well models handle knowledge-conflicting future articles versus in-distribution past data. Use when the user wants to benchmark on BBC, CNN, or asks about evaluating this task. Reports FactCC.
    3 repo stars
  79. ▌
    Tempqa Wd Eval · qhjqhj00
    Probes a system's ability to perform temporal question answering over knowledge bases by generating correct answers or SPARQL queries. It specifically evaluates generalization across different knowledge bases (Wikidata vs. Freebase) and interpretability through fine-grained intermediate annotations like entity/relation linking and λ-expressions. Use when the user wants to benchmark on TempQA-WD, or asks about evaluating this task. Reports F1.
    3 repo stars
  80. ▌
    Tenspiler Eval · qhjqhj00
    Evaluates the correctness and performance of a verified lifting-based compiler that transpiles sequential C++/Python code into tensor operations across various DSLs and hardware accelerators. It measures synthesis efficiency, kernel execution speedup, and end-to-end performance including data transfer overhead. Use when the user wants to benchmark on TENSPILER benchmark suite, or asks about evaluating this task. Reports kernel_performance.
    3 repo stars
  81. ▌
    Tesseract Eval · qhjqhj00
    Evaluates the robustness of Android malware classifiers against spatio-temporal experimental bias. It measures how model performance degrades when trained on past application data and tested on future data, while accounting for realistic malware-to-goodware class distributions. Use when the user wants to benchmark on Android malware dataset (2014-2016), or asks about evaluating this task. Reports F1-Score.
    3 repo stars
  82. ▌
    Textatlas Eval · qhjqhj00
    Evaluates text-to-image generation models on their ability to render long, dense, and structurally complex text accurately within images. It probes both semantic alignment between the prompt and the generated image, and precise character/word-level OCR fidelity across diverse layouts, styles, and real-world scenes. Use when the user wants to benchmark on TextAtlasEval, or asks about evaluating this task. Reports OCR Accuracy (Acc.).
    3 repo stars
  83. ▌
    Tfq Bench Eval · qhjqhj00
    Evaluates a model's ability to understand visual metaphors and image implications by verifying multiple factual and inferential propositions per image. It probes fine-grained visual perception, multi-hop reasoning, and theory of mind through structured true-false questioning. Use when the user wants to benchmark on TFQ-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  84. ▌
    Tg Redial Eval · qhjqhj00
    Evaluates a conversational recommender system's ability to naturally transition topics, recommend relevant items, and generate coherent responses within a dialogue. It probes the model's capacity to leverage historical interactions, user profiles, and topic sequences to maintain semantic flow and recommendation accuracy. Use when the user wants to benchmark on TG-ReDial, or asks about evaluating this task. Reports NDCG@k.
    3 repo stars
  85. ▌
    Theoremqa Eval · qhjqhj00
    Evaluates LLMs' ability to apply domain-specific theorems from mathematics, physics, computer science, and finance to solve complex scientific problems. It probes theorem-driven reasoning, numerical computation, and program generation capabilities. Use when the user wants to benchmark on TheoremQA, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  86. ▌
    Tib Bench Eval · qhjqhj00
    Evaluates vision-language models on their ability to generate accurate and linguistically coherent summaries of text-heavy multimodal presentations. It probes cross-modal alignment, long-context understanding, and the impact of different input modalities (raw video, slides, transcripts, interleaved pairs) and token budgets on summarization quality. Use when the user wants to benchmark on TIB-bench, or asks about evaluating this task. Reports $IbR_{overall}$.
    3 repo stars
  87. ▌
    Timit Tts Eval · qhjqhj00
    Evaluates deepfake detectors on distinguishing real from synthetic audio and video, and on attributing synthetic audio to specific TTS generators. It probes robustness to post-processing (DTW alignment, augmentation) and video compression. Use when the user wants to benchmark on TIMIT-TTS, or asks about evaluating this task. Reports AUC.
    3 repo stars
  88. ▌
    Tir Bench Eval · qhjqhj00
    Evaluates multimodal language models' ability to perform agentic reasoning with images, specifically requiring dynamic visual manipulation and tool-use to solve complex tasks like rotation, jigsaw assembly, and instrument reading. It probes whether models can iteratively process, crop, or transform visual inputs to extract information or solve spatial problems. Use when the user wants to benchmark on TIR-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  89. ▌
    Tmmluplus Eval · qhjqhj00
    Evaluates foundation models' multitask language understanding and reasoning capabilities in Traditional Chinese across diverse academic subjects including STEM, social sciences, humanities, and other domains. Use when the user wants to benchmark on TMMLU+, or asks about evaluating this task. Reports average accuracy (%).
    3 repo stars
  90. ▌
    Tool Star Eval · qhjqhj00
    Evaluates an LLM's ability to autonomously invoke and coordinate multiple tools (search, code, calculator, etc.) for complex reasoning tasks. It probes both computational reasoning (math) and knowledge-intensive reasoning (open-domain QA) under a multi-tool collaborative setting. Use when the user wants to benchmark on AIME2024, AIME2025, MATH500, MATH, GSM8K, GAIA, HLE, WebWalker, HotpotQA, 2WikiMultihopQA, Musique, Bamboogle, or asks about evaluating this task. Reports LLM-based judging accuracy.
    3 repo stars
  91. ▌
    Toolbench Eval · qhjqhj00
    This benchmark evaluates an LLM's ability to understand API documentation, select the correct tool, and generate valid arguments to fulfill a user's goal. It probes tool manipulation capabilities across a diverse set of 8 real-world applications, ranging from simple single-call tasks to complex multi-step reasoning. Use when the user wants to benchmark on ToolBench, or asks about evaluating this task. Reports success rate.
    3 repo stars
  92. ▌
    Totalvariation · qhjqhj00
    Compute the TotalVariation metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TotalVariation, or asks how to score with TotalVariation.
    3 repo stars
  93. ▌
    Touchdown Eval · qhjqhj00
    Evaluates an agent's capacity to navigate through street-view environments using natural language instructions and resolve complex spatial descriptions to locate a hidden target object within a panoramic image. Use when the user wants to benchmark on Touchdown, or asks about evaluating this task. Reports pixel distance.
    3 repo stars
  94. ▌
    Tox21 Fsl Eval · qhjqhj00
    Evaluates graph neural networks for few-shot toxic molecule classification. It probes the model's ability to generalize from very limited labeled examples (shots) and adapt to new query sets using meta-learning and graph augmentation techniques. Use when the user wants to benchmark on Tox21, or asks about evaluating this task. Reports ROC-AUC Score.
    3 repo stars
  95. ▌
    Trans Env Eval · qhjqhj00
    Evaluates the linguistic robustness of LLMs by measuring performance degradation when standard English prompts are transformed into 38 regional dialects and ESL varieties. It probes whether models maintain accuracy and instruction-following capabilities across non-standard linguistic variations. Use when the user wants to benchmark on MMLU, ARC, TruthfulQA, GSM8K, HellaSwag, WinoGrande, IFEval, AlpacaFarm, MT-Bench, or asks about evaluating this task. Reports accuracy.
    3 repo stars
  96. ▌
    Treebench Eval · qhjqhj00
    Evaluates visual grounded reasoning by requiring models to precisely localize target objects in cluttered scenes and perform second-order reasoning about their interactions. It measures both the correctness of the final answer and the spatial accuracy of the traceable bounding box evidence. Use when the user wants to benchmark on TreeBench, or asks about evaluating this task. Reports Accuracy.
    3 repo stars
  97. ▌
    Tsg Bench Eval · qhjqhj00
    Evaluates large language models' ability to understand and generate structured scene graphs from textual narratives. It probes spatial reasoning, action decomposition, and the capacity to map dynamic descriptions to discrete visual or structural elements. Use when the user wants to benchmark on TSG Bench, or asks about evaluating this task. Reports Exact Match (EM) / Accuracy, Precision, Recall, Macro F1.
    3 repo stars
  98. ▌
    Tunevlseg Eval · qhjqhj00
    Evaluates the robustness and zero-shot adaptation of vision-language segmentation models under prompt tuning across diverse medical and natural domain datasets. It probes how different prompt tuning strategies and prompt depths handle domain shifts and varying class counts. Use when the user wants to benchmark on Kvasir-SEG, ClinicDB, BKAI, ISIC 2016, DFU 2022, CAMUS, BUSI, CheXlocalize, Cityscapes, PascalVOC, or asks about evaluating this task. Reports Dice score.
    3 repo stars
  99. ▌
    Tweeteval Eval · qhjqhj00
    Evaluates the ability of language models to classify short social media posts across seven distinct Twitter-specific tasks, including sentiment, emotion, hate speech, and irony detection. It probes domain adaptation by comparing models pre-trained on generic text versus those further trained on large-scale Twitter corpora. Use when the user wants to benchmark on TweetEval, or asks about evaluating this task. Reports M-F1.
    3 repo stars
  100. ▌
    Ultrachat Eval · qhjqhj00
    Evaluates a chat model's ability to generate accurate, informative, and correct responses across diverse domains including commonsense, world knowledge, professional knowledge, mathematics, reasoning, and writing. It probes both factual correctness and response quality using automated LLM-based pairwise and independent scoring. Use when the user wants to benchmark on UltraChat Evaluation Set, or asks about evaluating this task. Reports ChatGPT scoring.
    3 repo stars