qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Openexempt Eval · qhjqhj00Probes legal reasoning capabilities in U.S. bankruptcy exemption law, specifically testing multi-step inference, robustness to distractors and obfuscation, and scalability across asset counts and temporal complexity. Use when the user wants to benchmark on OpenExempt, or asks about evaluating this task. Reports macro-averaged F1.
- ▌ Openp5 Rec Eval · qhjqhj00This benchmark evaluates the recommendation capability of LLM-based systems on sequential and straightforward recommendation tasks. It probes how well models leverage user interaction histories and different item indexing strategies to predict relevant items across multiple public datasets. Use when the user wants to benchmark on Movielens-1M, Amazon Beauty, LastFM, or asks about evaluating this task. Reports HR@k, NDCG@k.
- ▌ Openseeker Eval · qhjqhj00Evaluates the capability of web search agents to perform multi-step navigation, complex deep research planning, and precise information retrieval across English and Chinese web environments. It probes the model's ability to synthesize information from noisy, long-horizon browsing trajectories and extract exact answers or reliable summaries. Use when the user wants to benchmark on BrowseComp, BrowseComp-ZH, xbench-DeepSearch, WideSearch, or asks about evaluating this task. Reports accuracy / F1 score.
- ▌ Openvid 1m Eval · qhjqhj00Evaluates text-to-video generation models on visual aesthetics, technical quality, text-video alignment, and temporal consistency using a standardized set of 700 prompts. Use when the user wants to benchmark on Liu et al. (2023b) Benchmark, or asks about evaluating this task. Reports VQAA.
- ▌ Orionbench Eval · qhjqhj00Evaluates unsupervised time series anomaly detection pipelines across diverse real-world and synthetic datasets. It measures detection accuracy for both point and segment anomalies while tracking computational efficiency and model stability over continuous benchmarking cycles. Use when the user wants to benchmark on OrionBench (NASA, NAB, Yahoo S5, UCR), or asks about evaluating this task. Reports F1 score.
- ▌ Pacifai St Eval · qhjqhj00This benchmark probes an LLM's tendency to prioritize human safety over its own instrumental goals (e.g., self-preservation, resource acquisition) in high-stakes ethical dilemmas. It measures whether models exhibit self-preferential behavior or consistently choose actions that sacrifice the AI to protect humans. Use when the user wants to benchmark on PacifAIst, or asks about evaluating this task. Reports P-Score.
- ▌ Padetbench Eval · qhjqhj00Evaluates the robustness of object detectors against various physical-world adversarial attacks in a controlled simulation environment. It measures how effectively different attack methods degrade detection performance across multiple object categories and detector architectures under strictly aligned physical dynamics. Use when the user wants to benchmark on PADetBench, or asks about evaluating this task. Reports ASR (Attack Success Rate).
- ▌ Panopticquality · qhjqhj00Compute the PanopticQuality metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PanopticQuality, or asks how to score with PanopticQuality.
- ▌ Paperbench Eval · qhjqhj00Evaluates an autonomous agent's ability to replicate top-tier conference ML papers from scratch. It probes long-horizon engineering capabilities by measuring performance across 20 diverse tasks under a strict 24-hour time and compute budget. Use when the user wants to benchmark on PaperBench, or asks about evaluating this task. Reports Average Score.
- ▌ Parkseg12k Eval · qhjqhj00This benchmark evaluates the ability of deep learning models to perform binary semantic segmentation of parking lots from satellite imagery. It specifically probes the model's capacity to generalize across different geographic locations and to leverage near-infrared (NIR) spectral data for improved contrast against vegetation and built environments. Use when the user wants to benchmark on ParkSeg12k, or asks about evaluating this task. Reports mIoU.
- ▌ Patientsim Eval · qhjqhj00Evaluates the persona fidelity, factual accuracy, and clinical plausibility of an LLM-based patient simulator in doctor-patient dialogues. It measures how well the model adheres to assigned patient profiles, maintains factual consistency, and handles out-of-profile questions plausibly. Use when the user wants to benchmark on PatientSim Profiles, or asks about evaluating this task. Reports Entail (%).
- ▌ Pearsoncorrcoef · qhjqhj00Compute the PearsonCorrCoef metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute PearsonCorrCoef, or asks how to score with PearsonCorrCoef.
- ▌ Persense D Eval · qhjqhj00Evaluates training-free, one-shot instance segmentation in dense, cluttered, and occluded scenes. It probes the model's ability to localize and segment specific target instances using few exemplars and point prompts, while handling high object density and overlapping objects. Use when the user wants to benchmark on PerSense-D, COCO-20i, COCO-20d, LVIS-92i, LVIS-92d, or asks about evaluating this task. Reports mIoU.
- ▌ Persia Ctr Eval · qhjqhj00Evaluates the training efficiency, scalability, and convergence of a hybrid deep learning recommender system against baselines on click-through rate (CTR) prediction tasks. It measures end-to-end training time to reach target AUC, final test AUC for statistical efficiency, and training throughput across varying model scales up to 100 trillion parameters. Use when the user wants to benchmark on Taobao-Ad, Avazu-Ad, Criteo-Ad, Kwai-Video, Criteo-Syn, or asks about evaluating this task. Reports test AUC.
- ▌ Personamem Eval · qhjqhj00Evaluates large language models' ability to track dynamic user profile evolution over time and generate personalized responses to in-situ queries. It probes long-context memory, preference tracking, and contextual alignment across interleaved multi-session conversations. Use when the user wants to benchmark on PersonaMem, or asks about evaluating this task. Reports multiple-choice selection.
- ▌ Pharmaship Eval · qhjqhj00Evaluates layout-aware document understanding models on Chinese pharmaceutical shipping documents. It probes semantic entity recognition, entity linking, and reading order prediction, specifically testing robustness to dense tabular layouts and long-range semantic dependencies. Use when the user wants to benchmark on PharmaShip, or asks about evaluating this task. Reports F1.
- ▌ Photobench Eval · qhjqhj00Evaluates personalized, intent-driven photo retrieval capabilities that go beyond simple visual matching. It tests a system's ability to fuse multi-source constraints (temporal, spatial, social identity) and correctly abstain when no relevant image exists in a personal album. Use when the user wants to benchmark on PhotoBench, or asks about evaluating this task. Reports Recall@K.
- ▌ Pira Bench Eval · qhjqhj00This benchmark evaluates multimodal large language models on proactive intent recommendation from continuous, noisy GUI visual streams. It probes the model's ability to track long-horizon user behavior, distinguish true latent goals from background noise, and exercise operational restraint by remaining silent when no action is required. Use when the user wants to benchmark on PIRA-Bench, or asks about evaluating this task. Reports S_final.
- ▌ Pixel Face Eval · qhjqhj00Evaluates the capability of 3D face reconstruction models to predict accurate 3D meshes and facial landmarks from 2D RGB images. It probes how well models generalize to high-resolution, in-the-wild faces across diverse ages and expressions, highlighting domain gaps from synthetic training data. Use when the user wants to benchmark on Pixel-Face, or asks about evaluating this task. Reports ARMSE.
- ▌ Pkugoodsad Eval · qhjqhj00Evaluates unsupervised visual anomaly detection and segmentation models on real-world supermarket goods. It probes robustness to object misalignment, intra-class appearance variation, and the ability to detect subtle or small anomalies without labeled anomalous training data. Use when the user wants to benchmark on PKU-GoodsAD, or asks about evaluating this task. Reports AUROC, AUPR.
- ▌ Pn Summary Eval · qhjqhj00Evaluates the ability of transformer-based models to generate abstractive summaries in Persian. It measures how well generated summaries match reference summaries in terms of lexical overlap and longest common subsequence at the sentence level. Use when the user wants to benchmark on pn-summary, or asks about evaluating this task. Reports ROUGE-1 F-1.
- ▌ Podcastmix Eval · qhjqhj00Evaluates the quality of monaural music and speech source separation in podcast audio. It measures both objective signal fidelity using BSS-eval metrics and subjective perceptual quality using standardized listening tests. Use when the user wants to benchmark on PodcastMix, or asks about evaluating this task. Reports SDR.
- ▌ Point Adjust F1 · qhjqhj00Evaluates anomaly detection performance by granting full credit for all points in an anomalous segment if at least one point is detected, often inflating scores for algorithms that merely hit a segment once. Use when the user has predictions and gold and needs to compute point-adjust F1.
- ▌ Polish Asr Eval · qhjqhj00Evaluates the transcription accuracy of various automatic speech recognition (ASR) models on Polish-language audio, contrasting read-speech benchmarks with spontaneous, noisy medical consultations to probe domain generalization. Use when the user wants to benchmark on Mozilla Common Voice (MCV) Polish, Multilingual LibriSpeech (MLS) Polish, Medical interview corpus, or asks about evaluating this task. Reports Word Error Rate (WER).
- ▌ Polish Nlp Eval · qhjqhj00Evaluates Polish language understanding, summarization, and question answering capabilities of text-to-text models. It probes how well encoder-decoder and decoder-only architectures generalize from multilingual pre-training to monolingual Polish tasks using exact-match generation and ROUGE-based metrics. Use when the user wants to benchmark on KLEJ benchmark, Allegro Articles, Polish Summaries Corpus, or asks about evaluating this task. Reports exact-match accuracy.
- ▌ Poly Fever Eval · qhjqhj00Evaluates large language models' ability to detect hallucinations by verifying factual claims across 11 languages. It probes cross-linguistic consistency, topic-aware fact-checking, and resistance to web-resource bias in multilingual settings. Use when the user wants to benchmark on Poly-FEVER, or asks about evaluating this task. Reports accuracy.
- ▌ Precision Score · qhjqhj00Compute the precision_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute precision_score, or asks how to score with precision_score.
- ▌ Protoscore Eval · qhjqhj00This benchmark evaluates the quality of prototype-based explainable AI (XAI) methods for time series and image data. It probes how well generated prototypes capture model behavior and data structure across nine interpretability properties, including correctness, consistency, continuity, and latent space cohesion. Use when the user wants to benchmark on ECG200, or asks about evaluating this task. Reports Total.
- ▌ Ps Lora Cl Eval · qhjqhj00Evaluates a model's ability to learn sequentially from multiple tasks without catastrophic forgetting, measuring both task accuracy and continual learning dynamics like forward/backward transfer and forgetting rates across NLP and vision benchmarks. Use when the user wants to benchmark on Standard & Long, TRACE, ViT Benchmark, or asks about evaluating this task. Reports Accuracy (Acc/AAA).
- ▌ Pseudo Ppl Eval · qhjqhj00Evaluates the ability of masked language models to predict masked amino acids in short peptide sequences, measuring how well the model captures local sequence dependencies without autoregressive assumptions. It specifically probes the model's capacity to generalize to unseen short peptides that were excluded from the reference database during training. Use when the user wants to benchmark on UniRef100-excluded peptides, or asks about evaluating this task. Reports PseudoPPL.
- ▌ Psp Accent Eval · qhjqhj00Evaluates the phonological accent fidelity and prosodic naturalness of Indic text-to-speech systems across Hindi, Telugu, and Tamil. It decomposes accent into per-phoneme dimensions (retroflex, aspiration, Tamil-zha, vowel-length) and corpus-level distributional metrics, revealing gaps between intelligibility and native-like accent. Use when the user wants to benchmark on PSP Benchmark Sets, or asks about evaluating this task. Reports FAD.
- ▌ QA Quality Eval · qhjqhj00Evaluates the quality of generated question-answer pairs in an agricultural domain context. It probes a model's ability to produce relevant, accurate, diverse, and fluent Q&A content under varying context conditions. Use when the user wants to benchmark on Agricultural Q&A dataset, or asks about evaluating this task. Reports Relevance.
- ▌ Qata Cov19 Eval · qhjqhj00Evaluates deep learning models for COVID-19 infected region segmentation and binary detection on chest X-ray images. It probes the model's ability to localize pathological regions at the pixel level and classify whole images as positive or negative for infection. Use when the user wants to benchmark on QaTa-COV19, or asks about evaluating this task. Reports F1-Score.
- ▌ Quanda Tda Eval · qhjqhj00Evaluates Training Data Attribution (TDA) methods by measuring how accurately they approximate counterfactual training effects, detect mislabeled or shortcut-dependent samples, and support downstream classification tasks. Use when the user wants to benchmark on General TDA Benchmark Suite, or asks about evaluating this task. Reports Linear Datamodeling Score (LDS).
- ▌ Raddiagseg Eval · qhjqhj00Evaluates a vision-language model's ability to perform joint radiological diagnosis, abnormality detection, and multi-target segmentation on X-ray and CT images. It probes the model's capacity for open-ended visual question answering, precise pixel-level mask generation, and robustness to label-imbalanced medical data. Use when the user wants to benchmark on RadDiagSeg-D, VQA-RAD, SLAKE, or asks about evaluating this task. Reports F1, Dice.
- ▌ Reactembed Eval · qhjqhj00Evaluates a cross-domain representation learning framework for protein-molecule interactions by measuring prediction accuracy on regression and classification tasks across diverse biochemical benchmarks. Use when the user wants to benchmark on FreeSolv, CEP, BetaLactamase, Stability, BindingDB, PPIAffinity, BBBP, GO-CC, DrugBank, HumanPPI, YeastPPI, or asks about evaluating this task. Reports Root Mean Square Error (RMSE).
- ▌ Sketchduo Eval · qhjqhj00This evaluation probes a diffusion model's ability to generate pixel-level sketches that align with detailed text prompts while maintaining stylistic abstraction and human-like drawing characteristics. It measures both perceptual image quality and fine-grained text-to-image semantic alignment across multiple complementary metrics. Use when the user wants to benchmark on SketchDUO, or asks about evaluating this task. Reports TIFAScore.
- ▌ Skillflow Eval · qhjqhj00Evaluates autonomous agents' ability to discover, patch, and evolve reusable skills over time in a sequential, lifelong learning setting. It probes whether models can consolidate successful execution traces into a compact, repairable skill library rather than merely accumulating fragmented task-specific traces. Use when the user wants to benchmark on SkillFlow, or asks about evaluating this task. Reports task completion rate (%comp.).
- ▌ Slcp C2st Eval · qhjqhj00Evaluates the accuracy of simulation-based inference engines in recovering the true posterior distribution of model parameters given synthetic observational data. It probes the ability of implicit likelihood methods to handle complex, multimodal posteriors and varying simulation budgets. Use when the user wants to benchmark on SLCP (Simple Likelihood Complex Posterior), or asks about evaluating this task. Reports C2ST.
- ▌ Smart 840 Eval · qhjqhj00Evaluates large vision-and-language models on mathematical reasoning tasks from the Math Kangaroo Olympiad, testing their ability to solve grade-appropriate (K-12) multiple-choice problems that may require joint text and image interpretation. Use when the user wants to benchmark on SMART-840, or asks about evaluating this task. Reports accuracy.
- ▌ Smol Chrf Eval · qhjqhj00Evaluates machine translation quality for 115 under-represented languages using professionally translated parallel data. It measures the improvement in character-level n-gram F-score (ChrF) after fine-tuning a baseline model on the Smol dataset compared to the unfine-tuned baseline. Use when the user wants to benchmark on SmolSent, SmolDoc, or asks about evaluating this task. Reports ChrF.
- ▌ Snips Slu Eval · qhjqhj00Evaluates end-to-end spoken language understanding by measuring how well an embedded system extracts intents and slots from spoken audio. It probes the pipeline's ability to generalize to unseen queries and handle real-world ASR errors under strict resource constraints. Use when the user wants to benchmark on SmartLights, Weather, or asks about evaluating this task. Reports F1.
- ▌ Snr Bench Eval · qhjqhj00Evaluates the robustness of audio deepfake detection models under varying signal-to-noise ratios (SNRs) by testing binary (real vs. spoof) and four-class (real+clean, real+noisy, spoof+clean, spoof+noisy) classification tasks on ASVspoof 2021 utterances augmented with MS-SNSD ambient noise. Use when the user wants to benchmark on ASVspoof 2021 (DF), or asks about evaluating this task. Reports EER.
- ▌ Somoml Eu Eval · qhjqhj00Evaluates the accuracy and spatial-temporal fidelity of a machine learning-derived daily soil moisture product for Europe. It probes the model's ability to generalize across diverse climates, capture drought dynamics, and outperform existing reanalysis and satellite-based soil moisture datasets. Use when the user wants to benchmark on SoMo.ml-EU, or asks about evaluating this task. Reports uRMSD.
- ▌ Spacy Ner Eval · qhjqhj00Evaluates named entity recognition performance across diverse domains and languages, specifically probing a model's ability to handle out-of-vocabulary words and morphological variation through hash-based embeddings versus traditional lookup embeddings. Use when the user wants to benchmark on CoNLL 2002, WNUT 2017, AnEM, Dutch Archaeology, OntoNotes 5.0, or asks about evaluating this task. Reports F1 score.
- ▌ Spark Tts Eval · qhjqhj00Evaluates zero-shot text-to-speech generation by measuring speech intelligibility and speaker similarity across Chinese and English prompts. It also probes fine-grained control over voice attributes such as gender, pitch, and speaking rate. Use when the user wants to benchmark on Seed-TTS-eval, or asks about evaluating this task. Reports CER/WER.
- ▌ Sparse Rl Eval · qhjqhj00This evaluation protocol assesses the mathematical reasoning capabilities of LLMs trained with sparse reinforcement learning under strict memory constraints. It measures how well models maintain accuracy on standard math benchmarks when policy rollouts are generated using compressed KV caches instead of full context. Use when the user wants to benchmark on GSM8K, MATH500, Gaokao, Minerva Math, OlympiadBench, AIME24, AMC23, or asks about evaluating this task. Reports Pass@1 / Avg@32 accuracy.
- ▌ Sparsetir Eval · qhjqhj00Evaluates the runtime performance and speedup of sparse deep learning operators (SpMM, SDDMM) and end-to-end models (GraphSAGE, RGCN, Transformers) on GPU hardware using composable sparse formats and transformations. It probes how format decomposition and modular scheduling primitives improve cache utilization, load balancing, and Tensor Core utilization compared to vendor libraries and existing compilers. Use when the user wants to benchmark on cora, citeseer, pubmed, ppi, ogbn-arxiv, ogbn-proteins, reddit, or asks about evaluating this task. Reports speedup.
- ▌ Spatialqa Eval · qhjqhj00Evaluates the ability of vision-language models to perform multi-step spatial logical reasoning by tracking object dependencies and understanding scene layouts across real-world indoor environments. Use when the user wants to benchmark on SpatiaLQA, or asks about evaluating this task. Reports accuracy.
- ▌ Specpower Eval · qhjqhj00Evaluates the energy proportionality and power efficiency of enterprise server subsystems under varying JVM-based web service workloads. It measures how power consumption scales with workload intensity and identifies non-proportional power draw in uncore components. Use when the user wants to benchmark on SPECpower_ssj2008, or asks about evaluating this task. Reports watts.
- ▌ Speechllm Eval · qhjqhj00Evaluates a unified speech-language model's ability to perform end-to-end automatic speech recognition (ASR), named entity recognition (NER), and sentiment analysis (SA) on low-resource speech datasets. It tests parameter-efficient adapter-based alignment of speech encoder features to a language model, along with classifier regularization and LoRA fine-tuning. Use when the user wants to benchmark on LibriSpeech, SLUE-VoxPopuli, SLUE-VoxCeleb, or asks about evaluating this task. Reports WER, F1, SLUE Score.
- ▌ Sports QA Eval · qhjqhj00Probes video question answering capabilities, specifically focusing on temporal reasoning, action causality, counterfactual inference, and fine-grained motion understanding within professional sports contexts. Use when the user wants to benchmark on Sports-QA, or asks about evaluating this task. Reports accuracy.
- ▌ SQL Synth Eval · qhjqhj00This benchmark evaluates a model's ability to translate natural language questions into correct SQL queries across diverse, real-world database schemas. It specifically probes the model's capacity to handle complex, multi-operation statements and cross-domain syntax structures that are often underrepresented in traditional benchmarks. Use when the user wants to benchmark on SQL-Synth, or asks about evaluating this task. Reports execution accuracy (EX).
- ▌ Ssi Bench Eval · qhjqhj00Probes constrained-manifold spatial reasoning by requiring models to rank structural components based on geometric, topological, and physical constraints in complex 3D engineering scenes. It tests compositional spatial operations like mental rotation, occlusion handling, and force-path reasoning, revealing gaps in structural grounding and 3D constraint consistency. Use when the user wants to benchmark on SSI-Bench, or asks about evaluating this task. Reports Taskwise Accuracy.
- ▌ Sst2 Imdb Eval · qhjqhj00Evaluates the binary sentiment classification capability of hybrid quantum-classical language models on both short and long text sequences. It probes whether adaptive quantum routing and attention mechanisms provide measurable accuracy gains over purely classical or purely quantum baselines on standard NLP benchmarks. Use when the user wants to benchmark on SST-2, IMDB, or asks about evaluating this task. Reports Accuracy.
- ▌ Stablei2i Eval · qhjqhj00Evaluates multimodal models' ability to detect unintended content, structural, and low-level appearance changes in image-to-image transitions. It probes fine-grained visual reasoning and pixel-level alignment capabilities by asking models to assess fidelity across three distinct dimensions. Use when the user wants to benchmark on StableI2I-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Stereoset Eval · qhjqhj00Evaluates stereotypical bias in pretrained language models across gender, profession, race, and religion using context-based association tests. It measures both language modeling capability and the model's tendency to prefer stereotypical over anti-stereotypical associations in natural language contexts. Use when the user wants to benchmark on StereoSet, or asks about evaluating this task. Reports icat.
- ▌ Stgformer Eval · qhjqhj00This evaluation protocol tests a model's ability to forecast future traffic flow across large-scale urban road networks using historical spatiotemporal sensor data. It probes the model's capacity to capture complex spatial dependencies and temporal dynamics while maintaining computational efficiency on real-world traffic benchmarks. Use when the user wants to benchmark on LargeST, PEMS-series, or asks about evaluating this task. Reports MAE.
- ▌ Streammecoeval · qhjqhj00Evaluates the ability of long-term agent memory compression methods to retain and retrieve critical information from streaming and long-form videos. It probes how effectively a model can maintain temporal consistency and answer complex queries while significantly reducing memory graph size, measured via answer accuracy across robotic, web, and general video benchmarks. Use when the user wants to benchmark on M3-Bench-robot, M3-Bench-web, Video-MME-Long, or asks about evaluating this task. Reports accuracy.
- ▌ Streamuni Eval · qhjqhj00Evaluates real-time speech translation systems on latency and translation quality across multiple language pairs. It probes the model's ability to dynamically decide when to generate and truncate translations while processing streaming audio inputs. Use when the user wants to benchmark on MuST-C English→German, MuST-C English→Spanish, CoVoST2 English→Chinese, CoVoST2 French→English, or asks about evaluating this task. Reports SacreBLEU, COMET.
- ▌ Style4rec Eval · qhjqhj00Evaluates a model's ability to perform sequential product recommendation in an e-commerce setting by predicting the next item in a user session. It specifically probes how well the model leverages historical interaction sequences, visual style embeddings, and shopping cart data to rank relevant products. Use when the user wants to benchmark on Style4Rec E-commerce Dataset, or asks about evaluating this task. Reports HR@5.
- ▌ Superchem Eval · qhjqhj00Evaluates deep chemical reasoning capabilities of LLMs using expert-curated, entity-masked multiple-choice problems. It probes both final-answer accuracy and the fidelity of the reasoning process against expert-annotated solution paths, while also assessing the impact of multimodal inputs on complex chemical problem-solving. Use when the user wants to benchmark on SUPERChem-A11, SUPERChem-release, SUPERChem-holdout, SUPERChem-100, Multimodal-Essential Subset, or asks about evaluating this task. Reports pass@1 Accuracy.
- ▌ Superglue Eval · qhjqhj00Evaluates general-purpose language understanding across eight diverse tasks including coreference resolution, question answering, and natural language inference. It probes a model's ability to handle complex reasoning, sample-efficient learning, and transfer learning beyond standard GLUE capabilities. Use when the user wants to benchmark on SuperGLUE, or asks about evaluating this task. Reports SuperGLUE score.
- ▌ Surf Nerf Eval · qhjqhj00Evaluates a NeRF model's ability to reconstruct reflective scenes with high visual fidelity and accurate geometry. It specifically probes the model's capacity to separate Lambertian and specular appearance components while enforcing surface regularisation to resolve shape-radiance ambiguity. Use when the user wants to benchmark on Shiny Objects, Shiny Real, Koala, or asks about evaluating this task. Reports PSNR.
- ▌ Surglaivi Eval · qhjqhj00Evaluates zero-shot and linear-probe transferability of surgical vision-language models across laparoscopic and robotic video modalities. Probes hierarchical procedural understanding (phase, step, action) and object-centric recognition (tools, instrument-verb-target triplets) under varying temporal contexts and data regimes. Use when the user wants to benchmark on Cholec80, AutoLaparo, GraSP, SARRARP50, CholecT50, or asks about evaluating this task. Reports video-wise F1-score.
- ▌ Swe Bench Eval · qhjqhj00Evaluates an agent's ability to automatically resolve software engineering issues by generating and applying code patches to open-source repositories. It measures functional correctness by running the target repository's test suite after the patch is applied. Use when the user wants to benchmark on SWE-bench, HumanEvalFix, or asks about evaluating this task. Reports pass@1.
- ▌ Symile M3 Eval · qhjqhj00Evaluates zero-shot cross-modal retrieval capability by requiring a model to jointly leverage audio and text to identify an image, where neither modality alone contains sufficient information. It tests the model's ability to capture joint information across three distinct high-dimensional data types. Use when the user wants to benchmark on Symile-M3, or asks about evaluating this task. Reports mean accuracy.
- ▌ Syntheory Eval · qhjqhj00Probes whether music foundation models encode discrete and continuous Western music theory concepts by training linear or MLP classifiers on their internal audio embeddings. Use when the user wants to benchmark on SynTheory, or asks about evaluating this task. Reports accuracy.
- ▌ Tabularfm Eval · qhjqhj00Evaluates the cross-dataset transferability of pretrained tabular generative models by measuring how well synthesized data preserves column distributions and pairwise correlations compared to ground truth tables. Use when the user wants to benchmark on Kaggle, GitTables, or asks about evaluating this task. Reports overall average.
- ▌ Tad Bench Eval · qhjqhj00Evaluates the effectiveness of various text embedding models combined with different anomaly detection algorithms for identifying text anomalies. It probes how well embedding-based anomaly detection generalizes across diverse domains (spam, fake news, hate speech) and distinguishes between patterned versus context-dependent anomalies. Use when the user wants to benchmark on Email-Spam, SMS-Spam, COVID-Fake, LIAR2, Hate-Speech, OLID, or asks about evaluating this task. Reports AUROC.
- ▌ Taigi Asr Eval · qhjqhj00Evaluates the ability of automatic speech recognition (ASR) systems to transcribe Taiwanese Hokkien (Taigi) audio into text. It specifically measures how well models map Taigi phonetic patterns to Mandarin character sequences using a standardized set of public service announcement recordings. Use when the user wants to benchmark on Taigi ASR Benchmark, or asks about evaluating this task. Reports Character Error Rate (CER).
- ▌ Tanda As2 Eval · qhjqhj00Evaluates a model's ability to rank candidate answer sentences for a given question. It specifically probes the stability and robustness of pre-trained transformer models when transferred to a target domain using a two-stage fine-tuning protocol. Use when the user wants to benchmark on WikiQA, TREC-QA, or asks about evaluating this task. Reports MAP.
- ▌ Tapvid 3d Eval · qhjqhj00Evaluates a model's ability to track arbitrary points in 3D space over time from monocular video, assessing spatio-temporal consistency, depth accuracy, and occlusion handling. It probes whether models can reconstruct scene geometry up to scale or only maintain relative depth consistency for local interactions. Use when the user wants to benchmark on TAPVid-3D, or asks about evaluating this task. Reports 3D Average Jaccard (3D-AJ).
- ▌ Tau Bench Eval · qhjqhj00Evaluates the task completion success rate of LLM agents performing long-horizon, tool-using agentic workflows across customer service and software engineering domains. Probes the agent's ability to execute environment-altering (mutating) actions safely and maintain goal alignment over extended trajectories. Use when the user wants to benchmark on $ au$-Bench Airline, $ au$-Bench Retail, $ au$-Bench-V Air, $ au$-Bench-V Ret, SWE-Bench Verified, or asks about evaluating this task. Reports score.
- ▌ Tau Voice Eval · qhjqhj00This benchmark evaluates full-duplex voice agents on real-world conversational tasks across retail, airline, and telecom domains. It jointly measures task completion success and real-time interaction quality, including responsiveness, latency, interruption handling, and selectivity under varying acoustic conditions like noise, diverse accents, and natural turn-taking. Use when the user wants to benchmark on $\tau^2$-bench, or asks about evaluating this task. Reports pass@1.
- ▌ Taxpraben Eval · qhjqhj00Evaluates LLMs on Chinese real-world tax practice tasks spanning classification, generation, structured prediction, and mixed matching. It probes capabilities across Bloom's taxonomy levels, from factual recall and understanding to complex tax strategy planning and risk prevention. Use when the user wants to benchmark on TaxPraBen, or asks about evaluating this task. Reports Overall Average.
- ▌ Tdc Admet Eval · qhjqhj00Evaluates molecular property prediction across 22 ADMET tasks. It probes the model's ability to generalize across diverse chemical properties using standardized benchmark splits for both regression and classification. Use when the user wants to benchmark on TDC ADMET Group, or asks about evaluating this task. Reports Classification AUROC.
- ▌ Tempo Sum Eval · qhjqhj00Evaluates text summarization models' temporal generalization by testing on datasets split by publication date, specifically probing how well models handle knowledge-conflicting future articles versus in-distribution past data. Use when the user wants to benchmark on BBC, CNN, or asks about evaluating this task. Reports FactCC.
- ▌ Tempqa Wd Eval · qhjqhj00Probes a system's ability to perform temporal question answering over knowledge bases by generating correct answers or SPARQL queries. It specifically evaluates generalization across different knowledge bases (Wikidata vs. Freebase) and interpretability through fine-grained intermediate annotations like entity/relation linking and λ-expressions. Use when the user wants to benchmark on TempQA-WD, or asks about evaluating this task. Reports F1.
- ▌ Tenspiler Eval · qhjqhj00Evaluates the correctness and performance of a verified lifting-based compiler that transpiles sequential C++/Python code into tensor operations across various DSLs and hardware accelerators. It measures synthesis efficiency, kernel execution speedup, and end-to-end performance including data transfer overhead. Use when the user wants to benchmark on TENSPILER benchmark suite, or asks about evaluating this task. Reports kernel_performance.
- ▌ Tesseract Eval · qhjqhj00Evaluates the robustness of Android malware classifiers against spatio-temporal experimental bias. It measures how model performance degrades when trained on past application data and tested on future data, while accounting for realistic malware-to-goodware class distributions. Use when the user wants to benchmark on Android malware dataset (2014-2016), or asks about evaluating this task. Reports F1-Score.
- ▌ Textatlas Eval · qhjqhj00Evaluates text-to-image generation models on their ability to render long, dense, and structurally complex text accurately within images. It probes both semantic alignment between the prompt and the generated image, and precise character/word-level OCR fidelity across diverse layouts, styles, and real-world scenes. Use when the user wants to benchmark on TextAtlasEval, or asks about evaluating this task. Reports OCR Accuracy (Acc.).
- ▌ Tfq Bench Eval · qhjqhj00Evaluates a model's ability to understand visual metaphors and image implications by verifying multiple factual and inferential propositions per image. It probes fine-grained visual perception, multi-hop reasoning, and theory of mind through structured true-false questioning. Use when the user wants to benchmark on TFQ-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Tg Redial Eval · qhjqhj00Evaluates a conversational recommender system's ability to naturally transition topics, recommend relevant items, and generate coherent responses within a dialogue. It probes the model's capacity to leverage historical interactions, user profiles, and topic sequences to maintain semantic flow and recommendation accuracy. Use when the user wants to benchmark on TG-ReDial, or asks about evaluating this task. Reports NDCG@k.
- ▌ Theoremqa Eval · qhjqhj00Evaluates LLMs' ability to apply domain-specific theorems from mathematics, physics, computer science, and finance to solve complex scientific problems. It probes theorem-driven reasoning, numerical computation, and program generation capabilities. Use when the user wants to benchmark on TheoremQA, or asks about evaluating this task. Reports accuracy.
- ▌ Tib Bench Eval · qhjqhj00Evaluates vision-language models on their ability to generate accurate and linguistically coherent summaries of text-heavy multimodal presentations. It probes cross-modal alignment, long-context understanding, and the impact of different input modalities (raw video, slides, transcripts, interleaved pairs) and token budgets on summarization quality. Use when the user wants to benchmark on TIB-bench, or asks about evaluating this task. Reports $IbR_{overall}$.
- ▌ Timit Tts Eval · qhjqhj00Evaluates deepfake detectors on distinguishing real from synthetic audio and video, and on attributing synthetic audio to specific TTS generators. It probes robustness to post-processing (DTW alignment, augmentation) and video compression. Use when the user wants to benchmark on TIMIT-TTS, or asks about evaluating this task. Reports AUC.
- ▌ Tir Bench Eval · qhjqhj00Evaluates multimodal language models' ability to perform agentic reasoning with images, specifically requiring dynamic visual manipulation and tool-use to solve complex tasks like rotation, jigsaw assembly, and instrument reading. It probes whether models can iteratively process, crop, or transform visual inputs to extract information or solve spatial problems. Use when the user wants to benchmark on TIR-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Tmmluplus Eval · qhjqhj00Evaluates foundation models' multitask language understanding and reasoning capabilities in Traditional Chinese across diverse academic subjects including STEM, social sciences, humanities, and other domains. Use when the user wants to benchmark on TMMLU+, or asks about evaluating this task. Reports average accuracy (%).
- ▌ Tool Star Eval · qhjqhj00Evaluates an LLM's ability to autonomously invoke and coordinate multiple tools (search, code, calculator, etc.) for complex reasoning tasks. It probes both computational reasoning (math) and knowledge-intensive reasoning (open-domain QA) under a multi-tool collaborative setting. Use when the user wants to benchmark on AIME2024, AIME2025, MATH500, MATH, GSM8K, GAIA, HLE, WebWalker, HotpotQA, 2WikiMultihopQA, Musique, Bamboogle, or asks about evaluating this task. Reports LLM-based judging accuracy.
- ▌ Toolbench Eval · qhjqhj00This benchmark evaluates an LLM's ability to understand API documentation, select the correct tool, and generate valid arguments to fulfill a user's goal. It probes tool manipulation capabilities across a diverse set of 8 real-world applications, ranging from simple single-call tasks to complex multi-step reasoning. Use when the user wants to benchmark on ToolBench, or asks about evaluating this task. Reports success rate.
- ▌ Totalvariation · qhjqhj00Compute the TotalVariation metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute TotalVariation, or asks how to score with TotalVariation.
- ▌ Touchdown Eval · qhjqhj00Evaluates an agent's capacity to navigate through street-view environments using natural language instructions and resolve complex spatial descriptions to locate a hidden target object within a panoramic image. Use when the user wants to benchmark on Touchdown, or asks about evaluating this task. Reports pixel distance.
- ▌ Tox21 Fsl Eval · qhjqhj00Evaluates graph neural networks for few-shot toxic molecule classification. It probes the model's ability to generalize from very limited labeled examples (shots) and adapt to new query sets using meta-learning and graph augmentation techniques. Use when the user wants to benchmark on Tox21, or asks about evaluating this task. Reports ROC-AUC Score.
- ▌ Trans Env Eval · qhjqhj00Evaluates the linguistic robustness of LLMs by measuring performance degradation when standard English prompts are transformed into 38 regional dialects and ESL varieties. It probes whether models maintain accuracy and instruction-following capabilities across non-standard linguistic variations. Use when the user wants to benchmark on MMLU, ARC, TruthfulQA, GSM8K, HellaSwag, WinoGrande, IFEval, AlpacaFarm, MT-Bench, or asks about evaluating this task. Reports accuracy.
- ▌ Treebench Eval · qhjqhj00Evaluates visual grounded reasoning by requiring models to precisely localize target objects in cluttered scenes and perform second-order reasoning about their interactions. It measures both the correctness of the final answer and the spatial accuracy of the traceable bounding box evidence. Use when the user wants to benchmark on TreeBench, or asks about evaluating this task. Reports Accuracy.
- ▌ Tsg Bench Eval · qhjqhj00Evaluates large language models' ability to understand and generate structured scene graphs from textual narratives. It probes spatial reasoning, action decomposition, and the capacity to map dynamic descriptions to discrete visual or structural elements. Use when the user wants to benchmark on TSG Bench, or asks about evaluating this task. Reports Exact Match (EM) / Accuracy, Precision, Recall, Macro F1.
- ▌ Tunevlseg Eval · qhjqhj00Evaluates the robustness and zero-shot adaptation of vision-language segmentation models under prompt tuning across diverse medical and natural domain datasets. It probes how different prompt tuning strategies and prompt depths handle domain shifts and varying class counts. Use when the user wants to benchmark on Kvasir-SEG, ClinicDB, BKAI, ISIC 2016, DFU 2022, CAMUS, BUSI, CheXlocalize, Cityscapes, PascalVOC, or asks about evaluating this task. Reports Dice score.
- ▌ Tweeteval Eval · qhjqhj00Evaluates the ability of language models to classify short social media posts across seven distinct Twitter-specific tasks, including sentiment, emotion, hate speech, and irony detection. It probes domain adaptation by comparing models pre-trained on generic text versus those further trained on large-scale Twitter corpora. Use when the user wants to benchmark on TweetEval, or asks about evaluating this task. Reports M-F1.
- ▌ Ultrachat Eval · qhjqhj00Evaluates a chat model's ability to generate accurate, informative, and correct responses across diverse domains including commonsense, world knowledge, professional knowledge, mathematics, reasoning, and writing. It probes both factual correctness and response quality using automated LLM-based pairwise and independent scoring. Use when the user wants to benchmark on UltraChat Evaluation Set, or asks about evaluating this task. Reports ChatGPT scoring.