Model Training & Fine-tuning
-
qhjqhj00 Skill AurocComputes the AUROC metric using torchmetrics, handling binary, multiclass, and multilabel tasks with configurable thresholds and averaging.
Audited 3 -
qhjqhj00 Skill MenliEvaluates the robustness and alignment with human judgment of reference-based and reference-free evaluation metrics for machine translation and summarization, particularly under adversarial conditions.
Audited 3 -
qhjqhj00 Skill PolosScores generated image captions against reference captions and source images using the Polos metric, which is trained to align with human judgments and probes hallucination robustness and open-vocabulary evaluation.
Audited 3 -
qhjqhj00 Skill ScoreAudits medical LLM benchmarks across five lifecycle phases using 46 medically tailored criteria to assess clinical relevance, data integrity, safety-critical capabilities, validity, and governance.
Audited 3 -
qhjqhj00 Skill TtsdsEvaluates text-to-speech systems by measuring distributional distance between synthetic and real speech across five factors, producing a scalar score without subjective MOS ratings.
Audited 3 -
qhjqhj00 Skill VisorEvaluates text-to-image models on spatial relationship accuracy using the VISOR metric, separating object detection from spatial correctness to reveal biases like object priority and merging.
Audited 3 -
qhjqhj00 Bundle Umap LearnReduce high-dimensional data with UMAP for visualization, clustering preprocessing, and supervised or semi-supervised learning, including parameter tuning guidance.
Audited 3 -
qhjqhj00 Skill BleurtEvaluates the correlation between automatic text generation scores and human quality ratings, including robustness to domain and quality drift, using metrics like Kendall's Tau and Pearson correlation.
Audited 3 -
qhjqhj00 Skill InfolmComputes the InfoLM metric from torchmetrics for evaluating text generation against ground truth, with configurable information measures and sentence-level scoring.
Audited 3 -
qhjqhj00 Skill LogaucComputes the LogAUC metric using the torchmetrics implementation for binary, multiclass, or multilabel classification tasks.
Audited 3 -
qhjqhj00 Skill RecallComputes the Recall metric using torchmetrics, including configuration for binary, multiclass, and multilabel tasks.
Audited 3 -
qhjqhj00 Skill ReflexEvaluates machine-generated log summaries without human-written references, using LLM judgment and dense embeddings to score relevance, informativeness, and coherence.
Audited 3 -
qhjqhj00 Bundle TensorboardVisualize training metrics, debug models with histograms, compare experiments, visualize model graphs, and profile performance with TensorBoard.
Audited 3 -
qhjqhj00 Skill EpsilonEvaluates the correlation between a zero-cost NAS metric (epsilon) and actual training accuracy across different neural architecture search spaces, testing the metric's ability to rank architectures without training. It probes whether output dispersion from constant weight initializations can serve as a reliable.
Audited 3 -
qhjqhj00 Skill F1scoreCompute the F1Score metric using torchmetrics when predictions and ground-truth labels are available.
Audited 3 -
qhjqhj00 Skill R2scoreComputes the R2Score metric using torchmetrics, handling single and multi-output predictions with options for adjusted and variance-weighted scores.
Audited 3 -
qhjqhj00 Skill RuntimeBenchmarks inference latency and computational runtime of transformer models and MLX operations across Apple Silicon and NVIDIA GPU backends, with configurable input lengths and batch sizes.
Audited 3 -
qhjqhj00 Skill T5 EvalBenchmarks a text-to-text transformer across GLUE, SuperGLUE, CNN/Daily Mail, SQuAD, and WMT, reporting GLUE average, BLEU, ROUGE-2-F, and Exact Match scores.
Audited 3 -
qhjqhj00 Skill Abc EvalBenchmarks large language models on symbolic music understanding and instruction following using text-based ABC notation, covering syntax parsing, error detection, segment-level reasoning, and sequence-level musical analysis.
Audited 3 -
qhjqhj00 Skill AndersonComputes the Anderson-Darling test statistic and p-value using scipy.stats.anderson for evaluating predictions against ground truth.
Audited 3 -
qhjqhj00 Skill Ape EvalBenchmarks automatic post-editing (APE) models on WMT'18 SMT, SubEdits, and MLQE-PE datasets, reporting BLEU, ChrF, and TER scores computed with SacreBLEU and TERCOM.
Audited 3 -
qhjqhj00 Skill Arc EvalBenchmarks systems on the Abstraction and Reasoning Corpus (ARC) by requiring inference of abstract transformation rules from few input-output grid demonstrations and application to novel test cases, reporting the fraction of tasks solved.
Audited 3 -
qhjqhj00 Skill Aya EvalEvaluates open-ended generation quality of multilingual LLMs across brainstorming, planning, and long-form tasks, using AYA and DOLLY datasets with qualitative fluency and quality scoring.
Audited 3 -
qhjqhj00 Skill Bbh EvalBenchmarks zero-shot in-context learning on BIG-Bench Hard multiple-choice tasks, comparing self-generated demonstrations against direct prompting and chain-of-thought baselines, and reports accuracy.
Audited 3 -
qhjqhj00 Skill Bbq EvalEvaluates social bias in question-answering models using the BBQ benchmark, measuring accuracy and a bias score across ambiguous and disambiguated contexts to reveal reliance on stereotypes.
Audited 3 -
qhjqhj00 Skill Caa EvalBenchmarks large audio-language models against adversarial audio attacks using the CAA dataset, computing WER, ROUGE-L, cosine similarity, and coherence scores to assess robustness in conversational settings.
Audited 3 -
qhjqhj00 Skill Cab EvalBenchmarks LLM bias by scoring responses to automatically generated open-ended questions across sensitive attributes, producing a composite fitness score from 0 to 5.
Audited 3 -
schattenspiegel Bundle Numpyro PythonWrite, debug, and test NumPyro probabilistic programs on JAX with correct shapes, PRNG keys, and inference choice.
Audited 0 -
google Bundle Agent Platform TuningFine-tune open models or Gemini models using Agent Platform infrastructure, from environment setup through data preparation, job configuration, monitoring, and deployment.
14.4k -
nvidia Bundle Tao Train Mask2formerTrain, evaluate, export, quantize, and run inference on Mask2Former models for panoptic, instance, and semantic segmentation using NVIDIA TAO.
Audited 2.2k -
nvidia Bundle Nv Generate Vae FinetuneFinetune the NV-Generate-CTMR MAISI VAE/autoencoder on user-supplied CT or MRI NIfTI volumes using a staged config and datalist workflow.
Audited 2.2k -
nvidia Bundle Tao Train Deformable DetrTrain, evaluate, export, quantize, and run inference for a Deformable DETR 2D object detection model using TAO, with deformable attention for efficient multi-scale feature processing.
Audited 2.2k -
nvidia Bundle Tao Validate Dataset FormatValidates NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors using the `tao-daft validate` CLI tool.
Audited 2.2k -
nvidia Bundle Nemo Mbridge Multi Node SlurmConvert single-node PyTorch distributed scripts into multi-node Slurm sbatch jobs and debug common multi-node failures, covering srun-native and torch.distributed approaches, container setup, NCCL timeouts, and interactive allocation.
Audited 2.2k -
nvidia Bundle Nemo Mbridge Perf Moe Comm OverlapOptimizes MoE expert-parallel communication overlap in Megatron Bridge, covering dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.
Audited 2.2k -
google Skill Bigquery BigframesGenerates Python code using BigQuery DataFrames (BigFrames), the pandas/scikit-learn-style API over BigQuery, for dataframe and ML workflows.
Audited 14.4k