Model Training & Fine-tuning
-
qhjqhj00 Bundle ShapExplains machine learning model predictions using SHAP values, covering feature importance, visualization plots, model debugging, bias analysis, and production deployment.
Audited 3 -
qhjqhj00 Skill AucEvaluates machine learning classifiers on their ability to distinguish signal from background in particle physics simulations, measuring how well algorithms rank signal events above background ones using the AUC metric.
Audited 3 -
qhjqhj00 Bundle Ray DataProcess large ML datasets in parallel across CPU or GPU clusters, with streaming execution, multi-format I/O, and integration with Ray Train, PyTorch, and TensorFlow for batch inference and preprocessing pipelines.
Audited 3 -
qhjqhj00 Skill GecoEvaluates geometric consistency in text-to-video generation by measuring structural and motion coherence across camera trajectories, detecting deformation and occlusion artifacts in static scenes.
Audited 3 -
qhjqhj00 Skill TctbEvaluates the throughput and resource allocation efficiency of RIS-aided mobile edge computing systems by measuring the total computation task bits successfully completed under varying network conditions.
Audited 3 -
qhjqhj00 Skill CiderComputes CIDEr and related metrics to score how well generated image descriptions align with human consensus, using reference sentences and triplet annotations.
Audited 3 -
qhjqhj00 Skill SpiceEvaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
Audited 3 -
qhjqhj00 Bundle Distributed LLM Pretraining TorchtitanPretrains large language models at scale using PyTorch-native torchtitan with 4D parallelism, Float8, and distributed checkpointing.
3 -
qhjqhj00 Skill L EvalBenchmarks long-context language models across 20 sub-tasks spanning 3k–200k tokens, covering retrieval, reasoning, summarization, and instruction understanding, with exact-match accuracy as the primary metric.
Audited 3 -
qhjqhj00 Skill StreamEvaluates spatial realism and temporal flow consistency of AI-generated videos using embedding spaces and Fourier transforms, producing bounded STREAM-S and STREAM-T scores.
Audited 3 -
qhjqhj00 Skill LatencyMeasures inference latency of binarized, 8-bit, and 32-bit convolutional layers on edge devices to evaluate the efficiency and speedup of the Larq Compute Engine framework compared to standard implementations.
Audited 3 -
qhjqhj00 Skill TheilsuComputes Theil's U (uncertainty coefficient) between predictions and ground truth using the torchmetrics implementation, handling categorical data and NaN strategies.
Audited 3 -
qhjqhj00 Skill AccuracyEvaluates an AI judge system's pairwise ranking accuracy on generated commit messages against a heuristic ground truth from five automatic text metrics, using the MCMD dataset.
Audited 3 -
qhjqhj00 Skill Art EvalBenchmarks medical AI agents on synthetic EHR tasks, measuring success rates for data retrieval, temporal aggregation, and threshold-based conditional logic with exact-match scoring.
Audited 3 -
qhjqhj00 Skill Bis EvalBenchmarks energy-function-based safe control algorithms on the BIS (Benchmark of Interactive Safety) dataset, scoring safety, efficiency, and hybrid performance in human-robot and robot co-working scenarios.
Audited 3