qhjqhj00
- 7.6k skills
- 0 followers
- 3 repo stars
- 2 weeks ago last updated
- ▌ Pixel Art · qhjqhj00 bundleGenerates pixel art SVG illustrations for READMEs, docs, or slides, with character templates, chat bubbles, and scene composition guidance.
- ▌ Pennylane · qhjqhj00 bundleTrain quantum circuits with automatic differentiation and build hybrid quantum-classical models using PennyLane, including VQE, QAOA, and integration with PyTorch, JAX, and TensorFlow.
- ▌ Geopandas · qhjqhj00 bundlePerforms geospatial vector data analysis with GeoPandas, including reading/writing shapefiles, GeoJSON, GeoPackage, and PostGIS, geometric operations, spatial joins, overlays, coordinate transformations, and map visualization.
- ▌ Auroc · qhjqhj00Computes the AUROC metric using torchmetrics, handling binary, multiclass, and multilabel tasks with configurable thresholds and averaging.
- ▌ Cider · qhjqhj00Computes CIDEr and related metrics to score how well generated image descriptions align with human consensus, using reference sentences and triplet annotations.
- ▌ Menli · qhjqhj00Evaluates the robustness and alignment with human judgment of reference-based and reference-free evaluation metrics for machine translation and summarization, particularly under adversarial conditions.
- ▌ Polos · qhjqhj00Scores generated image captions against reference captions and source images using the Polos metric, which is trained to align with human judgments and probes hallucination robustness and open-vocabulary evaluation.
- ▌ Score · qhjqhj00Audits medical LLM benchmarks across five lifecycle phases using 46 medically tailored criteria to assess clinical relevance, data integrity, safety-critical capabilities, validity, and governance.
- ▌ Spice · qhjqhj00Evaluates image captions by converting them into scene graphs and computing an F-score over semantic propositions, measuring how well a generated caption captures the meaning of an image compared to human references.
- ▌ Squad · qhjqhj00Computes the SQuAD metric using torchmetrics, given predictions and ground truth. Use when evaluating question-answering outputs with exact match and F1 scores.
- ▌ Ttsds · qhjqhj00Evaluates text-to-speech systems by measuring distributional distance between synthetic and real speech across five factors, producing a scalar score without subjective MOS ratings.
- ▌ Visor · qhjqhj00Evaluates text-to-image models on spatial relationship accuracy using the VISOR metric, separating object detection from spatial correctness to reveal biases like object priority and merging.
- ▌ Umap Learn · qhjqhj00 bundleReduce high-dimensional data with UMAP for visualization, clustering preprocessing, and supervised or semi-supervised learning, including parameter tuning guidance.
- ▌ Distributed LLM Pretraining Torchtitan · qhjqhj00 bundlePretrains large language models at scale using PyTorch-native torchtitan with 4D parallelism, Float8, and distributed checkpointing.
- ▌ Matplotlib · qhjqhj00 bundleCreate publication-quality static, animated, and interactive plots with fine-grained control over every element, from basic charts to multi-panel figures, with export to PNG, PDF, and SVG.
- ▌ 125 Yuv Design · qhjqhj00 bundleApplies a battle-tested bilingual web-design system with typography, responsive rules, and performance patterns for Yuval Avidani's projects.
- ▌ Bleurt · qhjqhj00Evaluates the correlation between automatic text generation scores and human quality ratings, including robustness to domain and quality drift, using metrics like Kendall's Tau and Pearson correlation.
- ▌ Infolm · qhjqhj00Computes the InfoLM metric from torchmetrics for evaluating text generation against ground truth, with configurable information measures and sentence-level scoring.
- ▌ L Eval · qhjqhj00Benchmarks long-context language models across 20 sub-tasks spanning 3k–200k tokens, covering retrieval, reasoning, summarization, and instruction understanding, with exact-match accuracy as the primary metric.
- ▌ Lambre · qhjqhj00Scores generated text for morphosyntactic well-formedness by measuring how closely it adheres to language-specific dependency rules extracted from treebanks.
- ▌ Logauc · qhjqhj00Computes the LogAUC metric using the torchmetrics implementation for binary, multiclass, or multilabel classification tasks.
- ▌ Recall · qhjqhj00Computes the Recall metric using torchmetrics, including configuration for binary, multiclass, and multilabel tasks.
- ▌ Reflex · qhjqhj00Evaluates machine-generated log summaries without human-written references, using LLM judgment and dense embeddings to score relevance, informativeness, and coherence.
- ▌ Stream · qhjqhj00Evaluates spatial realism and temporal flow consistency of AI-generated videos using embedding spaces and Fourier transforms, producing bounded STREAM-S and STREAM-T scores.
- ▌ Vpeval · qhjqhj00Evaluates text-to-image generation models by decomposing assessment into five specialized skills (object presence, count, spatial relations, scale, and text rendering) and open-ended prompts, producing interpretable binary scores with visual and textual explanations.
- ▌ Paper Audit · qhjqhj00 bundleProvides reviewer-style audit and deep review for academic papers in LaTeX, Typst, and PDF formats, producing structured issue bundles, pass/fail gates, and revision roadmaps.
- ▌ Self Review · qhjqhj00 bundleReviews an academic paper using the NeurIPS review form with three reviewer personas, ensemble scoring, and reflection refinement. Extracts text from PDF, runs structured review, and outputs actionable feedback.
- ▌ Tensorboard · qhjqhj00 bundleVisualize training metrics, debug models with histograms, compare experiments, visualize model graphs, and profile performance with TensorBoard.
- ▌ A3 Eval · qhjqhj00Benchmarks mobile GUI agents on multi-step tasks across 20 Android apps, measuring task completion and essential-state navigation with Task Success Rate and Essential State Achieved Rate.
- ▌ Apessrc · qhjqhj00Evaluates the faithfulness of abstractive summaries by verifying if factual claims (masked as cloze questions) in the reference summary can be correctly answered using only the generated summary, compared against a gold-standard answer derived from the source context.
- ▌ Epsilon · qhjqhj00Evaluates the correlation between a zero-cost NAS metric (epsilon) and actual training accuracy across different neural architecture search spaces, testing the metric's ability to rank architectures without training. It probes whether output dispersion from constant weight initializations can serve as a reliable.
- ▌ Latency · qhjqhj00Measures inference latency of binarized, 8-bit, and 32-bit convolutional layers on edge devices to evaluate the efficiency and speedup of the Larq Compute Engine framework compared to standard implementations.
- ▌ Ndcg 10 · qhjqhj00Evaluates how well internal model representations (hidden states) predict token-level information importance in summarization tasks, using NDCG@10 and Spearman's rank correlation.
- ▌ R2score · qhjqhj00Computes the R2Score metric using torchmetrics, handling single and multi-output predictions with options for adjusted and variance-weighted scores.
- ▌ Runtime · qhjqhj00Benchmarks inference latency and computational runtime of transformer models and MLX operations across Apple Silicon and NVIDIA GPU backends, with configurable input lengths and batch sizes.
- ▌ T5 Eval · qhjqhj00Benchmarks a text-to-text transformer across GLUE, SuperGLUE, CNN/Daily Mail, SQuAD, and WMT, reporting GLUE average, BLEU, ROUGE-2-F, and Exact Match scores.
- ▌ Theilsu · qhjqhj00Computes Theil's U (uncertainty coefficient) between predictions and ground truth using the torchmetrics implementation, handling categorical data and NaN strategies.
- ▌ Tpr Fpr · qhjqhj00Evaluates speaker verification models by computing true positive rate at fixed false positive rate thresholds, probing embedding space separation of same-speaker versus different-speaker pairs.
- ▌ Usfiscaldata · qhjqhj00 bundleQuery the U.S. Treasury Fiscal Data API for federal financial data including national debt, government spending, revenue, interest rates, exchange rates, and savings bonds. Access 54 datasets and 182 data tables with no API key required.
- ▌ Copilot Docs · qhjqhj00 bundleConfigure repository-specific guidance for GitHub Copilot by creating and structuring .github/copilot-instructions.md files.
- ▌ Abc Eval · qhjqhj00Benchmarks large language models on symbolic music understanding and instruction following using text-based ABC notation, covering syntax parsing, error detection, segment-level reasoning, and sequence-level musical analysis.
- ▌ Accuracy · qhjqhj00Evaluates an AI judge system's pairwise ranking accuracy on generated commit messages against a heuristic ground truth from five automatic text metrics, using the MCMD dataset.
- ▌ Adp Eval · qhjqhj00Benchmarks LLM agents fine-tuned with the Agent Data Protocol across software engineering, web browsing, OS/database tool use, and reasoning tasks, reporting unit test pass rates and task success rates.
- ▌ Anderson · qhjqhj00Computes the Anderson-Darling test statistic and p-value using scipy.stats.anderson for evaluating predictions against ground truth.
- ▌ Ape Eval · qhjqhj00Benchmarks automatic post-editing (APE) models on WMT'18 SMT, SubEdits, and MLQE-PE datasets, reporting BLEU, ChrF, and TER scores computed with SacreBLEU and TERCOM.
- ▌ Arc Eval · qhjqhj00Benchmarks systems on the Abstraction and Reasoning Corpus (ARC) by requiring inference of abstract transformation rules from few input-output grid demonstrations and application to novel test cases, reporting the fraction of tasks solved.
- ▌ Art Eval · qhjqhj00Benchmarks medical AI agents on synthetic EHR tasks, measuring success rates for data retrieval, temporal aggregation, and threshold-based conditional logic with exact-match scoring.
- ▌ Jq · qhjqhj00 bundleQuery, filter, transform, and aggregate JSON data using jq, with practical patterns for shell pipelines and integration with CLI tools.
- ▌ Hqq Quantization · qhjqhj00 bundleQuantize LLMs to 8/4/3/2/1-bit precision without calibration data, using multiple backends and HuggingFace/vLLM integration.
- ▌ Adhx · qhjqhj00 bundleFetches any X/Twitter post as structured JSON via the ADHX API, including full article content, author info, and engagement metrics, without scraping or a browser.
- ▌ Tmux · qhjqhj00 bundleManage persistent terminal sessions, windows, and panes with tmux, including scripting multi-pane layouts and automating commands from bash.
- ▌ Aeon · qhjqhj00 bundlePerforms time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search using the aeon toolkit.
- ▌ Shap · qhjqhj00 bundleExplains machine learning model predictions using SHAP values, covering feature importance, visualization plots, model debugging, bias analysis, and production deployment.
- ▌ PPTX · qhjqhj00 bundleCreates, edits, and analyzes PowerPoint presentations by converting to markdown, unpacking raw XML, and applying design principles for new slides.
- ▌ Qutip · qhjqhj00 bundleSimulate open and closed quantum systems with QuTiP, covering master equations, Lindblad dynamics, decoherence, and quantum optics.
- ▌ Pymoo · qhjqhj00 bundleSolve single- and multi-objective optimization problems with NSGA-II/III, MOEA/D, and other evolutionary algorithms, including Pareto front analysis, constraint handling, and benchmarking on standard test problems.
- ▌ Mlflow · qhjqhj00 bundleTrack ML experiments, manage the model registry with versioning, deploy models, and reproduce experiments using MLflow's framework-agnostic platform.
- ▌ Review · qhjqhj00 bundleRoutes quality reviews to specialized critic agents based on file type or flags, covering peer review, code review, and manuscript polish.
- ▌ Depmap · qhjqhj00 bundleQuery the Cancer Dependency Map (DepMap) for CRISPR gene dependency scores, drug sensitivity data, and gene effect profiles to identify cancer-specific vulnerabilities, synthetic lethal interactions, and validate oncology drug targets.
- ▌ Hf MCP · qhjqhj00 bundleConnects AI assistants to the Hugging Face Hub via MCP server tools to search models, datasets, Spaces, and papers, retrieve repository details and documentation, run compute jobs, and use Gradio Spaces as AI tools.
- ▌ Slidev · qhjqhj00 bundleCreate and present web-based slidedecks for developers using Slidev with Markdown, Vue components, code highlighting, animations, and interactive features.
- ▌ Plotly · qhjqhj00 bundleCreates interactive Plotly visualizations in Python, covering Express and Graph Objects for scatter, line, bar, heatmap, 3D, and geographic charts, plus subplots, styling, and HTML export.
- ▌ Seaborn · qhjqhj00 bundleCreate publication-quality statistical graphics in Python with dataset-oriented plotting, semantic mapping, and built-in statistical estimation.
- ▌ Diagram Skills · qhjqhj00 bundleProvides guides for creating diagrams and visualizations using Mermaid, Excalidraw, PlantUML, TikZ, and other tools, covering flowcharts, architecture diagrams, and scientific illustrations.
- ▌ Pydicom · qhjqhj00 bundleRead, write, and manipulate DICOM medical imaging files, including pixel data extraction, metadata editing, anonymization, format conversion, and compression handling.
- ▌ Wiki QA · qhjqhj00 bundleAnswers questions about a code repository by analyzing source files and citing evidence with linked citations.
- ▌ Phoenix Observability · qhjqhj00 bundleSelf-hosted observability platform for LLM applications, providing tracing, evaluation, datasets, experiments, and real-time monitoring to debug and improve AI systems.
- ▌ Auc · qhjqhj00Evaluates machine learning classifiers on their ability to distinguish signal from background in particle physics simulations, measuring how well algorithms rank signal events above background ones using the AUC metric.
- ▌ Eas · qhjqhj00Validates the Emotional Attitude Score (EAS) metric by measuring its consistency with human judgment on word-level sentiment polarity, using the AmbGIMT dataset and pairwise score comparisons.
- ▌ Eer · qhjqhj00Compute the Equal Error Rate (EER) metric using torchmetrics for binary, multiclass, or multilabel classification tasks, with reference signatures and usage examples.
- ▌ Fid · qhjqhj00Measures distributional similarity between original GAN-generated images and their semantically manipulated counterparts using the Fréchet Inception Distance (FID) metric.
- ▌ Mos · qhjqhj00Evaluates the naturalness, speaker similarity, and real-time synthesis speed of a Mandarin speech cloning system across diverse practical application scenarios.
- ▌ Qqe · qhjqhj00Computes bibliometric indices for AI/NLP conferences, including QQE, average/median citations, and citation inequality, from annual publication and citation data.
- ▌ Roc · qhjqhj00Computes the Receiver Operating Characteristic (ROC) metric using torchmetrics, supporting binary, multiclass, and multilabel tasks.
- ▌ Sdr · qhjqhj00Quantifies audio source separation quality by computing the signal-to-distortion ratio (SDR) between ground-truth and estimated stems, with per-stem and record-level averaging.
- ▌ Tec · qhjqhj00Measures the trade-off between computation time and energy consumption in mobile edge computing by computing a weighted sum of the two objectives, given system configuration parameters and per-user task characteristics.
- ▌ Ara Compiler · qhjqhj00 bundleCompiles any research input — PDF papers, GitHub repositories, experiment logs, code directories, or raw notes — into a complete Agent-Native Research Artifact (ARA) with cognitive layer (claims, concepts, heuristics), physical layer (configs, code stubs), exploration graph, and grounded evidence. Use when ingesting a.
- ▌ Ray Data · qhjqhj00 bundleProcess large ML datasets in parallel across CPU or GPU clusters, with streaming execution, multi-format I/O, and integration with Ray Train, PyTorch, and TensorFlow for batch inference and preprocessing pipelines.
- ▌ Psnr · qhjqhj00Evaluates the trade-off between file size reduction and image fidelity when encoding radio astronomy data using JPEG2000, benchmarking both lossless and lossy compression modes to determine the compression ratio at which visual artifacts first appear.
- ▌ Flops · qhjqhj00Evaluates computational throughput and real-time efficiency of embedded CPU and GPU platforms by measuring peak FLOPS via a matrix rotation kernel and assessing inference latency and power consumption on a robotic vision pipeline.
- ▌ F1score · qhjqhj00Compute the F1Score metric using torchmetrics when predictions and ground-truth labels are available.
- ▌ Kruskal · qhjqhj00Compute the Kruskal-Wallis H-test using scipy.stats.kruskal for independent samples, returning the H statistic and p-value.