Results for “nlp-evaluation”

54 skills
More results
brycewang-stanford
Acl Experiments
Use when designing or auditing experiments for an ACL paper, covering tuned LLM baselines, multi-dataset and multilingual evaluation, statistical significance and variance, human evaluation with agreement reporting, contamination and prompt-sensitivity controls, ablations, and error-analysis expectations in NLP reviewing.
1k
orchestra-research
Nemo Evaluator Sdk
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution on local Docker, Slurm HPC, or cloud platforms.
10.4k · bundle
qcmuu
Evaluating Llms Harness
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
0 · bundle
shulkwisec
Network Assess
Internal network assessment. VLAN hopping, ARP spoofing detection, broadcast protocol abuse (LLMNR/NBT-NS/mDNS), network segmentation verification, SNMP enumeration, NFS exposure, router/switch audit, and internal service mapping. Assumes attacker has network access. Uses nmap, arp-scan, nbtscan, snmpwalk, onesixtyone, smbmap, nfs-common, masscan, hping3, and netexec.
21
qcmuu
Nemo Evaluator Sdk
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
0 · bundle
tianhao909
Nemo Evaluator Sdk
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
1 · bundle
samyakjhaveri
Paper Review Sim
Simulates a NeurIPS/SC/ICSE-style peer review with five reviewer personas (HPC, ML, Stats, Reproducibility, Devil's Advocate) that verify every claim against actual result data before submission.
0
omer-metin
Nlp Advanced
Use when extracting structured information from text - named entity recognition, relation extraction, coreference resolution, knowledge graph construction, and information extraction pipelinesUse when ", " mentioned.
128 · bundle
bytesagain
Nlp
Process text with NLP. Use when tokenizing, analyzing sentiment, extracting entities, summarizing documents, or measuring similarity.
12 · bundle
github
Phoenix Evals
Build and run evaluators for AI/LLM applications using Phoenix, covering error analysis, custom evaluators, experiments, and production monitoring.
36.2k · bundle
nvidia
Cuopt Numerical Optimization API
Model and solve LP, MILP, and QP problems using NVIDIA cuOpt's GPU-accelerated solver via Python, C/C++, or CLI interfaces.
2.2k · bundle
jiachen-t-wang
Nlvr2 A Visual Reasoning Benchmark For Natural Language Arxi
NLVR2: A Visual Reasoning Benchmark for Natural Language
6
seaworld008
Voice
Collecting user feedback via NPS surveys, review analysis, sentiment analysis, feedback classification, and insight extraction reports. Use when establishing feedback loops.
65 · bundle
lambenthan
Review
通用跨模型审查:Review LLM 对任意研究制品进行独立评审,输出结构化评分、wiki 实体映射与改进建议
77
qhjqhj00
Self Review
Reviews an academic paper using the NeurIPS review form with three reviewer personas, ensemble scoring, and reflection refinement. Extracts text from PDF, runs structured review, and outputs actionable feedback.
3 · bundle
qhjqhj00
T5 Eval
Benchmarks a text-to-text transformer across GLUE, SuperGLUE, CNN/Daily Mail, SQuAD, and WMT, reporting GLUE average, BLEU, ROUGE-2-F, and Exact Match scores.
3
ssrjkk
Nltk Ner
NER with Nltk. named entity recognition.
2 · bundle
brycewang-stanford
Cost Benefit
Cost-benefit analysis. Produces economic NPV and financial NPV side by side, with BCR, optimism bias (with mitigation), Marginal Excess Tax Burden, real-terms rebasing, WELLBY / QALY / VPF wellbeing valuation, sensitivity, switching values, EANC for unequal-life options, validation gate, and a one-line headline verdict (socially worthwhile vs financially self-sustaining). Backed by the greenbook R package (HM Treasury Green Book primitives) when available, with graceful fallback. Supports HMT Green Book, EU Better Regulation, World Bank, ADB, and Victorian HVHR. Reads a longlist markdown file directly via --from.
1k · bundle
mukul975
Performing Insider Threat Investigation
Investigates insider threat incidents involving employees, contractors, or trusted partners who misuse authorized access to steal data, sabotage systems, or violate security policies. Combines digital forensics, user behavior analytics, and HR/legal coordination to build an evidence-based case.
24.6k · bundle
brycewang-stanford
Auto Review Loop LLM
Autonomous research review loop using any OpenAI-compatible LLM API. Configure via llm-chat MCP server or environment variables. Trigger with "auto review loop llm" or "llm review".
1k
pranavnagrecha
Npsp Custom Rollups
Configures, troubleshoots, and extends NPSP Customizable Rollups, including rollup definitions, filter groups, batch job modes, and migration from legacy rollups.
15 · bundle
georgeqle
Eval Ideas
Loop feature-interviews over a brainstorm idea set and consolidate survivors into the roadmap
1 · bundle
seaworld008
Omen
Enumerating failure modes via pre-mortem analysis. Systematically identifies failure scenarios for plans, designs, and features, scoring them with RPN/AP. Does not write code.
65 · bundle
mukul975
Performing Privilege Escalation Assessment
Performs privilege escalation assessments on compromised Linux and Windows systems to identify paths from low-privilege access to root or SYSTEM-level control.
24.6k · bundle
ekatasingh1107
Lead Qualifier
Multi-dimensional lead qualification scoring. Evaluates leads against BANT criteria, firmographic fit, behavioral signals, and intent indicators. Outputs qualified/disqualified verdict with detailed reasoning.
2 · bundle
muratcankoylan
Advanced Evaluation
Provides production-grade techniques for evaluating LLM outputs using LLMs as judges, covering direct scoring, pairwise comparison, bias mitigation, rubric generation, and confidence calibration.
16.9k · bundle
casemark
Merit Review
Analyzes state merit review for non-covered securities offerings, applying NASAA Statements of Policy to cheap stock, promoter equity investment, voting rights, and promoter compensation. Produces examiner-ready comment responses with cap table analysis and negotiation strategy. Use when filing Reg A, Rule 504, intrastate, or direct public offerings in merit review states, responding to Blue Sky examiner comments, structuring offerings to avoid conditioning, or analyzing NASAA SOPs. Also trigger on cheap stock analysis, promoter equity tests, unequal voting rights review, state examiner correspondence, or phrases like "merit review issues" or "the state examiner sent comments."
34
fukukei23
Sentaku
選択肢(A/B/C)の深掘り比較→淘汰→推奨で判断負担を下げ判断の質を上げるスキル。5段階(L1固定3点/L1.5案拡張Diverge・自動/L2評価軸マトリクス/L3複数LLM弁証論/L4過去判断照合)。 「比較して」「深掘りして」「メリデメ教えて」「お勧めは?」「徹底的に」「過去の判断と照合」「前にどう決めたっけ」「/sentaku」等で発火。teian(浅)の深掘り要求を受け取り、brainstorming(深:設計全体)と棲み分け。
0
ssrjkk
LLM Eval
Evaluates LLM performance using BLEU, ROUGE metrics and LLM-as-judge. Use for model testing.
2 · bundle
diegosouzapw
Nlss
Runs R statistics analyses on local datasets, producing NLSS-format tables, narratives, and JSONL logs from CSV, SAV, RDS, RData, or Parquet files.
54 · bundle
gonglingrui
Novel Evaluator
严格细致判断与评分故事文本,从市场潜力、创新属性、内容亮点维度分析质量。适用于小说初筛选、多维度评估打分
349 · bundle
qhjqhj00
Cab Eval
Benchmarks LLM bias by scoring responses to automatically generated open-ended questions across sensitive attributes, producing a composite fitness score from 0 to 5.
3
whd4
Agent Evaluation
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
0
brycewang-stanford
B1
VS-Enhanced Literature Review Strategist - Comprehensive support for multiple review methodologies Full VS 5-Phase process: Prevents Mode Collapse and presents creative search strategies Supports: Systematic Review (PRISMA 2020), Scoping Review (JBI/PRISMA-ScR), Meta-Synthesis, Realist Synthesis, Narrative Review, Rapid Review Use when: conducting any type of literature review, systematic reviews, meta-analyses, scoping reviews, finding prior research Triggers: literature review, PRISMA, systematic review, scoping review, meta-synthesis, realist synthesis, narrative review, rapid review
1k