Hindi LLM Benchmark Eval

Evaluates instruction-following, mathematical reasoning, code/function-calling, and retrieval-augmented generation capabilities of LLMs in Hindi. The benchmark specifically probes the models' ability to handle culturally and linguistically nuanced prompts that go beyond direct English translation. Use when the user wants to benchmark on IFEval-Hi, MT-Bench-Hi, GSM8K-Hi, ChatRAG-Hi, BFCL-Hi, or asks about evaluating this task. Reports score.

qhjqhj00 1e7cffe 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/hindi-llm-benchmark-eval commit 1e7cffe5fc

Frequently asked questions

npx skillmds add qhjqhj00/hindi-llm-benchmark-eval