Multinrc Eval

This benchmark evaluates LLMs' ability to perform multi-step reasoning in native non-English languages (French, Spanish, Chinese) across linguistic, wordplay, cultural/tradition, and culturally-grounded math categories. It specifically probes whether models rely on translation bias or possess deep cultural and linguistic contextual knowledge required for accurate problem-solving. Use when the user wants to benchmark on MultiNRC, or asks about evaluating this task. Reports accuracy.

qhjqhj00 748b347 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multinrc-eval commit 748b347cd6

Frequently asked questions

npx skillmds add qhjqhj00/multinrc-eval