Unknown Token Rate

Evaluates scientific document representation models on multilingual abstracts by measuring tokenization coverage, language modeling perplexity, and embedding quality relative to citation networks. It probes whether models can meaningfully process non-Latin scripts and low-resource languages without degrading to English-only or graph-based heuristics. Use when the user has predictions and gold and needs to compute unknown_token_rate.

qhjqhj00 f398bf1 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/unknown_token_rate commit f398bf1664

Frequently asked questions

npx skillmds add qhjqhj00/unknown-token-rate