Smoldocling Doc Eval

This evaluation probes a vision-language model's ability to perform end-to-end document conversion, including text recognition, layout analysis, table and chart structure extraction, and code/formula parsing. It measures how accurately the model reconstructs document content and spatial structure from page images into standardized markup formats. Use when the user wants to benchmark on DocLayNet, SynthCodeNet, Im2Latex-230k, FinTabNet, PubTables-1M, or asks about evaluating this task. Reports mAP@0.5:0.95, TEDS.

qhjqhj00 dbc999c 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/smoldocling-doc-eval commit dbc999c5f4

Frequently asked questions

npx skillmds add qhjqhj00/smoldocling-doc-eval