narrativeqa-eval
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation — Chen et al. (2024) (arXiv:2402.03216, 2024)
What this evaluates
Evaluates English long-document retrieval on complex, narrative-style questions, probing deep comprehension and information extraction from lengthy texts.
Datasets
- NarrativeQA — total ?; splits: test (-1)
Metrics
nDCG@10(primary) — range: [0, 1]- Normalized Discounted Cumulative Gain at rank 10, assessing the ranking quality of retrieved narrative documents.
Input / output format
Input: Complex questions and corresponding English narrative documents.
Output: A ranked list of retrieved documents.
Scoring recipe
def compute_ndcg_at_10(retrieved_ids, relevant_ids):
dcg = 0.0
for i, doc_id in enumerate(retrieved_ids[:10]):
if doc_id in relevant_ids:
dcg += 1.0 / math.log2(i + 2)
idcg = sum(1.0 / math.log2(i + 2) for i in range(min(len(relevant_ids), 10)))
return dcg / idcg if idcg > 0 else 0.0
Common pitfalls
- Performance advantage over baselines grows with sequence length, indicating sensitivity to input context window.
- Only evaluates English documents, unlike the other benchmarks in the paper.
Evidence (verbatim from paper)
We make further analysis with NarrativeQA (Table[4]), where we can make a similar observation as MLDR. Besides, with the growth of sequence length, our method gradually expands its advantage over baseline methods (Figure [5]), which reflects its proficiency in handling long inputs. Table 4: Evaluation on NarrativeQA (nDCG@10).
Citation
@misc{chen2024m3embedding,
title={M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation},
author={Chen et al. (2024)},
year={2024},
note={arXiv:2402.03216}
}
- arXiv: 2402.03216