LLM Search Agent Eval

Evaluates the end-to-end latency and answer accuracy of LLM-based search agents operating in a ReAct workflow with external Wikipedia API calls. It probes the system's ability to balance speculative action execution with verification to reduce inference time while maintaining multi-hop reasoning quality. Use when the user wants to benchmark on HotPotQA, 2WikiMultihopQA, TriviaQA, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 7eb5d99 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/llm-search-agent-eval commit 7eb5d99bfe

Frequently asked questions

npx skillmds add qhjqhj00/llm-search-agent-eval