Agents
Mr.LHDR benchmark: the best deep-research agent gets 43.1% of final answers right but only 34.3% with a correct intermediate chain
arXiv 2609.11318 builds multimodal research questions from hidden node-relation graphs. Each needs an average of 12.1 intermediate conclusions at a dependency depth of 10.4, and each includes at least one image, map, PDF, chart or video frame. The strongest system scored 43.1% overall accuracy and 34.3% strict accuracy, so final-answer grading overstates how often the research was actually complete. Removing images cost 12.6 points on the dependency-aware checklist score. Grading only final answers hides broken evidence chains, so deep-research evals should score intermediate conclusions too.
Source
↳ Follow the thread