Academic

Academic

Academic · 약 1분

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models

arXiv:2603.13966v1 Announce Type: new Abstract: Vision Language Action VLA models are typically evaluated using per benchmark scripts maintained independently by each model repository, leading to …

Suhwan Choi, Yunsung Lee, Yubeen Park, Chris Dongjoo Kim, Ranjay Krishna, Dieter Fox, Youngjae Yu
조회수 28회
Academic · 약 1분

The Phenomenology of Hallucinations

arXiv:2603.13911v1 Announce Type: new Abstract: We show that language models hallucinate not because they fail to detect uncertainty, but because of a failure to integrate …

Valeria Ruscio, Keiran Thompson
조회수 20회
Academic · 약 1분

LiveWeb-IE: A Benchmark For Online Web Information Extraction

arXiv:2603.13773v1 Announce Type: new Abstract: Web information extraction (WIE) is the task of automatically extracting data from web pages, offering high utility for various applications. …

Seungbin Yang, Jihwan Kim, Jaemin Choi, Dongjin Kim, Soyoung Yang, ChaeHun Park, Jaegul Choo
조회수 32회
Academic · 약 1분

LLM Routing as Reasoning: A MaxSAT View

arXiv:2603.13612v1 Announce Type: new Abstract: Routing a query through an appropriate LLM is challenging, particularly when user preferences are expressed in natural language and model …

Son Nguyen, Xinyuan Liu, Ransalu Senanayake
조회수 33회