분류 전체보기421 [2026-1] 강민정, 염제원 - GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks Paperhttps://arxiv.org/abs/2510.04374 GDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksWe introduce GDPval, a benchmark evaluating AI model capabilities on real-world economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDParxiv.orgArticlehttps://.. 2026. 3. 20. [2026-1] 김다정, 황징아이 - TaU2-Benchmark 1. 데이터셋의 구성 의의$\tau^2$-Benchmark는 대화형 에이전트와 시뮬레이션된 사용자 사이에서 이루어지는 multi-turn interaction을 체계적으로 연구하기 위해 제안된 밴치마크이다. 기존 Single-Control 벤치마크의 한계기존의 대화형 AI 에이전트 벤치마크는 대부분 single-control 환경을 가정한다. 에이전트만이 도구(tool)를 사용하여 환경과 상호작용하고, 사용자는 단순히 정보나 선호도를 제공하는 역할에 그친다. 하지만 이러한 설정은 실제 상황의 복잡성을 충분히 반영하지 못한다는 한계가 있다. 사용자와 에이전트의 협업 필요성 (Dual-Control의 도입 배경)실생활에서는 에이전트와 사용자가 함께 문제를 해결하는 협업 상황이 자주 발생한다.예를 들어 Te.. 2026. 3. 18. [2026-1] 백승우 - OpenClaw-RL: Train Any Agent Simply by Talking OpenClaw-RL: Train Any Agent Simply by TalkingEvery agent interaction generates a next-state signal, namely the user reply, tool output, terminal or GUI state change that follows each action, yet no existing agentic RL system recovers it as a live, online learning source. We present OpenClaw-RL, a fraarxiv.org 2026. 3. 17. [2026-1] 백승우 - AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines https://arxiv.org/abs/2602.14296 AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State MachinesThe performance of autonomous Web GUI agents heavily relies on the quality and quantity of their training data. However, a fundamental bottleneck persists: collecting interaction trajectories from real-world websites is expensive and difficult to verify. Tarxiv.org 2026. 3. 10. 이전 1 ··· 4 5 6 7 8 9 10 ··· 106 다음