Source

$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

peer-reviewed-track preprint Yao, Shinn, Razavi, Narasimhan, arXiv:2406.12045, 2024

Claim this source supports

“State-of-the-art agents succeed on under half of realistic multi-step tasks and are inconsistent across repeated trials”

View original source

Cited in