OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interaction
arXiv:2602.15197v1 Announce Type: new Abstract: Tool-calling is essential for Large Language Model (LLM) agents to complete real-world tasks. While most existing benchmarks assume simple, perfectly …