Reply with commands, one per line:
/approve all · /approve 1,3-5 · /reject 2 · /edit 3 category=world_terminal tags=benchmark (edit implies approve; tags=- clears tags). Valid category keys: see taxonomy.json.
1. E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
Weihuang Zheng, Tianyuan Zou, Eileen Ye, et al. · arXiv 2026/07 · paper
proposed: world_apps · tags: benchmark
E-Bench is a synthetic benchmark of 323 state-changing, multi-step tool-use tasks across three real product domains (Honor of Kings, QQ Music, Tencent Meeting), decoupling environment and task synthesis for controllable, scalable evaluation. It shows current LLM agents struggle to reliably discover hidden information and compose tool calls to change state.
Agents complete tasks by operating real product/enterprise software via tool calls, matching world_apps.
machine payload (do not edit)
[{"paper": {"id": "2607.23722", "title": "E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios", "authors": ["Weihuang Zheng", "Tianyuan Zou", "Eileen Ye", "Alphet Liu", "Youyong Kong", "Ya-Qin Zhang", "Duran Zheng", "Maxm Pan"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-26", "links": {"paper": "https://arxiv.org/abs/2607.23722", "github": "", "website": ""}}, "category": "world_apps", "tags": ["benchmark"], "summary": "E-Bench is a synthetic benchmark of 323 state-changing, multi-step tool-use tasks across three real product domains (Honor of Kings, QQ Music, Tencent Meeting), decoupling environment and task synthesis for controllable, scalable evaluation. It shows current LLM agents struggle to reliably discover hidden information and compose tool calls to change state.", "reason": "Agents complete tasks by operating real product/enterprise software via tool calls, matching world_apps.", "source": "crawl"}]
Reply with commands, one per line:
/approve all·/approve 1,3-5·/reject 2·/edit 3 category=world_terminal tags=benchmark(edit implies approve;tags=-clears tags). Valid category keys: see taxonomy.json.1. E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
Weihuang Zheng, Tianyuan Zou, Eileen Ye, et al. · arXiv 2026/07 · paper
proposed:
world_apps· tags:benchmarkmachine payload (do not edit)
[{"paper": {"id": "2607.23722", "title": "E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios", "authors": ["Weihuang Zheng", "Tianyuan Zou", "Eileen Ye", "Alphet Liu", "Youyong Kong", "Ya-Qin Zhang", "Duran Zheng", "Maxm Pan"], "venue": "arXiv 2026/07", "category": "", "published": "2026-07-26", "links": {"paper": "https://arxiv.org/abs/2607.23722", "github": "", "website": ""}}, "category": "world_apps", "tags": ["benchmark"], "summary": "E-Bench is a synthetic benchmark of 323 state-changing, multi-step tool-use tasks across three real product domains (Honor of Kings, QQ Music, Tencent Meeting), decoupling environment and task synthesis for controllable, scalable evaluation. It shows current LLM agents struggle to reliably discover hidden information and compose tool calls to change state.", "reason": "Agents complete tasks by operating real product/enterprise software via tool calls, matching world_apps.", "source": "crawl"}]