Recent Papers / arXiv:2605.19099

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

arXiv:2605.19099Submitted May 20, 20260 benchmark results

Authors pending

View PDF ↗arXiv page ↗Edit

Abstract

Emergent delegation evaluation across GAIA, BFCL, and tau-bench

Tasks

Results

No benchmark results recorded yet.

Benchmark results referencing this paper haven't been added to the registry yet. If you have a reproduction, submit it →

CodeSOTA extraction

Benchmark evidence

Verify that DecisionBench routing fidelity-at-1 ranges from 7.5% to 29.5% across conditions.

Add or update benchmark results

Logged-in editor · benchmark trail

Read next

Three places to go from here.

All tracked papers in the registry, with benchmark result, model, and leaderboard linkage where available.

Papers with Code is dead — alternatives

What replaced PWC for each use case: LLMs, OCR, speech, vision, robotics.

Every benchmark in Agentic AI.