Agents that finish multi-hour work unsupervised
The bottleneck moves from model quality to reliability over long horizons: checkpointing, verification and rollback rather than raw intelligence.
First visible signal · Pricing that shifts from per-token to per-completed-task
Strongest case against · Long-horizon reliability may not be an engineering problem at all. If error compounds per step, a 99% step accuracy still fails most hundred-step tasks, and no amount of checkpointing fixes a base rate.
Watch pricing model changes