Frontier LLMs drop from 83% to 43% once reasoning has to chain across domains

https://arxiv.org/abs/2607.18438

Comments

westurnerJul 26, 2026, 10:14 PM
Is that the model, all models, the agent?

Does something like this make a difference?

"Schema Harness Achieves ~99% on Arc‑AGI‑3 Public" (2026-07) https://news.ycombinator.com/item?id=48935905