GPT-5.5 Still Misses One in Three: Reviewing Wuying-Browser-Agent
The strongest closed agent still misses a third of 37.9-step BrowserBench tasks. We check the source.
- Wuying-Browser-Agent-27B tops open-source models at 70.8% average (80.6% WebVoyager, 66.7% Online-Mind2Web, 65.1% BrowserBench), beating Qwen3.8-Max and GPT-5
- Even GPT-5.5, the strongest closed model tested, misses 32.6 points on the 37.9-step BrowserBench — a gap short benchmarks don't expose
- Online RL (DAO-GRPO) gains scale with difficulty and length: +12.4 pts on hard tasks, +13.4 pts (13.3%→26.7%) beyond 50 steps




















