91.4% success across 116 AndroidWorld tasks. Open source, GPT-5, direct accessibility API. Reproducible from the eval harness in our public repo.
Highest reported score on AndroidWorld across system apps, third-party apps, and multi-step workflows.
Calendar, Files, Notes, Markor, Audio, Browser, System, Expense, Recipe, Clock, and more.
Dynamic feedback loops between planning and execution. Robust under long-horizon failure modes.
AndroidWorld success rates across publicly reported mobile agents. Mobilerun leads by ~7 points. Accessibility API, not screenshots.
Scores are pulled from their public reports and the AndroidWorld leaderboard. The eval harness is in our public repo so anyone can reproduce. Updated 29.10.2025
Smaller payloads, richer semantics, faster steps.
| Per-step input | Screenshot | Accessibility tree | Difference |
|---|---|---|---|
| Payload size | ~1,024 KB | ~2 KB | 500× smaller |
| Semantic structure | Pixels only | Labels & roles | Higher accuracy |
| Step latency | High | Low | Faster agents |
| Vision fallback | n/a | When tree is incomplete | Robust |
Successful categories from the 106 / 116 run.
Event creation, navigation, recurring schedules.
Settings, intents, system-level workflows.
Note editing, search, formatting.
Multi-step content workflows with state.
Form completion, calculation, lookups.
Across the remaining task families.
Open any task to replay its full execution trajectory.
The eval harness is open source. Reproduce the result, or build your own benchmark on the same primitives.