Blog
Every public run, reported in full
One post per public benchmark run. These are the same reports the leaderboard is built from — manual-audit verdicts for all 60 tasks, cost and token breakdowns, per-app battery drain, device thermals, and the hallucination-control results. Published unedited, including the runs that went badly.
-
2026-09-16 01:13 → ~06:19 IST (2026-09-15 19:43 → 2026-09-16 00:34 UTC) — **interrupted by phone battery / ADB death**
Public 3-Day Sample — 60-Task Run Report (gemma-4-E2B-it, TEXT) — INTERRUPTED
`gemma-4-E2B-it` (local llama-server `@127.0.0.1:8088`) — **TEXT** (`--no-tracing`, `--temperature 0.0`)
**⚠️ Partial run.** Battery died at **~0–1%** during `hard__google-meet-files__070` (day3); wireless ADB went **offline** before `easy__messages__010` preflight. **53 tasks finalized** (day1 20/20 + day2 20/20 + day3 13/20), **2 orphans** (…
Read the report → -
2026-09-14 06:18 → 2026-09-15 ~03:11 local IST — **interrupted by phone battery / ADB death** (three tasks re-run 16 Sep)
Public 3-Day Sample — 60-Task Run Report (Qwen3.5-4B, TEXT) — INTERRUPTED
[`Qwen3.5-4B`](https://huggingface.co/Qwen/Qwen3.5-4B) (local llama-server `@127.0.0.1:8088`) — **TEXT** (`--reasoning off`, `--no-tracing`)
**⚠️ Partial run.** Battery / wireless ADB died mid–Day 2. **27 tasks finalized** (day1 20/20 + day2 7/20), **2 orphans** (`hard__bookmyshow__005`, `easy__settings__014`), **31 never started** (11 remaining day2 + all 20 day3). Day 3 has **…
Read the report → -
2026-09-10 04:15 → ~13:04 local IST (**8.59 h** wall / **8.43 h** agent time)
Public 3-Day Sample — 60-Task Run Report (openai/gpt-5.6-luna, VISION)
`openai/gpt-5.6-luna` (OpenRouter) — **VISION mode** (screenshot-driven)
**Weak VISION run — weaker than Luna TEXT (30%).** 60/60 finalized, avg **44.8 steps/task**, ~**$5.72**. Only **3** agent self-successes (calendar conflicts, Camera VIDEO, outbound call) — all verified. **7/7 HC honest fails** → PASS. **0 h…
Read the report → -
2026-09-09 04:34 → 2026-09-10 01:35 local IST (resume after pause; **≈7.27 h** summed wall / **7.10 h** agent time)
Public 3-Day Sample — 60-Task Run Report (qwen/qwen3.8-27b, VISION)
`qwen/qwen3.8-27b` (OpenRouter) — **VISION mode** (screenshot-driven; no a11y tree)
**Solid mid-pack VISION run.** 60/60 finalized, avg **26.3 steps/task**, ~**$5.07**. Official self-report after HC/ASK gates: **36/60 (60.0%)**. Deep audit keeps the same headline **36 PASS** but reclassifies **`easy__calendar__008` as HALL…
Read the report → -
2026-09-06 06:33 → 2026-09-06 ~13:30 local IST (≈6.74 h wall / 6.58 h agent time)
Public 3-Day Sample — 60-Task Run Report (openai/gpt-5.6-luna, TEXT)
`openai/gpt-5.6-luna` (OpenRouter) — **TEXT mode** (a11y-tree-driven)
**Weak TEXT run.** 60/60 finalized, avg **41.8 steps/task**, ~$3.84. Luna occasionally completes single-app reads (Camera video mode, invoice ₹1,240, Docs copy, BookMyShow, YouTube resume) but burns most budgets on **step-cap-60** launcher …
Read the report → -
2026-09-05 05:21 → 2026-09-05 15:18 local IST (≈4.67 h wall / 4.51 h agent time)
Public 3-Day Sample — 60-Task Run Report (bytedance-seed/seed-2.0-lite, VISION)
`bytedance-seed/seed-2.0-lite` (OpenRouter) — **VISION mode** (screenshot-driven)
**Vision sibling of the 08-30 text seed run (51.7%).** 60/60 finalized, avg **19.7 steps/task**, ~$3.75. Stronger HC honesty than text seed (only **1** fabricated control; none destructive) and solid easy-bucket accuracy — but the deep audi…
Read the report → -
2026-08-31 21:58 → 2026-09-01 03:09 local IST (≈7.55 h wall / 7.39 h agent time)
Public 3-Day Sample — 60-Task Run Report (xiaomi/mimo-v2.5-pro, TEXT)
`xiaomi/mimo-v2.5-pro` (OpenRouter) — **TEXT mode** (a11y-tree-driven)
**DIAGNOSTIC run.** mimo-v2.5-pro drives single-app deterministic work competently (many clean read-and-report PASSes, strong a11y-tree reading, several ADB-verified end-states) — but fails **every Telegram-message deliverable (0/5, Send-bu…
Read the report → -
2026-08-26 10:52 → 2026-08-26 12:15 local IST (≈1.72 h wall / 1.55 h agent time)
Public 3-Day Sample — 60-Task Run Report (gemini-3.1-flash-lite)
`google/gemini-3.1-flash-lite` (OpenRouter)
**⚠️ Swiggy rerun (2026-08-28) — merged in place, no verdict change:** the two Swiggy tasks (`hard__swiggy__005`, `easy__swiggy__001`) were re-run on 2026-08-28 on a freshly reset phone (same model, updated "last three months" prompt). Both…
Read the report → -
2026-08-30 14:35 → 2026-08-30 18:00 local IST (≈2.93 h wall / 2.77 h agent time)
Public 3-Day Sample — 60-Task Run Report (bytedance-seed/seed-2.0-lite, TEXT)
`bytedance-seed/seed-2.0-lite` (OpenRouter) — **TEXT mode** (a11y-tree-driven)
**Cleanest run so far** — 60/60 finalized (no orphans), avg **13.8 steps/task**, and the lowest cost yet ($2.06). The model is decisive and completes read-and-report tasks well, but the manual audit found **8 false passes** + **3 hallucinat…
Read the report → -
2026-08-30 02:18 → 2026-08-30 11:12 local IST (≈8.9 h wall) — **interrupted by phone battery death**
Public 3-Day Sample — 60-Task Run Report (moonshotai/kimi-k2.6, VISION-ONLY) — INTERRUPTED
`moonshotai/kimi-k2.6` (OpenRouter) — **VISION-ONLY mode** (`--vision-only`: screenshots only, NO accessibility tree)
**⚠️ Run interrupted — battery died mid-Day-2.** This is a **partial run**, not a full 60-task run. The OnePlus battery died at task `easy-google-meet-004` (day2), the batch wedged on a post-task ADB call, and the run was killed. Result: **…
Read the report → -
2026-08-29 15:36 → 2026-08-29 22:44 local IST (≈6.41 h wall / 6.25 h agent time)
Public 3-Day Sample — 60-Task Run Report (moonshotai/kimi-k2.6, TEXT)
`moonshotai/kimi-k2.6` (OpenRouter) — **TEXT mode** (no `--vision`; a11y-tree-driven)
**⚠️ API-key expiry — mid-run interruption + in-place resume:** the OpenRouter key expired part-way through Day 3 (`401 API key expired` at `medium-calculator-001`). The run was cancelled, the key refreshed, and the remaining tasks were re-…
Read the report → -
2026-08-28 00:24 → ~09:29 local IST (6.45 h wall / 6.29 h agent time, no gaps)
Public 3-Day Sample — 60-Task Run Report (qwen3.8-27b, TEXT)
`qwen/qwen3.8-27b` (OpenRouter) — **TEXT mode** (no `--vision-only`; a11y tree, no screenshots)
**⚠️ Music-Obsidian rerun (2026-08-29) — merged in place, no verdict change:** `hard__music-obsidian__077` was re-run on 2026-08-29 (freshly reset phone, redesigned "music app I used the most lately … stops by itself around my asleep time" …
Read the report → -
2026-08-26 18:49 → 2026-08-27 ~13:00 local IST (≈8.87 h wall / 8.71 h agent time, incl. the 5-task resume on 2026-08-27)
Public 3-Day Sample — 60-Task Run Report (qwen3.8-27b, VISION-ONLY)
`qwen/qwen3.8-27b` (OpenRouter) — **vision-only** (`--vision-only`, screenshots, no a11y tree)
**⚠️ Swiggy rerun (2026-08-28):** the two Swiggy tasks (`hard__swiggy__005`, `easy__swiggy__001`) were **re-run on 2026-08-28** on a freshly reset phone with the updated `easy__swiggy__001` prompt ("last **three** months", `public.md` 08-27…
Read the report →