
The first publicly published macOS-native computer-use benchmark. 369 task slots across 15 categories (Finder, Safari, Mail, Notes, Calendar, Reminders, Settings, Terminal, Pages, Numbers, Keynote, Music, Photos, Maps, Multi-app), agent-agnostic Go runner, dual scoring (IMPLEMENTED + STRICT), per-task PID-snapshot isolation. First reference run: kinclaw v1.15.0 + Kimi-K2.5 = 67.3% IMPLEMENTED. Documents the full 49.3 -> 62 -> 67.3 debugging trajectory as methodology contribution.Note (2026-05-09): This version bundles English + 中文 in a single PDF (English first, then Chinese), generated directly from the canonical Markdown source files.
macOS, benchmark, task isolation, AppleScript, methodology, autonomous agents, OSWorld, computer-use agents
macOS, benchmark, task isolation, AppleScript, methodology, autonomous agents, OSWorld, computer-use agents
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
