Advances in models, chips, or devices will continually improve local-first capabilities. We will open our benchmarks soon.
Read the full post: https://t.co/SHYlWB56zA
We post-trained the model to purpose. PPLX 27B is trained inside the Computer harness on synthetic tasks developed from how people actually use Computer, no real user data. https://t.co/0Z16bihCdg
The local model can escalate to frontier advisor. It's user-gated, PII-flagged, text guidance only.
On Terminal Bench 2.1, escalation lifts the score from 59.6% to 73.0% at $0.415 per rollout, recovering about three-fifths of the frontier gap at two-thirds of the cost. https://t…
Parsing runs entirely on-device, so sensitive documents never leave the machine.
On ParseBench-100, Computer scores 65.1% vs 34.6% for Hermes and 13.9% for Pi, in least time with fewest tokens. https://t.co/GQ590BYEn5
Web research: on 1,266 BrowseComp tasks, Computer reaches 66.7% accuracy vs 50.2% for Pi and 43.9% for Hermes on their respective search providers.
Portable uses the least wall time and the fewest tokens. Inference and private documents stay local; only search touches the web. h…
The model and harness are designed together because small models fail in harnesses built for frontier models.
Portable is a minimal system prompt, skills that load on demand, connectors as compact CLI tools instead of MCP servers, self-verification, and an always-on sandbox. htt…
New research: Portable Computer is a local-first agent for private and cost-effective work.
With an on-device 27B model, our harness scores 82.6% on real knowledge work, beating open-source harnesses Pi and Hermes. Our post-trained PPLX 27B reaches 85.4%. https://t.co/Rb6d47clCI