67Stories Tracked
99Peak Impact
6724h Fresh Stories
14+Linked Entities

ChatGPT Work Computers Can Now Sign In to Websites Without Exposing User Credentials to the Model

ChatGPT Work's cloud computer environment can now sign in to websites on web and mobile for Plus, Pro, and Business users, storing credentials inside isolated browser forms without exposing passwords to the model.

Key Takeaways

  • Authenticated computer-use unlocks high-value workflows behind logins that were previously unreachable;
  • Secure browser credential isolation ensures the underlying LLM never sees or ingests plaintext passwords;
  • Simultaneous web and mobile rollout positions ChatGPT Work as an autonomous workplace operator.
Linked Models & Agents:

OpenAI Publishes First JalapeΓ±o Custom Inference Chip Results: Up to 1.9x Work per Watt and 3.6x Lower Latency

OpenAI released InferenceX benchmarks for JalapeΓ±o, its first custom inference silicon, showing 1.5–1.9x higher work per watt and 1.7–3.6x lower end-to-end latency across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5.

Key Takeaways

  • OpenAI officially expands into first-party silicon, publishing public model-agnostic benchmarks across open weights;
  • Designed around agentic serving, unified architecture handles both prefill and decode stages simultaneously;
  • AI-assisted chip development achieved tapeout in 9 months and enabled rapid Codex-driven model porting in 2 months.

vLLM 0.28.0 Released: Full-Stack Kimi-K3, DeepSeek-V4 Sparse MLA, and Tiered Disk KV Cache

vLLM 0.28.0 landed with 584 commits, bringing stack-wide Kimi-K3 support, end-to-end DeepSeek-V4 sparse MLA with DFlash2/DSpark speculative decoding, Model Runner V2 disaggregation, and tiered disk KV cache.

Key Takeaways

  • Adaptive speculative token budgets slash end-to-end TTFT by 55–65% with 1.5–3x faster sequence-parallel kernels;
  • End-to-end DeepSeek-V4 sparse MLA and Kimi-K3 shared-expert sharding save ~17 GiB VRAM per GPU;
  • Tiered KV cache introduces a disk tier, doubles max batched tokens to 16,384, and supports diverse hardware backends.

Google DeepMind Launches Gemini 3.7 Flash as Workhorse Model for Coding and Agents

Google DeepMind released Gemini 3.7 Flash, delivering massive gains in agentic coding and debugging (FrontierCode 43.6%, DeepSWE 65.3%) with Day-0 availability in Antigravity and AI Studio at introductory rates.

Key Takeaways

  • Substantial accuracy gains over 3.6 Flash on complex debugging, issue resolution, and single-shot UI layouts;
  • Benchmark scores: FrontierCode 1.1 jumps to 43.6% (vs 34.4%) and DeepSWE v1.1 to 65.3% (vs 49.0%);
  • Day-0 integration in Antigravity, AI Studio, and Android Studio at promotional $0.75 / $3.75 per 1M rates.

Alibaba Qwen Open-Weights Qwen3.8-Flash: 125B MoE Previewing Qwen4 Architecture

Alibaba Qwen released Qwen3.8-Flash, an open-weights 125B MoE model (6B active) previewing the Qwen4 architecture with 1/9 training cost, 262K native context, and ultra-cheap API pricing.

Key Takeaways

  • New core architecture combining GDN gated residuals, sparse attention, N-gram embeddings, and Muon optimizer;
  • Coding & SWE benchmarks: 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, and 73.9 on CoWorkBench;
  • Day-0 support across QwenCloud, OpenRouter, vLLM, SGLang, and Unsloth (runnable locally with ~75GB RAM via GGUF).

DeepSeek Open-Sources Multi-Token Prediction (MTP) Agent Harness Framework: 300% Multi-File Generation Speedup

DeepSeek open-sources its Multi-Token Prediction (MTP) agent acceleration framework. By speculatively predicting syntax blocks in parallel, multi-file code diff and test generation in Cline and Aider achieves over 3x speedup with zero degradation in AST validity.

Key Takeaways

  • DeepSeek releases Multi-Token Prediction (MTP) framework tailored for coding agents;
  • Delivers 300% throughput acceleration during multi-file refactor and test suites;
  • Directly compatible with local vLLM and Ollama deployments.

OpenRouter Reveals Ox Alpha as GLM-5.3-Flash: Native Multimodal Model Surpasses 20T Tokens in 6 Days

OpenRouter confirmed the viral mystery Ox Alpha model is Zhipu's GLM-5.3-Flash, processing over 20 trillion tokens in six days as the fastest-growing model in router history, now live with 1M context.

Key Takeaways

  • Processed 20T tokens in six days, setting an all-time adoption record for new models on OpenRouter;
  • Hybrid sparse-linear attention delivers ultra-cheap long-horizon coding and agent inference across 1M context;
  • Promotional launch rates set at $0.075 / $0.25 / $0.015 per million tokens (input / output / cache) through Sept 9.

xAI Launches Grok 4.6 for Autonomous Coding Agents at $2 / $6 per Million Tokens

xAI announced Grok 4.6, delivering major speed and reasoning enhancements tailored for long-running coding agents, with Day-0 availability in Grok Build, Cursor, and the API at $2 / $6 per million tokens.

Key Takeaways

  • Deeply optimized via RL for multi-file refactoring, dependency tracking, and autonomous agent loops;
  • Priced at $2/M input and $6/M output, keeping Grok 4.5 parity while boosting reasoning and speed;
  • Available Day-0 in Grok Build, Cursor, and the API with 2x usage allowance during launch week.
Linked Models & Agents:

Perplexity Launches Portable Computer: Local-First Autonomous Agent Stack on NVIDIA DGX Spark

Perplexity introduced Portable Computer on NVIDIA DGX Spark, providing a fully local-first agent stack powered by on-device PPLX 27B where private documents never leave the machine by default.

Key Takeaways

  • Commercializes local-first agents: private documents stay on-device unless explicitly approved for cloud calls;
  • Lightweight harness custom-engineered for 27B models with compact CLI connectors and on-device sandboxing;
  • Hybrid escalation allows optional pay-per-task cloud frontier calls (~$0.42) lifting Terminal Bench to 73.0%.

DeepSeek-V4-Flash-Vision-Exp Goes Live on API with Near-Opus 4.8 Multimodal Agent Capabilities

DeepSeek launched deepseek-v4-flash-vision-exp on its API, maintaining ultra-low V4-Flash pricing while achieving multimodal coding and agent performance approaching Claude Opus 4.8.

Key Takeaways

  • Preserves ultra-cheap V4-Flash token rates for both text and vision multimodal inputs;
  • Achieves massive leaps on multimodal agent benchmarks, rivaling Claude Opus 4.8 on GUI and diagram reasoning;
  • Day-0 support added in DeepSeek Harness 0.1.1 for seamless agent orchestration and benchmarking.
Page 1 of 7(67 stories in total)