This Week in All Things AI covers key developments in models, agents, tools, infrastructure, and policy curated from discussions in the All Things AI Telegram group.
If you follow AI for work, research, investing, or just to understand where the technology is heading, this weekly brief is a concise way to scan the most important launches, risks, and resources in a few focused minutes.
During the week of May 31 to June 6, 2026, the AI landscape was defined by a wave of model releases, escalating cost concerns, and infrastructure advances. Major model launches came from MiniMax (M3), NVIDIA (Nemotron 3 Ultra), Alibaba (Qwen3.7-Plus), Google DeepMind (Gemma 4 12B), and Microsoft (seven MAI models), while Anthropic reported Claude now authors 80% of its own codebase and OpenAI updated ChatGPT's memory system. On the cost front, an unnamed firm's $500M Claude AI bill dominated headlines, spurring broader cost-control moves — Uber capped AI spending at $1,500 per employee monthly, and tools like Factory Router and Lindy's migration to DeepSeek v4 demonstrated significant savings. Infrastructure and tooling saw updates from Perplexity's "Search as Code," SambaNova's disaggregated inference, Supabase's $500M raise, OpenAI's Codex plugins, Grok Composer 2.5, and Scaledown.ai launch of small-language-model services for compression and summarization
The sections that follow walk through these items day by day, with short context and links so you can dive deeper into the pieces most relevant to your work or interests.
Unnamed firm racks up $500M Claude AI bill in one month
An unnamed enterprise client spent more than $500 million in just 30 days on Anthropic's Claude AI platform after failing to set any usage limits or spending caps for employees, according to an Axios investigation published on May 28. The disclosure, made by an AI consultant whose client generated the bill, may represent one of the costliest IT governance failures on record
Xiaomi just published a deep dive into the end-to-end inference engineering behind MiMo-V2.5 and MiMo-V2.5-Pro — and it’s a strong example of how architecture + systems work together to unlock real production gains.
MiniMax Launches M3 AI Model for Coding and Agents
Chinese AI lab MiniMax released its M3 model excelling in coding, agentic reasoning, tool use, multimodal chat, and long-context tasks. OpenCode integrated it instantly for free early access, offering 1,400 requests every five hours via terminal, IDE, or desktop app on macOS, Windows, and Linux. It posts strong benchmark results including 59.0% on SWE-Bench Pro, 66.0% on Terminal Bench 2.1, 74.2% on MCP Atlas, and competitive scores against models like GPT 5.5 and Gemini 3.1 Pro across agentic and tool-use tasks. Weights and technical report expected in about 10 days; API is immediately available with 50% off standard usage for the first 7 days to encourage early testing.
NVIDIA Unveils Nemotron 3 Ultra
The model tops U.S. open-weights rankings with an Intelligence Index score of 48, beating rivals like Gemma 4 31B while delivering over 300 output tokens per second in tests. It shines in agentic AI tasks, boasting 91% agent productivity, strong instruction following, and top coding skills, with 5x faster inference and 30% lower costs than leading competitors. Huang highlighted its openness for developers building everything from search tools to molecular simulations, paired with new agent deployment tools, as Nemotron 4 looms on the horizon.
What I was able to one-shot with Minimax-M3.
SuperGrok and X Premium+ users now can use Composer 2.5 model from Cursor via Grok Build
Qwen3.7-Plus is Alibaba's new multimodal agent model unifying vision and language for hybrid GUI/CLI operations, visual perception, reasoning, grounding, coding, and search-augmented tasks.
Benchmarks show it delivers competitive results across text, multimodal reasoning, visual understanding, agentic coding, and real-world QA, often matching or exceeding models like Claude Opus-4.6, GPT-5.4, and Gemini-3.1-Pro.
Released via API on Alibaba Cloud Model Studio with demos of browser and interactive agents;
Some discussion between me and Anya Shapina related to Composer2-5 model availablility via Grok
=====
Anya
IMO Grok is the most under-appreciated model. I know, my comment is out of context (I haven't used Grok Build enough). But just as a model, it roxx! Testing very complicated agents and workflows across models and it kinda wins most scenarios. In my experience, it excels at walking and chewing gum at the same time (while also looking for shadows lurking behind bushes) - this is where most models fail for me. Not to mention native search, X, and Reddit. Anyone else in my Grok camp? Sorry, Elon haters :)
Yusuf
Which model of Grok Anya ? I think grok-build-0.1 is highy underrated and really look forward to the release of v9. Elon has a tendency to hype up particularly given the SpaceX IPO but the pace of updates of Grok Build CLI is relentless
Anya
Grok 4.3 - in my specific tests powering Agents (Not Grok Build which you post was about; look fwd to reporting on this later). Grok 4.3 - Strongest agentic tool calling + lowest hallucination rate in the lineup, from what I know. Worked for me.
=====
Perplexity AI launched Search as Code, enabling AI agents to generate Python code that directly calls atomic search primitives like fanout queries, deduplication, filtering, and ranking within a secure sandbox.
Available in the Perplexity Agent API, and now default in Computer.
Uber says it has limited all employees to $1,500 in monthly token spending per AI coding tool “to responsibly encourage agentic AI adoption”
Microsoft Launches MAI AI Models at Build 2026
The lineup covers reasoning, coding, images, voice, and transcription, all built as a multimodal system optimized for Microsoft's efficient MAIA 200 chips. MAI-Thinking-1, a 35 billion active-parameter model trained on 30 trillion tokens, scores 97% on math benchmark AIME 2025 and tops tough coding tests like SWE-Bench Pro, even beating models from Anthropic and matching Claude in blind evaluation
Microsoft AI CEO Mustafa Suleyman calls it a step toward 'humanist superintelligence' under human control.
OpenAI unveils new Codex plugins for tasks related to public equity investment, banking and sales, and other roles, and plans to integrate Codex into ChatGPT
Factory Router automatically selects the best AI model for each task in the Droid coding agent, slashing costs by 20-25% while matching 99% of premium performance on benchmarks like Terminal-Bench 2. It routes routine edits to cheaper models and escalates complex work to heavy-hitters like Claude Opus 4.7, even switching providers mid-task for reliability. Tech leaders like Keith Rabois and Garry Tan praised it as essential for enterprises facing rising AI bills, with CEO Matan Grinberg comparing it to hiring the right expert instead of Einstein for basic math. Currently in private preview for CLI and desktop users.
FWIW, Droid from Factory is one of my goto model-agnostic harness alongside Opencode Go. Very good Claude Code compatability in terms of support for skills, plugins, hooks, sub-agents etc and one of the best if not the best context management I've come across
via James Chan
Does anyone have experience with multiplayer PRD like ChatPRD? Any good?
Google DeepMind releases Gemma 4 12B, a lightweight 12B-parameter multimodal model under Apache 2.0 license designed to run locally on laptops with just 16GB VRAM or unified memory.
It introduces an encoder-free unified architecture where vision uses a tiny 35M-parameter embedding module and audio projects raw signals directly into the LLM backbone, eliminating separate encoders for efficiency.
Delivers advanced reasoning and multimodal capabilities nearing the larger Gemma 4 26B model's benchmarks, with native 256K context, MTP drafters, and immediate support across Hugging Face, Kaggle, llama.cpp, MLX, and vLLM.
via my high-school classmate Abhi Ingle who works at Sambanova and lurks on this group via the newsletter. Took me a few readings of the blog post and had to go back and refresh my knowledge on the prefill and decode steps in auto reggressive transformer architectures before the penny dropped on the how this architecture is unique
====
At Computex, SambaNova demoed a “disaggregated inference” architecture where GPUs handle prefill and their RDUs handle decode, yielding faster, cheaper long-horizon agent workloads than GPU-only setups, and they now have this running in production-style environments with partners like VC2 and Together AI.
OpenAI updates ChatGPT memory with a “more capable and compute-efficient” architecture and a summary page that lets users review and steer what it remembers
Supabase, which provides backend tools for building AI apps, raised a $500M Series F led by GIC at a $10B pre-money valuation, up from $5B in October 2025. Also released Multigres v0.1 alpha to the open source community, Multigres tries to bring Vitess-grade horizontal scaling, high availability, and operational simplicity to Postgres.
Response from Anya Shapina
===
I believe Supabase (and possibly its philosophical rival Convex) will be some of the biggest winners in the AI race. I use it as the shared memory and persistent backend of my multi-agent system - it's everything to my project, and all other projects I've ever built. But then there is the Goliath Convex with its built-in support for AI Agents... and NO SQL (love/hate?). Which camp are you people in?
====
I've not used Lindy but have heard good things about it from others. For those unaware of Lindy, it describs itself productivity and automation platform that positions itself as an "AI executive assistant" or "AI employee." It is particularly strong for professionals who want proactive help with email, meetings, and calendar management.
Saw this post from Flo Crivello, founder of Lindy who migrated 100% of its traffic to DeepSeek v4 from Anthropic models, reporting millions in cost savings alongside performance improvements on core use cases. The switch required building substantial new infrastructure and internal tooling, described as 100x more work than anticipated, emphasizing the value of swappable model architectures.
Response from Ms Macarena Correa to a post dated 30th May 2026 about the announcement of Kirkland and Ellis committing $500 million over the next 3-4 years to build its own proprietary AI platform and custom tools
=====
On this one, I think its the right way forward. If all Magic Circle / Silver Circle law firms are using the same AI tools to help drafting contracts (Harvey / Legora), and review them too, as well as providing advice, all solutions will become kind of standard, so ultimately the difference between law firms will be reduce only to 1) pricing and 2) charisma of each particular lawyer (e.g., building the relation with the client). this 2nd point cannot be replaced by AI, so probably most of the efforts need to shit into building more and better relations of trust.
=====
Anthropic details its progress toward recursive self-improvement, and its implications, and says Claude has authored 80%+ of the code merged into its codebase
via Imran Muthuvappa
made a free community for people who want to build and ship their first agent
Came across Scaledown which has developed purpose-built models for compression, summarization, extraction, and classification. Pricing is 0.05/M tokens with self-deployment options also available
Below is my personal website which aggregates links to many of my socials as well as the various content and community that I curate. Feel free to share this link to others who you think may find this content/community useful to them
The cover image of this newsletter via generated via the Krea 2 Large model within the Krea tool via the following prompt
An oil painting in the style of H.R. Giger and Zdzisław Beksiński of an ancient Greek temple on the side of Mount Olympus, surrounded by cacti, a rainbow in the sky, grey clouds overhead, blue lilies around the base, and birds flying above. The painting is highly detailed.







Multimodal interactive hybrid agent: unified GUI & CLI operation across visual and text tasks

















