# This Week in All Things AI - Week 23-2026

*Sunday 31st May 2026 to Saturday 6th June 2026*

By [This Week in All Things AI](https://paragraph.com/@twiata) · 2026-06-07

---

This Week in All Things AI covers key developments in models, agents, tools, infrastructure, and policy curated from discussions in the [**All Things AI Telegram group.**](https://t.me/+PnCnwhgH8V4yMTFl)  
  
If you follow AI for work, research, investing, or just to understand where the technology is heading, this [**weekly brief**](https://paragraph.com/@twiata) is a concise way to scan the most important launches, risks, and resources in a few focused minutes.

**During the week of May 31 to June 6, 2026, the AI landscape was defined by a wave of model releases, escalating cost concerns, and infrastructure advances.** Major model launches came from MiniMax (M3), NVIDIA (Nemotron 3 Ultra), Alibaba (Qwen3.7-Plus), Google DeepMind (Gemma 4 12B), and Microsoft (seven MAI models), while Anthropic reported Claude now authors 80% of its own codebase and OpenAI updated ChatGPT's memory system. On the cost front, an unnamed firm's $500M Claude AI bill dominated headlines, spurring broader cost-control moves — Uber capped AI spending at $1,500 per employee monthly, and tools like Factory Router and Lindy's migration to DeepSeek v4 demonstrated significant savings. Infrastructure and tooling saw updates from Perplexity's "Search as Code," SambaNova's disaggregated inference, Supabase's $500M raise, OpenAI's Codex plugins, Grok Composer 2.5, and [Scaledown.ai](http://Scaledown.ai) launch of small-language-model services for compression and summarization

The sections that follow walk through these items day by day, with short context and links so you can dive deeper into the pieces most relevant to your work or interests.

* * *

Sunday 31st May 2026
--------------------

Unnamed firm racks up $500M Claude AI bill in one month

An unnamed enterprise client spent more than $500 million in just 30 days on Anthropic's Claude AI platform after failing to set any usage limits or spending caps for employees, according to an Axios investigation published on May 28. The disclosure, made by an AI consultant whose client generated the bill, may represent one of the costliest IT governance failures on record

[

When AI costs spiral: A company accidentally spent $500 million in one month on Claude AI- what went wrong? | Company Business News
-----------------------------------------------------------------------------------------------------------------------------------

An enterprise client spent $500 million in a single month on Claude AI after failing to set employee usage limits, exposing a growing crisis in corporate AI cost governance.

https://www.livemint.com

![When AI costs spiral: A company accidentally spent $500 million in one month on Claude AI- what went wrong? | Company Business News](https://storage.googleapis.com/papyrus_images/ca7badf04fafb3f1a7bd0712d3d096c783d7ea17d4e4bdca078e375a2a7f7933.jpg)

](https://www.livemint.com/companies/news/when-ai-costs-spiral-a-company-accidentally-spent-500-million-in-one-month-on-claude-ai-what-went-wrong-11780022792096.html)

Xiaomi just published a deep dive into the end-to-end inference engineering behind MiMo-V2.5 and MiMo-V2.5-Pro — and it’s a strong example of how architecture + systems work together to unlock real production gains.

[![User Avatar](https://storage.googleapis.com/papyrus_images/0c8630743c0f7dee330a7190946ce21ca1623d789f6c901b98b390b513d3cbff.jpg)](https://twitter.com/_LuoFuli)

[Fuli Luo](https://twitter.com/_LuoFuli)

[@\_LuoFuli](https://twitter.com/_LuoFuli)

[](https://twitter.com/_LuoFuli/status/2060672928367497480)

Inference Optimizations Behind the MiMo-V2.5 Series API Price Reductions  
  
Read the full technical blog: [mimo.xiaomi.com/blog/mimo-v2-5…](https://t.co/pnRvk1NWyq)  
  
The V2.5 model family, including MiMo-V2.5 and MiMo-V2.5-Pro, is built on a Hybrid Sliding Window Attention (Hybrid SWA) architecture, which

[933](https://twitter.com/_LuoFuli/status/2060672928367497480)[

10:41 AM • May 30, 2026

](https://twitter.com/_LuoFuli/status/2060672928367497480)

[![User Avatar](https://storage.googleapis.com/papyrus_images/616904168fdcbc83d6e7291d50661dc4cd9dfd8ea739d4c9deb333e5b497eada.jpg)](https://twitter.com/ng_thanh8)

[Thanh Nguyen](https://twitter.com/ng_thanh8)

[@ng\_thanh8](https://twitter.com/ng_thanh8)

[](https://twitter.com/ng_thanh8/status/2060706632980644214)

Xiaomi just published a deep dive into the end-to-end inference engineering behind MiMo-V2.5 and MiMo-V2.5-Pro — and it’s a strong example of how architecture + systems work together to unlock real production gains.  
  
The big story: Hybrid Sliding Window Attention (Hybrid SWA). By

[![User Avatar](https://storage.googleapis.com/papyrus_images/29ec46febc6b0f0b74ad155ac4427abc394c06da62c703ab55a17af48c52b6c9.jpg)](https://twitter.com/XiaomiMiMo)

[Xiaomi MiMo](https://twitter.com/XiaomiMiMo)

[@XiaomiMiMo](https://twitter.com/XiaomiMiMo)

[](https://twitter.com/XiaomiMiMo/status/2060682249742471650)

What’s new with MiMo-V2.5 series inference?  
  
We just published a blog on our full pipeline inference optimizations for MiMo-V2.5 series, including how we pushed hybrid SWA efficiency to the limit.  
  
Read the full blog here:  
[mimo.xiaomi.com/blog/mimo-v2-5…](https://t.co/lYBEcgaVhU)

[2](https://twitter.com/ng_thanh8/status/2060706632980644214)[

12:55 PM • May 30, 2026

](https://twitter.com/ng_thanh8/status/2060706632980644214)

Monday 1st June 2026
--------------------

MiniMax Launches M3 AI Model for Coding and Agents 

Chinese AI lab MiniMax released its M3 model  excelling in coding, agentic reasoning, tool use, multimodal chat, and long-context tasks. OpenCode integrated it instantly for free early access, offering 1,400 requests every five hours via terminal, IDE, or desktop app on macOS, Windows, and Linux.   It posts strong benchmark results including 59.0% on SWE-Bench Pro, 66.0% on Terminal Bench 2.1, 74.2% on MCP Atlas, and competitive scores against models like GPT 5.5 and Gemini 3.1 Pro across agentic and tool-use tasks.   Weights and technical report expected in about 10 days; API is immediately available with 50% off standard usage for the first 7 days to encourage early testing.

[![User Avatar](https://storage.googleapis.com/papyrus_images/657e3052af6a561a528b48b3675844a76f1076a770c78daaabe49a313b95df1c.jpg)](https://twitter.com/MiniMax_AI)

[MiniMax (official)](https://twitter.com/MiniMax_AI)

[@MiniMax\_AI](https://twitter.com/MiniMax_AI)

[](https://twitter.com/MiniMax_AI/status/2061266317815296322)

Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities  
  
\- Coding & Agentic Frontier: 59.0% SWE-Bench Pro, 66.0% Terminal Bench 2.1, 34.8% SWE-fficiency, 28.8% KernelBench Hard, 74.2% MCP Atlas  
\- MiniMax Sparse Attention scales context to 1M  
\-

![](https://storage.googleapis.com/papyrus_images/d5819adc1c4a4999a92915d98e5b16985abaeb9109d73e86b8cf2cf79b8dcd0d.jpg)

[8,878](https://twitter.com/MiniMax_AI/status/2061266317815296322)[

1:59 AM • Jun 1, 2026

](https://twitter.com/MiniMax_AI/status/2061266317815296322)

[![User Avatar](https://storage.googleapis.com/papyrus_images/d0993c56f46ef9ef8b27a1420a78ad0a2323457d31006110d21d7cfb3af60a6c.jpg)](https://twitter.com/scaling01)

[Lisan al Gaib](https://twitter.com/scaling01)

[@scaling01](https://twitter.com/scaling01)

[](https://twitter.com/scaling01/status/2061254285849964649)

Full Table:

![](https://storage.googleapis.com/papyrus_images/33909fe1fd3f3f01e860374334dac784b84c6caa71ea5e90c70798057d97762c.jpg)

[46](https://twitter.com/scaling01/status/2061254285849964649)[

1:11 AM • Jun 1, 2026

](https://twitter.com/scaling01/status/2061254285849964649)

[![User Avatar](https://storage.googleapis.com/papyrus_images/996d2635b565461ab4c660268cbab5f26705efe7f3be324041d55100b39288d9.png)](https://twitter.com/opencode)

[OpenCode](https://twitter.com/opencode)

[@opencode](https://twitter.com/opencode)

[](https://twitter.com/opencode/status/2061233503337906187)

MiniMax M3 will be launching soon  
  
You can try it right now in OpenCode  
  
For free

[5,634](https://twitter.com/opencode/status/2061233503337906187)[

11:48 PM • May 31, 2026

](https://twitter.com/opencode/status/2061233503337906187)

NVIDIA Unveils Nemotron 3 Ultra 

The model tops U.S. open-weights rankings with an Intelligence Index score of 48, beating rivals like Gemma 4 31B while delivering over 300 output tokens per second in tests. It shines in agentic AI tasks, boasting 91% agent productivity, strong instruction following, and top coding skills, with 5x faster inference and 30% lower costs than leading competitors. Huang highlighted its openness for developers building everything from search tools to molecular simulations, paired with new agent deployment tools, as Nemotron 4 looms on the horizon.

[![User Avatar](https://storage.googleapis.com/papyrus_images/51db31a2ffc8cf9397bcc6fe71f8af7338664b4c3dced759d79bc3bcc2d9ec3a.jpg)](https://twitter.com/ArtificialAnlys)

[Artificial Analysis](https://twitter.com/ArtificialAnlys)

[@ArtificialAnlys](https://twitter.com/ArtificialAnlys)

[](https://twitter.com/ArtificialAnlys/status/2061304911565144230)

NVIDIA just announced the release of Nemotron 3 Ultra in Jensen Huang's Computex keynote: at 550B parameters (55B active), this is the largest Nemotron 3 model to date, and it is the most intelligent US open weights model  
  
We partnered with [@nvidia](https://twitter.com/nvidia) to evaluate this model for

![](https://storage.googleapis.com/papyrus_images/6fbd7d8b605fd6ba560071e30ffa1420aa7162a21249841bbf49cc253c0e2230.jpg)

[929](https://twitter.com/ArtificialAnlys/status/2061304911565144230)[

4:32 AM • Jun 1, 2026

](https://twitter.com/ArtificialAnlys/status/2061304911565144230)

What I was able to one-shot with Minimax-M3.   

[![User Avatar](https://storage.googleapis.com/papyrus_images/907e534a328e5ec602a790953457a1cb3063a9ae5cc3dad3ce84bd9f265ef11f.jpg)](https://twitter.com/yusufg)

[Yusuf Goolamabbas](https://twitter.com/yusufg)

[@yusufg](https://twitter.com/yusufg)

[](https://twitter.com/yusufg/status/2061429526585106486)

So I gave [@MiniMax\_AI](https://twitter.com/MiniMax_AI) M3 the following task to test its multimodal capabilities, I found the following video via some website. It's basically a video scroll of a geocities page (some of you maybe old enough to remember geocities ![😇](https://abs-0.twimg.com/emoji/v2/72x72/1f607.png)

![](https://pbs.twimg.com/amplify_video_thumb/2061428261687795712/img/cqMk0eWW1F6_oTZh.jpg)

[17](https://twitter.com/yusufg/status/2061429526585106486)[

12:47 PM • Jun 1, 2026

](https://twitter.com/yusufg/status/2061429526585106486)

Tuesday 2nd June 2026
---------------------

SuperGrok and X Premium+ users now can use Composer 2.5 model from Cursor via Grok Build

[![User Avatar](https://storage.googleapis.com/papyrus_images/68f65108c8e4d564dcec1dbe6a418fcc467d9a49eb0e13b073d85c5ebd1b98ae.jpg)](https://twitter.com/skcd42)

[skcd](https://twitter.com/skcd42)

[@skcd42](https://twitter.com/skcd42)

[](https://twitter.com/skcd42/status/2061513477966225522)

We are excited for all of you to try out Composer 2.5 in Grok Build starting today!  
  
To use composer-2-5 do \`/model\` in Grok Build and type in Composer to switch  
  
Composer 2.5 comes with 200k context window and supports: subagents, MCPs, skills and additionally also works with

![](https://pbs.twimg.com/amplify_video_thumb/2061493918185971712/img/bvB9ssvjGPj0B41w.jpg)

[647](https://twitter.com/skcd42/status/2061513477966225522)[

6:21 PM • Jun 1, 2026

](https://twitter.com/skcd42/status/2061513477966225522)

Qwen3.7-Plus is Alibaba's new multimodal agent model unifying vision and language for hybrid GUI/CLI operations, visual perception, reasoning, grounding, coding, and search-augmented tasks. 

Benchmarks show it delivers competitive results across text, multimodal reasoning, visual understanding, agentic coding, and real-world QA, often matching or exceeding models like Claude Opus-4.6, GPT-5.4, and Gemini-3.1-Pro. 

Released via API on Alibaba Cloud Model Studio with demos of browser and interactive agents;

[![User Avatar](https://storage.googleapis.com/papyrus_images/6868acecf957db005a4e15398feb441591f5598882c603dc3b50c5f3068111e3.jpg)](https://twitter.com/Alibaba_Qwen)

[Qwen](https://twitter.com/Alibaba_Qwen)

[@Alibaba\_Qwen](https://twitter.com/Alibaba_Qwen)

[](https://twitter.com/Alibaba_Qwen/status/2061506641120641494)

![👏](https://abs-0.twimg.com/emoji/v2/72x72/1f44f.png)![👏](https://abs-0.twimg.com/emoji/v2/72x72/1f44f.png) Introducing Qwen3.7-Plus — a multimodal agent model that unifies vision and language into one versatile agent foundation.  
  
![✅](https://abs-0.twimg.com/emoji/v2/72x72/2705.png) Multimodal interactive hybrid agent: unified GUI & CLI operation across visual and text tasks  
![✅](https://abs-0.twimg.com/emoji/v2/72x72/2705.png) Versatile coding agent & productivity assistant with

![](https://storage.googleapis.com/papyrus_images/50026be6e4e2ce25ef0e98bc208188077dd2a03a80c8102fa3b0deacfc577540.jpg)

[3,860](https://twitter.com/Alibaba_Qwen/status/2061506641120641494)[

5:54 PM • Jun 1, 2026

](https://twitter.com/Alibaba_Qwen/status/2061506641120641494)

[

Qwen Studio
-----------

Qwen Studio offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.

https://qwen.ai



](https://qwen.ai/blog?id=qwen3.7-plus)

Some discussion between me and [Anya Shapina](https://t.me/anyasha89) related to Composer2-5 model availablility via Grok  
\=====  
Anya

IMO Grok is the most under-appreciated model. I know, my comment is out of context (I haven't used Grok Build enough). But just as a model, it roxx! Testing very complicated agents and workflows across models and it kinda wins most scenarios. In my experience, it excels at walking and chewing gum at the same time (while also looking for shadows lurking behind bushes) - this is where most models fail for me.  Not to mention native search, X, and Reddit. Anyone else in my Grok camp? Sorry, Elon haters :)

Yusuf

Which model of Grok Anya ?  I think grok-build-0.1 is highy underrated and really look forward to the release of v9.  Elon has a tendency to hype up particularly given the SpaceX IPO but the pace of updates of Grok Build CLI is relentless

Anya

Grok 4.3 - in my specific tests powering Agents (Not Grok Build which you post was about; look fwd to reporting on this later). Grok 4.3 - Strongest agentic tool calling + lowest hallucination rate in the lineup, from what I know. Worked for me.  
\=====

Perplexity AI launched Search as Code, enabling AI agents to generate Python code that directly calls atomic search primitives like fanout queries, deduplication, filtering, and ranking within a secure sandbox.

Available in the Perplexity Agent API, and now default in Computer.

[![User Avatar](https://storage.googleapis.com/papyrus_images/e08d2690b4212e9790d6fa3fdc9353c01ef6436bab5b299cf996d52c0b685bfc.jpg)](https://twitter.com/perplexity_ai)

[Perplexity](https://twitter.com/perplexity_ai)

[@perplexity\_ai](https://twitter.com/perplexity_ai)

[](https://twitter.com/perplexity_ai/status/2061506359326384319)

Introducing Search as Code, our new search architecture for AI agents.  
  
It writes Python that calls our search stack directly, instead of looping through function calls one at a time.  
  
Available in the Perplexity Agent API, and now default in Computer.  
  
[research.perplexity.ai/articles/rethi…](https://t.co/ut6GGWQTVO)

![](https://storage.googleapis.com/papyrus_images/b7626faa89595a816f1098f8662284c97c471436ba4a63104f59b809634d0df0.png)

[1,860](https://twitter.com/perplexity_ai/status/2061506359326384319)[

5:53 PM • Jun 1, 2026

](https://twitter.com/perplexity_ai/status/2061506359326384319)

[![User Avatar](https://storage.googleapis.com/papyrus_images/00f2d99db81294245d07d83223ee5451484db56635e3b058e492ac974d4f4657.jpg)](https://twitter.com/AravSrinivas)

[Aravind Srinivas](https://twitter.com/AravSrinivas)

[@AravSrinivas](https://twitter.com/AravSrinivas)

[](https://twitter.com/AravSrinivas/status/2061575845056278971)

We’re moving away from search as a web fetch tool call to search as codegen to be future proof in a world where code execution inside agent harnesses is the way to do almost all of our knowledge work.  
  
Doing this lets you compose multi-step primitives far more naturally and be

[![User Avatar](https://storage.googleapis.com/papyrus_images/e08d2690b4212e9790d6fa3fdc9353c01ef6436bab5b299cf996d52c0b685bfc.jpg)](https://twitter.com/perplexity_ai)

[Perplexity](https://twitter.com/perplexity_ai)

[@perplexity\_ai](https://twitter.com/perplexity_ai)

[](https://twitter.com/perplexity_ai/status/2061506359326384319)

Introducing Search as Code, our new search architecture for AI agents.  
  
It writes Python that calls our search stack directly, instead of looping through function calls one at a time.  
  
Available in the Perplexity Agent API, and now default in Computer.  
  
[research.perplexity.ai/articles/rethi…](https://t.co/ut6GGWQTVO)

![](https://storage.googleapis.com/papyrus_images/b7626faa89595a816f1098f8662284c97c471436ba4a63104f59b809634d0df0.png)

[991](https://twitter.com/AravSrinivas/status/2061575845056278971)[

10:29 PM • Jun 1, 2026

](https://twitter.com/AravSrinivas/status/2061575845056278971)

[![User Avatar](https://storage.googleapis.com/papyrus_images/bd090f8b2ca675a841c9a689e0a44948536aceb4bfde9e4ba7113d7d4fc7c00f.jpg)](https://twitter.com/denisyarats)

[Denis Yarats](https://twitter.com/denisyarats)

[@denisyarats](https://twitter.com/denisyarats)

[](https://twitter.com/denisyarats/status/2061535918998347776)

we're going beyond traditional tool calls / MCPs to interact with the search stack. codegen is the most natural way for an LLM to drive search: the research tasks our users need require complex pipelines, customized per task.  
  
so we're exposing the search stack as composable

[![User Avatar](https://storage.googleapis.com/papyrus_images/e08d2690b4212e9790d6fa3fdc9353c01ef6436bab5b299cf996d52c0b685bfc.jpg)](https://twitter.com/perplexity_ai)

[Perplexity](https://twitter.com/perplexity_ai)

[@perplexity\_ai](https://twitter.com/perplexity_ai)

[](https://twitter.com/perplexity_ai/status/2061506359326384319)

Introducing Search as Code, our new search architecture for AI agents.  
  
It writes Python that calls our search stack directly, instead of looping through function calls one at a time.  
  
Available in the Perplexity Agent API, and now default in Computer.  
  
[research.perplexity.ai/articles/rethi…](https://t.co/ut6GGWQTVO)

![](https://storage.googleapis.com/papyrus_images/b7626faa89595a816f1098f8662284c97c471436ba4a63104f59b809634d0df0.png)

[134](https://twitter.com/denisyarats/status/2061535918998347776)[

7:50 PM • Jun 1, 2026

](https://twitter.com/denisyarats/status/2061535918998347776)

[![User Avatar](https://storage.googleapis.com/papyrus_images/e53391249d485a3ce9e9e6f660781d73fd50556aa15e606a8d2ac02e614f3f40.jpg)](https://twitter.com/VaibhavSisinty)

[Vaibhav Sisinty](https://twitter.com/VaibhavSisinty)

[@VaibhavSisinty](https://twitter.com/VaibhavSisinty)

[](https://twitter.com/VaibhavSisinty/status/2061557154537206186)

Perplexity just rebuilt how AI searches the internet from scratch, and it's a big deal. ![😨](https://abs-0.twimg.com/emoji/v2/72x72/1f628.png)  
  
Here's what it means in simple words.  
  
Every AI today searches the web the same way. It sends one query, waits for results, reads them, sends another query, waits again. One at a time.

![](https://storage.googleapis.com/papyrus_images/6c2a4e32c978aeae96d25982ed636a2b2d60ac36dc353df0477b1ed9b6b1e153.jpg)

[![User Avatar](https://storage.googleapis.com/papyrus_images/e08d2690b4212e9790d6fa3fdc9353c01ef6436bab5b299cf996d52c0b685bfc.jpg)](https://twitter.com/perplexity_ai)

[Perplexity](https://twitter.com/perplexity_ai)

[@perplexity\_ai](https://twitter.com/perplexity_ai)

[](https://twitter.com/perplexity_ai/status/2061506359326384319)

Introducing Search as Code, our new search architecture for AI agents.  
  
It writes Python that calls our search stack directly, instead of looping through function calls one at a time.  
  
Available in the Perplexity Agent API, and now default in Computer.  
  
[research.perplexity.ai/articles/rethi…](https://t.co/ut6GGWQTVO)

![](https://storage.googleapis.com/papyrus_images/b7626faa89595a816f1098f8662284c97c471436ba4a63104f59b809634d0df0.png)

[58](https://twitter.com/VaibhavSisinty/status/2061557154537206186)[

9:15 PM • Jun 1, 2026

](https://twitter.com/VaibhavSisinty/status/2061557154537206186)

Wednesday 3rd June 2026
-----------------------

Uber says it has limited all employees to $1,500 in monthly token spending per AI coding tool “to responsibly encourage agentic AI adoption”

[

Uber caps employee AI spending after blowing through budget in 4 months | TechCrunch
------------------------------------------------------------------------------------

Uber's cutback has occurred after the company had reportedly encouraged staff to use AI as much as possible.

https://techcrunch.com

![Uber caps employee AI spending after blowing through budget in 4 months | TechCrunch](https://storage.googleapis.com/papyrus_images/943e9757f04f90d60ec98d6a7c9027de9aa9bdd17a297f2daf29761204f3f4ca.jpg)

](https://techcrunch.com/2026/06/02/uber-caps-employee-ai-spending-after-blowing-through-budget-in-four-months/)

Microsoft Launches MAI AI Models at Build 2026

The lineup covers reasoning, coding, images, voice, and transcription, all built as a multimodal system optimized for Microsoft's efficient MAIA 200 chips. MAI-Thinking-1, a 35 billion active-parameter model trained on 30 trillion tokens, scores 97% on math benchmark AIME 2025 and tops tough coding tests like SWE-Bench Pro, even beating models from Anthropic and matching Claude in blind evaluation

Microsoft AI CEO Mustafa Suleyman calls it a step toward 'humanist superintelligence' under human control.

[![User Avatar](https://storage.googleapis.com/papyrus_images/ade41e3f78e69664c6e6ac383c5b4b6f3c0b4b4deb16ea367094dc37d416aa6a.jpg)](https://twitter.com/mustafasuleyman)

[Mustafa Suleyman](https://twitter.com/mustafasuleyman)

[@mustafasuleyman](https://twitter.com/mustafasuleyman)

[](https://twitter.com/mustafasuleyman/status/2061880164498428188)

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier.  
First is our text foundation model, MAI-Thinking-1, exceptionally strong on reasoning and SWE tasks.  
\- It’s a

![](https://storage.googleapis.com/papyrus_images/dc11b70290d79c7c35d62c6053e2264abaf9b0612aea363503eb9b1e13e36892.jpg)

[3,771](https://twitter.com/mustafasuleyman/status/2061880164498428188)[

6:38 PM • Jun 2, 2026

](https://twitter.com/mustafasuleyman/status/2061880164498428188)

[![User Avatar](https://storage.googleapis.com/papyrus_images/80580ab09813b4367cace67e2035e149b6c7a2128f6ba76ff9559da3e02129e5.jpg)](https://twitter.com/MicrosoftAI)

[Microsoft AI](https://twitter.com/MicrosoftAI)

[@MicrosoftAI](https://twitter.com/MicrosoftAI)

[](https://twitter.com/MicrosoftAI/status/2061887500541366489)

Seven new models launching at Build: let’s go!  
Reasoning. Code. Image. Transcribe. Voice.  
  
Built from scratch on a clean data lineage, designed for efficiency, working seamlessly as a family of models  
  
Thread ![🧵](https://abs-0.twimg.com/emoji/v2/72x72/1f9f5.png)  
[#MSBuild](https://twitter.com/hashtag/MSBuild)

![](https://storage.googleapis.com/papyrus_images/53d03194f2c1257b86e69d9273ef8521abce373a4175669c6e01a5ee84e0553a.jpg)

[3,312](https://twitter.com/MicrosoftAI/status/2061887500541366489)[

7:07 PM • Jun 2, 2026

](https://twitter.com/MicrosoftAI/status/2061887500541366489)

[

Building a hill-climbing machine: Launching seven new MAI models | Microsoft AI
-------------------------------------------------------------------------------

Alongside distribution on Foundry and optimization for our 1P products, our models are also going to be widely available for developers on OpenRouter, as well as and . For the first time developers will be able to tune the weights of the model themselves.

https://microsoft.ai



](https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/)

OpenAI unveils new Codex plugins for tasks related to public equity investment, banking and sales, and other roles, and plans to integrate Codex into ChatGPT

[![User Avatar](https://storage.googleapis.com/papyrus_images/029dab812e4c272f2eda74f86a3f50a4047d80160e62cb1c72bc4536a142b9b9.jpg)](https://twitter.com/OpenAI)

[OpenAI](https://twitter.com/OpenAI)

[@OpenAI](https://twitter.com/OpenAI)

[](https://twitter.com/OpenAI/status/2061887650391625870)

We’re making Codex more useful for your work by expanding plugins beyond individual tools.  
  
These plugins turn Codex into a specialist for a specific role with a single install, no coding required.  
  
Codex can access 62 popular apps and 110 skills for work across sales, data[

![](https://storage.googleapis.com/papyrus_images/2d1a254bb2dbb1d8bc769e3a097ecccbe7709fe6104624fda3f8e78f7294dd51.jpg)

openai.com

Codex for every role, tool, and workflow
----------------------------------------

Discover new Codex plugins, sites, and annotations that help analysts, marketers, designers, investors, and other teams get more done with AI.





](https://t.co/nunrYP2uMI)

[4,649](https://twitter.com/OpenAI/status/2061887650391625870)[

7:08 PM • Jun 2, 2026

](https://twitter.com/OpenAI/status/2061887650391625870)

[

Codex for every role, tool, and workflow
----------------------------------------

Discover new Codex plugins, sites, and annotations that help analysts, marketers, designers, investors, and other teams get more done with AI.

https://openai.com

![Codex for every role, tool, and workflow](https://storage.googleapis.com/papyrus_images/b0f539ca1af2ff4e9f2f945e10eabc2172c4b11f8307de5a3867dc7800a13c03.png)

](https://openai.com/index/codex-for-every-role-tool-workflow/)

Factory Router automatically selects the best AI model for each task in the Droid coding agent, slashing costs by 20-25% while matching 99% of premium performance on benchmarks like Terminal-Bench 2. It routes routine edits to cheaper models and escalates complex work to heavy-hitters like Claude Opus 4.7, even switching providers mid-task for reliability. Tech leaders like Keith Rabois and Garry Tan praised it as essential for enterprises facing rising AI bills, with CEO Matan Grinberg comparing it to hiring the right expert instead of Einstein for basic math. Currently in private preview for CLI and desktop users.

FWIW,  Droid from Factory is one of my goto model-agnostic harness alongside Opencode Go.  Very good Claude Code compatability in terms of support for skills, plugins, hooks, sub-agents etc and one of the best if not the best context management I've come across

[![User Avatar](https://storage.googleapis.com/papyrus_images/7439c32f6d22a34990166003176081931d761ac2cde767760616f14bcd1eb6af.jpg)](https://twitter.com/FactoryAI)

[Factory](https://twitter.com/FactoryAI)

[@FactoryAI](https://twitter.com/FactoryAI)

[](https://twitter.com/FactoryAI/status/2061862733126275549)

Introducing model routing to Factory.  
  
Factory Router picks the right model for every task, automatically.  
  
Maintain frontier performance while cutting costs by 25%.

![](https://pbs.twimg.com/amplify_video_thumb/2061862679506219008/img/BXUojooRIZxOXo_w.jpg)

[2,165](https://twitter.com/FactoryAI/status/2061862733126275549)[

5:29 PM • Jun 2, 2026

](https://twitter.com/FactoryAI/status/2061862733126275549)

[![User Avatar](https://storage.googleapis.com/papyrus_images/7439c32f6d22a34990166003176081931d761ac2cde767760616f14bcd1eb6af.jpg)](https://twitter.com/FactoryAI)

[Factory](https://twitter.com/FactoryAI)

[@FactoryAI](https://twitter.com/FactoryAI)

[](https://twitter.com/FactoryAI/status/2061862733126275549)

Introducing model routing to Factory.  
  
Factory Router picks the right model for every task, automatically.  
  
Maintain frontier performance while cutting costs by 25%.

![](https://pbs.twimg.com/amplify_video_thumb/2061862679506219008/img/BXUojooRIZxOXo_w.jpg)

[2,165](https://twitter.com/FactoryAI/status/2061862733126275549)[

5:29 PM • Jun 2, 2026

](https://twitter.com/FactoryAI/status/2061862733126275549)

via [James Chan](https://t.me/motochan)

Does anyone have experience with multiplayer PRD like ChatPRD? Any good?

Thursday 4th June 2026
----------------------

Google DeepMind releases Gemma 4 12B, a lightweight 12B-parameter multimodal model under Apache 2.0 license designed to run locally on laptops with just 16GB VRAM or unified memory. 

It introduces an encoder-free unified architecture where vision uses a tiny 35M-parameter embedding module and audio projects raw signals directly into the LLM backbone, eliminating separate encoders for efficiency. 

Delivers advanced reasoning and multimodal capabilities nearing the larger Gemma 4 26B model's benchmarks, with native 256K context, MTP drafters, and immediate support across Hugging Face, Kaggle, llama.cpp, MLX, and vLLM.

[![User Avatar](https://storage.googleapis.com/papyrus_images/0066abc3139d624ed73b5ffaa0f2002fbdad6160ae95a99d2e19cf76e57df32d.png)](https://twitter.com/googlegemma)

[Google Gemma](https://twitter.com/googlegemma)

[@googlegemma](https://twitter.com/googlegemma)

[](https://twitter.com/googlegemma/status/2062202706882883696)

Meet Gemma 4 12B!  
  
A unified, encoder-free multimodal model designed to bring high-performance intelligence directly to your laptop, and released under an Apache 2.0 license.  
  
Bridging the gap between edge efficiency and advanced reasoning. Here is what’s new with Gemma 4 12B: ![👇](https://abs-0.twimg.com/emoji/v2/72x72/1f447.png)

![](https://storage.googleapis.com/papyrus_images/7388410ad88beb0d6602c3f7b936b113ee877e13310156a4ca6a14605303569b.jpg)

[11.9K](https://twitter.com/googlegemma/status/2062202706882883696)[

4:00 PM • Jun 3, 2026

](https://twitter.com/googlegemma/status/2062202706882883696)

[

Bringing Gemma 4 12B to your Laptop: Unlocking Local, Agentic Workflows with Google AI Edge
-------------------------------------------------------------------------------------------

Google DeepMind's Gemma 4 12B model brings agentic, multimodal AI capabilities to everyday laptops with 16GB of RAM, enabling local data processing and visual insight generation. Users can leverage this model on macOS through the Google AI Edge Gallery for dynamic Python code execution and visualization, as well as via Google AI Edge Eloquent for completely offline voice dictation and text editing.

https://developers.googleblog.com

![Bringing Gemma 4 12B to your Laptop: Unlocking Local, Agentic Workflows with Google AI Edge](https://storage.googleapis.com/papyrus_images/3779276d28edb7a23b92e70e3d0633ee03883e1dadb2e36de00359ba3d670160.png)

](https://developers.googleblog.com/bringing-gemma-4-12b-to-your-laptop-unlocking-local-agentic-workflows-with-google-ai-edge/)

Friday 5th June 2026
--------------------

via my high-school classmate [Abhi Ingle](https://www.linkedin.com/in/ingle-abhi/) who works at Sambanova and lurks on this group via the newsletter.  Took me a few readings of the blog post and had to go back and refresh my knowledge on the prefill and decode steps in auto reggressive transformer architectures before the penny dropped on the how this architecture is unique  
\====

At Computex, SambaNova demoed a “disaggregated inference” architecture where GPUs handle prefill and their RDUs handle decode, yielding faster, cheaper long-horizon agent workloads than GPU-only setups, and they now have this running in production-style environments with partners like VC2 and Together AI.

[

3.5 Bn compute commitment to SambaNova ! Taking on the lofty goal of making agentic inference faster and more efficient, Vista Equity Partners, Cambium Capital Management announced Vector Core... | Abhi I.
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------

3.5 Bn compute commitment to SambaNova ! Taking on the lofty goal of making agentic inference faster and more efficient, Vista Equity Partners, Cambium Capital Management announced Vector Core Compute - "The World's first heterogenous disaggregated Inference Cloud".

https://www.linkedin.com

![3.5 Bn compute commitment to SambaNova ! Taking on the lofty goal of making agentic inference faster and more efficient, Vista Equity Partners, Cambium Capital Management announced Vector Core... | Abhi I.](https://storage.googleapis.com/papyrus_images/f8b36019382ba1a7edaac0387c904e67670d33d47f0360b1ada35a0ce6946758.jpg)

](https://www.linkedin.com/posts/ingle-abhi_35-bn-compute-commitment-to-sambanova-share-7467969914008440833-Z-Ej/)

[

The First Disaggregated Inference Demo for AI Agents Is Live
------------------------------------------------------------

SambaNova demonstrates how GPUs and RDUs work together to deliver premium inference for agent workloads using the right chip for the right workload.

https://sambanova.ai

![The First Disaggregated Inference Demo for AI Agents Is Live](https://storage.googleapis.com/papyrus_images/39a6d68717d1c40c7309a079b6c0fd92a1e97ab5e45a692b8ef1c5ea795ffb40.jpg)

](https://sambanova.ai/blog/first-disaggregated-inference-demo-for-ai-agents-live)

OpenAI updates ChatGPT memory with a “more capable and compute-efficient” architecture and a summary page that lets users review and steer what it remembers

[

Dreaming: Better memory for a more helpful ChatGPT
--------------------------------------------------

ChatGPT introduces a new memory system to better remember preferences, keeping context fresh and relevant across conversations.

https://openai.com

![Dreaming: Better memory for a more helpful ChatGPT](https://storage.googleapis.com/papyrus_images/bc6be47821022e21613cf4ecdc10ffe71d73426f46196d0afce29d139350595b.png)

](https://openai.com/index/chatgpt-memory-dreaming/)

Supabase, which provides backend tools for building AI apps, raised a $500M Series F led by GIC at a $10B pre-money valuation, up from $5B in October 2025.  Also released Multigres v0.1 alpha to the open source community,  Multigres tries to bring Vitess-grade horizontal scaling, high availability, and operational simplicity to Postgres.

[![User Avatar](https://storage.googleapis.com/papyrus_images/5cc7c51379725800f9a0ec5cfc6f9ced854e1aa4a949ede9ba6274aefb9b4d0b.jpg)](https://twitter.com/coatuemgmt)

[COATUE](https://twitter.com/coatuemgmt)

[@coatuemgmt](https://twitter.com/coatuemgmt)

[](https://twitter.com/coatuemgmt/status/2062604734373023955)

We led @Supabase's Seed and Series A. Five years in, they're the open source Postgres platform powering the majority of AI app builders — and 9 million developers strong.  
  
As agents reshape how software gets made, Supabase has become the default backend underneath it. Proud to

![](https://storage.googleapis.com/papyrus_images/046349721b40e74ce75dac0364cdd0aba186e52aef485ed36ff094bc4f1f46c7.jpg)

[28](https://twitter.com/coatuemgmt/status/2062604734373023955)[

6:37 PM • Jun 4, 2026

](https://twitter.com/coatuemgmt/status/2062604734373023955)

[

Supabase Series F
-----------------

Supabase has raised a $500M Series F at a $10B pre-money valuation, led by GIC.

https://supabase.com

![Supabase Series F](https://storage.googleapis.com/papyrus_images/ab80716c60b0ee19131a969930b200275afe77044ea87fa290e75244292088b6.png)

](https://supabase.com/blog/supabase-series-f)

[

Multigres v0.1 Alpha: an operating system for Postgres
------------------------------------------------------

Today we're releasing Multigres v0.1 alpha to the open source community, bringing Vitess-grade horizontal scaling, high availability, and operational simplicity to Postgres.

https://supabase.com

![Multigres v0.1 Alpha: an operating system for Postgres](https://storage.googleapis.com/papyrus_images/9236d6c32e205195891d92f0bc21bb0080ffdb483d33a4febeb373f875a95907.jpg)

](https://supabase.com/blog/multigres-v0-1-alpha)

Response from [Anya Shapina](https://t.me/anyasha89)  
\===  
I believe Supabase (and possibly its philosophical rival Convex) will be some of the biggest winners in the AI race. I use it as the shared memory and persistent backend of my multi-agent system - it's everything to my project, and all other projects I've ever built.  But then there is the Goliath Convex with its built-in support for AI Agents... and NO SQL (love/hate?).  Which camp are you people in?  
\====

I've not used Lindy but have heard good things about it from others.   For those unaware of Lindy, it describs itself productivity and automation platform that positions itself as an "AI executive assistant" or "AI employee." It is particularly strong for professionals who want proactive help with email, meetings, and calendar management.

Saw this post from Flo Crivello, founder of Lindy who migrated 100% of its traffic to DeepSeek v4 from Anthropic models, reporting millions in cost savings alongside performance improvements on core use cases.  The switch required building substantial new infrastructure and internal tooling, described as 100x more work than anticipated, emphasizing the value of swappable model architectures.

[![User Avatar](https://storage.googleapis.com/papyrus_images/aa7949e665cc7a38df2d4062dd5e5a51fc693e13d1b8af70567b32c59c94b0a8.jpg)](https://twitter.com/Altimor)

[Flo Crivello](https://twitter.com/Altimor)

[@Altimor](https://twitter.com/Altimor)

[](https://twitter.com/Altimor/status/2062389885437366342)

Pulled the trigger today and switched 100% of Lindy traffic to DeepSeek v4, churning from Anthropic models.  
Saves us millions of $ and we're actually seeing an \*increase\* in performance on many core use cases. Transformative for the business.

[2,177](https://twitter.com/Altimor/status/2062389885437366342)[

4:24 AM • Jun 4, 2026

](https://twitter.com/Altimor/status/2062389885437366342)

Response from [Ms Macarena Correa](https://t.me/RHmacx) to a post dated 30th May 2026 about the announcement of [Kirkland and Ellis committing $500 million over the next 3-4 years to build its own proprietary AI platform and custom tools](https://paragraph.com/@twiata/this-week-in-all-things-ai-week-22-2026#h-saturday-30th-may-2026)  
\=====  
On this one, I think its the right way forward. If all Magic Circle / Silver Circle law firms are using the same AI tools to help drafting contracts (Harvey / Legora), and review them too, as well as providing advice, all solutions will become kind of standard, so ultimately the difference between law firms will be reduce only to 1) pricing and 2) charisma of each particular lawyer (e.g., building the relation with the client). this 2nd point cannot be replaced by AI, so probably most of the efforts need to shit into building more and better relations of trust.  
\=====

Anthropic details its progress toward recursive self-improvement, and its implications, and says Claude has authored 80%+ of the code merged into its codebase

[

When AI builds itself
---------------------

Our progress toward recursive self-improvement, and its implications.

https://www.anthropic.com

![When AI builds itself](https://storage.googleapis.com/papyrus_images/18b3c70b6298faa3457d1d79c8cd03e7ab16342d5f857cd72079459793e5bd35.jpg)

](https://www.anthropic.com/institute/recursive-self-improvement)

via [Imran Muthuvappa](https://t.me/formation313ye)

made a free community for people who want to build and ship their first agent

[

agentmaxxing
------------

No AI background? I'll help you build, deploy, and sell your first AI agent to a small business in 90 days.

https://www.skool.com

![agentmaxxing](https://storage.googleapis.com/papyrus_images/ae6506c247e4c2f9afdeca3f38226e85503538d5ac60076969864775f447bd3e.jpg)

](https://www.skool.com/agentmaxxing/about?ref=13439e0d58404aefbebb7c7dbde2e3c1)

Saturday 6th June 2026
----------------------

Came across Scaledown which has developed purpose-built models for compression, summarization, extraction, and classification.  Pricing is 0.05/M tokens with self-deployment options also available

[![User Avatar](https://storage.googleapis.com/papyrus_images/1b7caee663b1e77ec876d9dfbfaa3f439e978afe0aa987c621385a59fcdc7e86.jpg)](https://twitter.com/neal_k_patel)

[Neal Patel](https://twitter.com/neal_k_patel)

[@neal\_k\_patel](https://twitter.com/neal_k_patel)

[](https://twitter.com/neal_k_patel/status/2062534030638141695)

Introducing ScaleDown  
  
15x cheaper. 63x faster. 5.1% more accurate than GPT-5.4 Mini.  
  
Task-specific SLMs for the 70-80% of AI workloads that don't need a frontier model.  
  
From NeurIPS '25 to [scaledown.ai](https://t.co/imZLOYCcTA)

![](https://pbs.twimg.com/amplify_video_thumb/2062533530547101696/img/_1x78_bxY7o4ERHr.jpg)

[952](https://twitter.com/neal_k_patel/status/2062534030638141695)[

1:56 PM • Jun 4, 2026

](https://twitter.com/neal_k_patel/status/2062534030638141695)

[

Scaledown.ai - Task-Specific Small Language Models
--------------------------------------------------

Purpose-built models for compression, summarization, extraction, and classification. Frontier quality at a fraction of the cost.

https://scaledown.ai

![Scaledown.ai - Task-Specific Small Language Models](https://storage.googleapis.com/papyrus_images/788ef8e17963028e8c2abb4424be90236f41de958e9c498417d293528338ab8f.png)

](https://scaledown.ai/)

[https://scaledown.ai/blog](https://scaledown.ai/blog)

* * *

Below is my personal website which aggregates links to many of my socials as well as the various content and community that I curate. Feel free to share this link to others who you think may find this content/community useful to them

[**https://linktr.ee/goolamabbas**](https://linktr.ee/goolamabbas)

The cover image of this newsletter via generated via the **Krea 2 Large** model within the [**Krea**](https://www.krea.ai/refer/EJQQP9FJ) tool via the following prompt

![](https://paragraph.com/editor/callout/information-icon.png)

An oil painting in the style of H.R. Giger and Zdzisław Beksiński of an ancient Greek temple on the side of Mount Olympus, surrounded by cacti, a rainbow in the sky, grey clouds overhead, blue lilies around the base, and birds flying above. The painting is highly detailed.

---

*Originally published on [This Week in All Things AI](https://paragraph.com/@twiata/this-week-in-all-things-ai-week-23-2026)*
