This Week in All Things AI covers key developments in models, agents, tools, infrastructure, and policy curated from discussions in the All Things AI Telegram group.
If you follow AI for work, research, investing, or just to understand where the technology is heading, this weekly brief is a concise way to scan the most important launches, risks, and resources in a few focused minutes.
I would like to begin by sharing the recording of the recent roundtable I conducted with Shamir Allibhai, CEO/co-founder of Eddie AI. Shamir was very eloquent in describing his entrepreneurial journey and what the world of professional video editors looks like and why a personal itch drove him to build a product which is gaining accolades from industry peers.
The week of 20th to 26th September 2026 saw all three frontier labs—OpenAI, Anthropic, and xAI—release upgraded models. GPT-6 Sol and Luna, Claude Opus 5.5, and Grok 4.7 brought improvements in coding, reasoning, and agent work, with lower costs or stronger performance at existing prices. StepFun’s Step 5 Preview and Xiaomi’s MiMo-V2.6 added competitive, lower-cost options, with Xiaomi releasing open weights and StepFun announcing theirs for October. Speech and audio followed with their own wave of releases: Qwen-Audio-3.1 combined broader capabilities with substantial price cuts, Google’s Gemini 3.8 TTS models expanded expressive voice creation, and NVIDIA’s Nemotron 3 Diarization improved speaker identification, including through its integration into MacWhisper. Free previews of Space Bunny Alpha and Pixel Canary gave developers further options to explore, with different strengths and data-use terms.
Making these capabilities practical was another recurring theme. Google’s emerging AX orchestrator explored how agents can preserve and resume work at scale, while Cursor reduced agent costs through tighter prompts, selective tool loading, and better caching. Firecrawl’s Alexandria expanded access to specialised datasets, and Perplexity’s Fast Search targeted faster, cheaper retrieval. Alibaba’s roadmap connected models and agents to ambitious chip and cloud expansion plans, while discussion of Nscale’s filing and alternative computing architectures raised questions about how efficiently AI investment becomes usable capacity. Perplexity’s SPACE research highlighted a separate challenge: agents can remain inside a sandbox while exceeding intended network permissions. Chinese local-government incentives for AI filmmaking extended the week’s focus to adoption, as cities sought to turn creative experimentation into an industry.
The sections that follow walk through these items day by day, with short context and links so you can dive deeper into the pieces most relevant to your work or interest
via Shivameet
Built a free browser-based tool that removes the visible sparkle watermark from Gemini / Nano Banana images and Veo videos — sharing here in case it's useful.
The value proposition: the sparkle is composited with a known alpha-blending formula, so instead of AI inpainting (which leaves a smudge), the tool reverses the math and recovers the original pixels underneath exactly — pixel-perfect, no blur. Everything runs client-side in the browser: no uploads, no login, no server processing. Works on both images and videos, batch supported.
Honest caveat, since this group will ask anyway: it removes only the visible mark, not SynthID — no tool honestly does that.
Web: tool.removegenteam.workers.dev
Telegram bot: @removegen_gemini_watermark_bot
Happy to answer questions about the approach or hear critiques.
I truly wasn't expecting this huge of an intelligence jump from StepFun's upcoming model. This will be interesting when its weights become available and more inference providers are able to serve it
StepFun announces Step 5 Preview, its new flagship MoE model for agentic workflows, emphasizing frontier performance in software engineering, professional knowledge work, and finance with sustained long-horizon execution.
The 600B total / 27B active parameter model offers 1M context, vision support, and substantially lower per-task costs while charts in the post position it as advancing the Pareto frontier in intelligence versus cost compared to models like GPT-6 and Claude Opus 5.
StepFun, the Shanghai-based AI company behind the Step series, will release open weights on October 15, making the preview available now via their platform for testing in coding, analysis, and financial applications.
Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.
Pricing via its API is currently at 1.00/2.70/0.05 [input (cache-miss), output, input (cache-hit)]
https://artificialanalysis.ai/#intelligence
via Alex
The Nscale S-1 has been released: https://www.sec.gov/Archives/edgar/data/2110365/000119312526395475/ck0002110365-20260918.htm
Operating loss margin of -3000% on a non-GAAP accounting. Peak bubble?
Operating loss margin of -3000% on a non-GAAP accounting. Peak bubble?
The S-1 also confirms recent speculation that only a small fraction of sold GPUs are currently in use. 95% of NSCALE inventory is sitting idle according to the filing
Google ML/AI engineer Jaana Dogan announced ongoing work on AX, Google's open agentic orchestrator and runtime, which reinvents Kubernetes concepts for stateful AI agent workloads with emphasis on fast resumption, suspension, and reproducibility.
The demo video shows the ax CLI in action: applying YAML manifests for workspaces and tasks (e.g., setting up Go environments from Git), monitoring status, SSH access for commands like builds, and stateful operations like suspend/resume/delete to preserve progress.
AX targets scaling to billions of agents per cluster with 10-20x higher sandbox density, declarative primitives for isolation and networking, and open development to support agentic applications, RL workloads, and evaluations beyond traditional batch or stateless jobs.
“The weights do not need to move.” — Dave Blundin
Why spend so much energy moving an AI model’s parameters between memory and processors, over and over?
In this discussion, Dave Blundin joins Cerebras CEO Andrew Feldman and semiconductor veteran Atiq Raza to explore what happens when we rethink the hardware running AI.
Blundin makes a striking claim: architectures that keep those parameters in place could be 1,000–1,000,000× more efficient An ambitious prospect—and the panel digs into the engineering and supply-chain hurdles.
They also explore why faster AI could unlock entirely new businesses, how chip restrictions have pushed China toward greater efficiency, and where founders should look for opportunities.
One revealing detail: Feldman says OpenAI initially gave its new Cerebras computing capacity to its own engineers, prioritising faster development over customer revenue.
What could you build if AI were 10× faster? That’s the question to keep in mind while watching.
SpaceXAI Launches Grok 4.7 for Coding and Knowledge Tasks
SpaceXAI has released Grok 4.7, a larger frontier model than Grok 4.6, trained with extended reinforcement learning for complex, multi-hour tasks. It excels in self-verification, long context handling up to 500,000 tokens, and agent integrations, with strong benchmarks on coding tests like CursorBench 4.0 and DeepSWE. Priced at $2 per million input tokens and $6 per million output tokens, it's now available in Cursor, Grok Build, the SpaceXAI API, and select tools, including a faster variant.
Xiaomi Launches MiMo-V2.6-Pro as Top Open Source AI Model
The flagship MiMo-V2.6-Pro is a Mixture-of-Experts model with 1.02 trillion total parameters and 42 billion active ones, handling text, images, video, and audio natively. A smaller MiMo-V2.6-Flash variant uses 310 billion total and 15 billion active parameters. Both are open-sourced under MIT license on Hugging Face, with the same low API pricing—$0.435 per million input tokens—and reinforcement learning that boosted training pass rates after public livestreamed runs costing under $3 million total. Early tests show strong results in coding, agents, and cybersecurity against rivals.
Claude AI announces Opus 5.5 as the first in the Claude 5.5 family, matching Fable 5.1 performance on most tasks while costing 40% less and generating output over 30% faster than Opus 5.
The model features benchmark leadership in agentic coding, computer use, knowledge work, and multidisciplinary reasoning, with a pricing table showing lower per-token rates including reduced cache costs.
Released after external safety testing by METR and Frontier Design, it sets a new high on alignment evaluations following Anthropic's call to pace frontier development, with an artistic video highlighting themes of discovery.
Following is snippet of guidance from Anthropic
Instructions written for an older model can make Opus 5.5 write more and repeat tool calls. Run /claude-api prompt-audit in Claude Code to check your Claude Code setup, such as your skills and CLAUDE.md file, for these prompting anti-patterns. It also checks the code of an app you build on the Claude Platform.
🚀 OpenAI launches GPT‑6 Sol and Luna: more capable AI at lower prices
The two models bring advances from GPT‑6 Astra to more affordable options for everyday work, coding and automation. Astra remains OpenAI’s most capable model.
What’s new?
• Stronger performance: Improvements in coding, business workflows and computer use.
• Better factual accuracy: OpenAI reports Sol makes roughly half as many factual mistakes as its predecessor on its internal evaluation.
• Clearer answers: Less jargon and more concise communication.
• Cheaper API usage: Sol costs $2 input / $10 output, and Luna $0.10 input / $0.50 output, per million tokens.
• Better caching: Reusing context helps make long conversations and AI agents faster and cheaper, with a 90% discount on cached input reads.
Where can you use them?
They’re available through the API and rolling out in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. Free and Go users can access Luna in the desktop app. They aren’t yet available in regular Chat.
The practical appeal: more room to use capable AI for ongoing work without paying Astra-level prices.
Alibaba’s main AI news this week was a full-stack roadmap announced at its September 22 Apsara Conference—covering Qwen models, chips, cloud infrastructure, and AI agents.
- Models: Qwen 4 is in training; Qwen 4.5 and Qwen 5 are on the roadmap, with future models projected to reach 5–10 trillion parameters. Alibaba also introduced speech and audio updates, including LiveTranslate and Qwen-Audio tools. Qwen-Image 3.1 is planned for later this year.
- AI agents: Alibaba announced Qwen Intelligence, a Qwen-powered agent platform for phone makers, plus enterprise cloud tools for building and managing agents, security, and persistent context.
- Chips and data centers: Alibaba unveiled the Zhenwu V900 accelerator, with commercial release planned for Q1 2027, and Yitian CPU plans for 2027. It also set a goal of exceeding 20 GW of Alibaba Cloud data center capacity by 2032.
The timing matters: several items are future plans, not products available now. Alibaba’s performance and efficiency figures are company-reported. The Apsara Conference runs through September 24, so further announcements may follow.
Firecrawl Raises $75M for Alexandria AI Knowledge Library
Firecrawl raised $75M in Series B funding led by Smash Capital to launch Alexandria, a knowledge library designed for AI agents and superintelligence with instant access to extensive datasets from 100+ providers and custom indexes.
It enables direct searches across specialized data including research papers, US homes/rentals, professional profiles, and product prices, outperforming standard web tools by 21% in answer quality based on internal evaluations of 1,000 questions.
Alexandria focuses on compensating data providers and content creators while offering quick agent integration via CLI, API, and MCP, with a call to build toward solving major problems through better data infrastructure.
via Raw Mind
Quick builder share for the AI-tools crowd: I’m building RAWMIND, an uncensored/unfiltered multimodal AI workspace for chat, image, and video. It supports wallet or seed login—no email—and I’m looking for honest feedback on the experience. If you’re curious, it’s here: https://rawmind.xyz/
via Matthew Weigand
For those interested in agent security, this one is worth a read. https://www.perplexity.ai/hub/blog/escaping-space-part-i
An AI agent can stay inside its sandbox and still exceed its permissions.
Perplexity’s SPACE research shows how shared IP addresses can let agents reach blocked services through allowed network connections.
The lesson for anyone building agents- verify which service the agent actually reaches. An allowed IP alone isn’t enough.
Cursor Cuts AI Coding Agent Token Costs by 7% Without Quality Loss
The company achieved the cost reduction through precise changes like shortening the system prompt by 66%, loading tool definitions only when needed to cut static tokens by 60%, and optimizing cache to reduce misses by 20%. Production data reveals read tools dominate at 91% of conversations, with grep, glob, and shell close behind, while line numbers in reads now appear every tenth line to trim another 1.6%.
Eric Zakariasson even shared the following
here's a prompt to improve your agent harness based on what we've learned at cursor. enjoy
After Monday and Tuesday torrent of LLM model updates, it's time for the speech-to-text and text-to-speech world to do the same so stay tuned for a series of posts on updates on models which deal with ASR and TTS
Alibaba's Qwen team announces Qwen-Audio-3.1, expanding their audio AI lineup with upgraded ASR, TTS, and Realtime models plus new ASR-Next and TTS-Next variants to deliver full-stack capabilities in speech understanding, generation, interaction, and creative audio production.
Notable upgrades focus on multilingual/dialect handling, multi-speaker transcription with timestamps and emotion detection, ambient sound recognition for audio QA, and unified TTS-Next generation of voices, effects, and backgrounds using an LM-diffusion hybrid.
The release features steep price cuts of 70-95% across the models to boost accessibility, with immediate API links provided and community replies showing keen interest in open weights availability.
Google AI announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, its most expressive text-to-speech models enabling custom voice creation in 100+ languages from prompts or 2,000+ presets, with support for natural cues like <laughs> or |mhm|, line-by-line direction, and hours of consistent audio.
Flash TTS suits high-fidelity creative work such as gaming, audiobooks, and podcasts with granular control for bespoke personas, while Flash-Lite optimizes for cost-efficient real-time voice agents, dynamic tone adjustments, and bulk dubbing.
Both models are rolling out today to developers in Google AI Studio and Gemini API, consumers in Gemini Notebook and Google Vids, with enterprise access soon; they lead benchmarks in voice quality, accent modeling, and long-form consistency.
via Jordi Bruin whose MacWhisper product is something I use as my workhorse for generating transcripts locally
We've made some really nice improvements to MacWhisper in the last month and today we're releasing MacWhisper 15.2 with a lot of core improvements to speaker recognition thanks to a new NVIDIA model
Today NVIDIA is releasing their latest speaker recognition model Nemotron 3 Diarization, and MacWhisper has support for it from day 1! Thanks to our collaboration with Argmax we've been testing the new diarization model and it performs a lot better than our previous model
Perplexity AI introduced Fast Search in its API, powered by Photon, a new Rust-based retrieval and ranking service built by a small engineering team using hundreds of agents. It is now the default search in Hermes Agent for Nous Portal subscribers, and Nous Research says it is free on every Portal tier.
Fast Search is priced at $1 per 1,000 requests and is described as the lowest cost per task among other search APIs. It returns 95% of results in 230 ms or less, with a median single-search latency of 160 ms. On six agent benchmarks it cut estimated model-plus-search cost by about 68% ($59.73 versus $187.60 on 3,554 tasks) at comparable aggregate quality.
Photon replaces the prior open-source engine with compact indexes, asynchronous disk reads, and rolling index updates that warm caches before taking traffic. In production, retrieval-and-ranking p99 fell from about 800 ms to about 65 ms, with roughly 20% fewer serving machines and 2.5× more data per document. The fast preset trades a little internal relevance and answer availability on long-tail queries for that speed and cost.
Chinese local governments are offering subsidies like computing vouchers, rent waivers, and dedicated funding to lure AI filmmakers as part of China's AI push
Two free stealth models dropped this week. Both are preview-only — try them before the window closes.
Space Bunny Alpha
Listed 23 Sep on OpenRouter as stealth/space-bunny-alpha
• 1M context, ~524k max output
• Text + image + video in, text out
• Always-on reasoning (adjustable effort)
• Tools, fast inference, strong at coding / agent loops
• Free during preview. Provider says it does not train on your prompts.
Try: https://openrouter.ai/stealth/space-bunny-alpha
🐤 Pixel Canary
Listed ~25 Sep on Vercel AI Gateway as stealth/pixel-canary
• Coding model aimed at frontend + mobile UI
• ~262k context
• Vercel Next.js eval: 28/31 baseline, 30/31 with docs in context
• Free for a limited time
• No ZDR — prompts/responses may be used for training. Don’t send secrets.
Try: https://vercel.com/changelog/pixel-canary-is-now-available-in-stealth-for-free-on-ai-gateway
Quick split
Bunny → long context, multimodal, agents
Canary → Next.js / UI / app-building
Both → free now, no official vendor name, can vanish or start costing money
Read more on X
OpenRouter announce: https://x.com/OpenRouter/status/2102772427105407270
Bunny demo (3D solar system): https://x.com/HeyZaraKhan/status/2103451247785480221
Both models in one roundup: https://x.com/VaibhavSisinty/status/2103772515701260461
Pixel Canary drop: https://x.com/Ved_CJ/status/2103755573217128922
Canary usage note: https://x.com/CommandCodeAI/status/2103717089869808029
Below is my personal website which aggregates links to many of my socials as well as the various content and community that I curate. Feel free to share this link to others who you think may find this content/community useful to them
The cover image of this newsletter via generated via the GPT Image 2.5 model within the Krea tool via the following prompt
a magnificent modern ecotourism village with many different and singular residential buildings designed based in style of architects Oscar Niemeyer and Lucio Costa, human eye view, central avenue, central garden with lake, tree-lined streets, pedestrians, square garden, lake, modern cars, drone photo 8k, realistic, golden hour, volumetric lights,



Excited to start revealing what we've been working on in the last few months. First, we decided to reinvent Kubernetes for agentic workloads with statefulness and fast resumption. Secondly, we are building an agentic orchestrator that will be Google's open agentic orchestrator


Two omnimodal models, advancing through scaled reinforcement learning








Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding.
up to 8 speakers
21 languages
15000 RTFx
variable latency down to 0.32s
commercial and non-commercial use


