This Week in All Things AI covers key developments in models, agents, tools, infrastructure, and policy curated from discussions in the All Things AI Telegram group.
If you follow AI for work, research, investing, or just to understand where the technology is heading, this weekly brief is a concise way to scan the most important launches, risks, and resources in a few focused minutes.
The week of 23rd to 29th August 2026 saw Chinese AI labs push the open-weight frontier with a cluster of multimodal, long-context MoE releases. Z.ai launched GLM-5.3-Flash—previously previewed as Ox Alpha—while Alibaba introduced Qwen3.8-Flash and Tencent released Hy4 preview; all target coding, tool use, and agentic productivity with million-token-scale context windows and aggressive pricing. Video generation also advanced with Wan 3.0 becoming available on ArtArch, while Google expanded Gemini Notebook with ebook-grounded “Expert Intelligence.”
AI infrastructure and agent tooling were equally prominent: OpenAI reported early efficiency and latency gains from its Jalapeño inference chip, Nvidia’s Groq 3 LPX entered production, and Perplexity launched Portable Computer for fully local agent execution on NVIDIA DGX Spark. Developer discussions focused on making agents more capable and interoperable, from external search/fetch skills and free agent web access to Shopify CEO Tobi Lütke’s criticism of proprietary instruction-file conventions; OpenAI’s decision to end Cursor’s direct model access after its SpaceX acquisition further underscored intensifying competition across the coding-agent stack.
The sections that follow walk through these items day by day, with short context and links so you can dive deeper into the pieces most relevant to your work or interest
via David An
In our latest Provocation Talk, I speak with Advait, from Opengradient.AI about the challenging topic of AI Privacy
https://www.linkedin.com/feed/update/urn:li:activity:7497552042186403840/
I created the below public Notion page to articulate the benefits of adding external search/fetch providers in an AI harness and to suggest a prompt prefix for users to incorporate into prompts that utilise some of these providers.
These approaches can be useful if you are looking to take advantage of specialized data that external search providers can bring to the table within your AI harness.
The mechanism described on the below Notion page can be useful for GTM intelligence , lead-generation, finance or investment research, competitive intelligence or market-mapping workflows, among others.
The below page is best viewed on a desktop or laptop
I humbly welcome your comments.
I updated the above Notion page to include a skill sub-directory that you can drop into your skill folder as well as some snippets to augment into AGENTS.md / CLAUDE.md and provides some guidance on invoking one or more of these providers in a prompt.
So some upfront investment in creating accounts with these providers, crediting balances beyond what they provide initially and integrating them in your harness but subsequently the skill file as well as some lines in AGENTS.md/CLAUDE.md should make this very easy to incorporate external search providers in a prompt
Nvidia says its inference accelerator Groq 3 LPX has entered full production, with Nebius signing on as the first customer, and SpaceX will deploy Vera CPUs
Shopify CEO Tobi Lütke criticizes Anthropic's Claude Code for ignoring the emerging open standard AGENTS.md and .agents/skills directory, sticking only to its proprietary CLAUDE.md file.
This creates "split brain" inconsistencies in large teams where developers use multiple AI coding tools, requiring workarounds like symlinks or @imports
that add unnecessary friction.
The post highlights growing industry pressure for AI agents to adopt shared, vendor-neutral conventions to streamline codebase instructions across tools like Cursor, Codex, and Copilot.
PS: This is also my pet peeve with Anthropic that their harness will read only in their directories and filename.
OpenAI's post shares initial test results for Jalapeño, its first custom AI inference chip, highlighting gains in power efficiency and response speed by delivering higher throughput and lower latency simultaneously.
Performance benchmarks show Jalapeño achieving 1.5-1.9x more AI work per watt and 1.7-3.6x lower end-to-end latency compared to prior systems across large models like 120B, 670B, and 1T parameter sizes.
Deployment in OpenAI's infrastructure is planned by the end of 2026, as the first step in a multi-generation hardware roadmap with Gen 2 already in development to further boost efficiency for products like ChatGPT.
Some snippets on SemiAnalysis longish article on Jalapeno
In general first generation chips are not competitive, but OpenAI bucks the trend by being industry leading and beating every Nvidia, AMD, and Google chip we have been able to test on multiple top open source models. OpenAI does this with extreme hardware software codesign. Surprisingly, OpenAI is not over specialization on any specific part of model inference, but instead by focusing on being a general chip that delivers high performance in all scenarios.
Everyone says that OpenAI’s chip is specialized for OpenAI models, but that’s wrong, OpenAI made a generalized chip for AI inference.
Dwarkesh Patel shares his latest podcast with Dylan Patel of SemiAnalysis, analyzing AI lab economics including the projected shift of compute from inference to training as recursive self-improvement nears.
The discussion forecasts Anthropic and OpenAI controlling most global usable FLOPs within years by outbidding others through superior monetization, driven by economies of scale, compute scarcity, and continual learning.
They examine risks from over $10T in AI capex by decade's end potentially triggering sovereign debt crises via elevated interest rates, hyperscaler borrowing, and economic fallout for non-AI sectors and countries.
Perplexity Launches Fully Local AI Agent on NVIDIA DGX Spark
Perplexity released Portable Computer, a local version of its AI agent platform that handles orchestrator, subagents, and tasks on NVIDIA DGX Spark without cloud reliance for core functions. It starts tasks locally by default and only seeks user approval for cloud use on complex reasoning, ensuring privacy by flagging personal data and avoiding direct file access. Available now for Pro and Max subscribers with compatible hardware, it excels in benchmarks like 85.4% on real-world tasks and supports app connectors for Gmail, Slack, and GitHub. NVIDIA's Jensen Huang even gifted a DGX Station to the team after an early demo, fueling excitement for private, affordable AI.
"Even under the best circumstances, the transition to this new AI era will be one of the most turbulent times in human history," Bill Gates writes in a new, 12-page 5,784 word essay.
He says that “we are not preparing adequately”, calling for a regulatory framework
He also says tech executives are privately “very worried” about AI disruption but publicly downplay the risks to protect fundraising and planned IPOs
Looks like it's going to be a China open-weights night/US early morning with Alibaba's Qwen3.8-Flash-Next as well as GLM-5.3-Flash which was unmasked as the stealth model 0xalpha
https://x.com/i/trending/2092526498800718248
Z.ai announces GLM-5.3-Flash, a 320B total / 18B active parameter natively multimodal model with 1M-token context, released under MIT license and optimized to run on Chinese AI chips after preview as Ox Alpha.
The model offers strong coding and agentic performance at low cost, with API pricing of $0.15 input / $0.50 output per million tokens, outperforming GLM-5.2 across benchmarks like DeepSWE and AutomationBench. 50% discount till Sep 9th via OpenRouter
Architectural changes including hybrid sparse-linear attention and specialized pre-training enable efficient scaling, achieving results near Claude Opus 4.8 on select tasks while reducing compute demands significantly.
From OpenRouter
We expect many providers to onboard this model throughout the week
via Alba Chung
Hi this is Alba with Alibaba ☁️
Wan 3.0 is now on ArtArch AI
→ Reality-grade rendering
→ Cinematic motion
→ Stronger scene consistency
Limited time: Aug 25–Sep 7: 100 FREE generations for every user.
No duration limit.
Start creating:
Curating some links about Qwen-3.8-Flash-Next which is now available via OpenRouter at 0.16/0.47/0.016 so priced competitively with GLM-5.3-Flash.
Shall be interesting when other providers join the fray in serving this model.
An impressive technical breakdown by Ryan Smith on Sambanova's SN50 RDU
Nvidia Acquires Hugging Face for $12.9 Billion
The agreement gives Nvidia control of Hugging Face, often called the GitHub for AI, where millions of models and datasets are shared by researchers and companies worldwide. Reports from The Information, Bloomberg, and Business Insider pegged the valuation around $13 billion, down slightly from earlier talks, following Nvidia's prior investment in the startup's 2023 funding round at $4.5 billion. The move expands Nvidia beyond chips into the open-source ecosystem that powers AI on its GPUs, even as Hugging Face's revenue sits around $150 million annually. It highlights hardware giants racing to own more of the AI stack amid strong demand.
https://x.com/i/trending/2092825795001790544
For the video editors in da house via Shamir Allibhai of Hey Eddie fame who is tentatively my upcoming guest in my conversation series. Luma link for folks to register for the virtual event shall be posted once dates/times are finalised with Shamir
The iOS version of Eddie AI is now available
https://apps.apple.com/app/eddie-ai-video-editor/id6792122289
Shamir also pointed me towards this completely free product he built https://99.life/
Here's a brief blurb about the above product
99 is a free arena for AI video editing. You bring the clips. Two AI models each cut their own edit. You watch both without knowing which model made which, and pick the winner. Every vote moves the public scoreboard.
As to why the above product is Free
Why it's free
No trick, just a trade. Blind votes from real people on real footage are the honest measure of an AI editor, and model makers pay for that measure. 99 owns the edits the models generate, and the evaluation data the arena produces — votes, watch behavior, and the content they judge — is used to train AI models and is licensed to the companies that build them. That deal funds the arena. You get free edits and a real say in which models win; the model makers get judged. We think that is a fair trade, and we would rather say it in bold than bury it in fine print. The details live in our Terms and Privacy Policy.
Tencent Hunyuan released Hy4 preview, a 770B parameter MoE model with 49B active parameters and 1M context window optimized for productivity applications.
The model is fully open-sourced with weights on Hugging Face and GitHub, offered at consistent affordable pricing as a frontier option.
Benchmark visuals show Hy4 preview leading or matching top models like Claude Opus 5 across coding, agentic tasks, tool use, and math evaluations including Terminal Bench, SWE tasks, and OneMillionBench.
https://hy.tencent.ai/research/hy4-preview
via Diamond Hands Dig
Your agent can now search & fetch any webpage for 100% FREE.
Them: $7 per 1,000 searches.
Us: $0. No subscriptions, no quotas.
Humans search Google for free. Agents shouldn't have to pay either.
Made possible by Tiny_Fish and MonidHQ
Google Launches Expert Intelligence for Ebooks in Gemini Notebook
Google's new Expert Intelligence feature lets users add eligible Google Play Books ebooks to Gemini Notebook for AI-powered summaries, infographics, audio overviews, and quizzes all grounded in the book's text. People can mix these insights with personal notes, like applying management lessons to work scenarios, with over 100,000 titles from publishers such as Penguin Random House and Bloomsbury available at launch. Featured Notebooks pair bestselling authors like Michael Pollan and Steven Pinker with extra interactive elements, and Google plans expansions to subscriptions and textbooks soon.
OpenAI is terminating its partnership with the AI code editor Cursor following SpaceX's acquisition of the tool, with direct model access ending on November 12 to allow transition time,
The decision centers on trust issues, while OpenAI offers workarounds like using personal API keys or its own IDE extensions and commits to supporting diverse developer tools and open-source initiatives.
This move impacts developers who integrated GPT models into Cursor workflows, reflecting competitive dynamics in the AI tooling space between major players.
Below is my personal website which aggregates links to many of my socials as well as the various content and community that I curate. Feel free to share this link to others who you think may find this content/community useful to them
The cover image of this newsletter via generated via the Seedream 5 Pro model within the Krea tool via the following prompt
A futuristic city on an alien planet, surrounded by sand dunes and strange spires, painted in the style of Frank Frazetta. The buildings have domes with intricate details, and there's a sense that something mysterious is waiting to be discovered within them. In the background, there's a vast desert landscape with orange hues under a blue sky. A small red light shines from one building, adding contrast against the dark tones of the scene. It feels like you could find treasures or secrets here.



Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!



51-billion-param N-gram Embedding to look up a table with very little extra computation, which means the embedding



Leads every compared model on SWE-bench Pro (62.5 vs 53.4 for Claude-Opus-4.6 Max), SWE-bench

Hy4 preview is here.

.



