# This Week in All Things AI - Week 35-2026

*Sunday 23rd August 2026 to Saturday 29th August 2026*

By [This Week in All Things AI](https://paragraph.com/@twiata) · 2026-08-30

---

This Week in All Things AI covers key developments in models, agents, tools, infrastructure, and policy curated from discussions in the [**All Things AI Telegram group.**](https://t.me/+PnCnwhgH8V4yMTFl)

If you follow AI for work, research, investing, or just to understand where the technology is heading, this [**weekly brief**](https://paragraph.com/@twiata) is a concise way to scan the most important launches, risks, and resources in a few focused minutes.

The week of 23rd to 29th August 2026 saw Chinese AI labs push the open-weight frontier with a cluster of multimodal, long-context MoE releases. [Z.ai](http://Z.ai) launched GLM-5.3-Flash—previously previewed as Ox Alpha—while Alibaba introduced Qwen3.8-Flash and Tencent released Hy4 preview; all target coding, tool use, and agentic productivity with million-token-scale context windows and aggressive pricing. Video generation also advanced with Wan 3.0 becoming available on ArtArch, while Google expanded Gemini Notebook with ebook-grounded “Expert Intelligence.”

AI infrastructure and agent tooling were equally prominent: OpenAI reported early efficiency and latency gains from its Jalapeño inference chip, Nvidia’s Groq 3 LPX entered production, and Perplexity launched Portable Computer for fully local agent execution on NVIDIA DGX Spark. Developer discussions focused on making agents more capable and interoperable, from external search/fetch skills and free agent web access to Shopify CEO Tobi Lütke’s criticism of proprietary instruction-file conventions; OpenAI’s decision to end Cursor’s direct model access after its SpaceX acquisition further underscored intensifying competition across the coding-agent stack.

The sections that follow walk through these items day by day, with short context and links so you can dive deeper into the pieces most relevant to your work or interest

* * *

Sunday 23rd August 2026
-----------------------

Monday 24th August 2026
-----------------------

via [David An](https://t.me/davidkendo)

In our latest Provocation Talk, I speak with Advait, from [Opengradient.AI](http://Opengradient.AI) about the challenging topic of AI Privacy

[https://www.linkedin.com/feed/update/urn:li:activity:7497552042186403840/](https://www.linkedin.com/feed/update/urn:li:activity:7497552042186403840/)

I created the below public Notion page to articulate the benefits of adding external search/fetch providers in an AI harness and to suggest a prompt prefix for users to incorporate into prompts that utilise some of these providers.

These approaches can be useful if you are looking to take advantage of specialized data that external search providers can bring to the table within your AI harness.

The mechanism described on the below Notion page can be useful for GTM intelligence , lead-generation, finance or investment research, competitive intelligence or market-mapping workflows, among others.

The below page is best viewed on a desktop or laptop 

I humbly welcome your comments.

Tuesday 25th August 2026
------------------------

I updated the above Notion page to include a skill sub-directory that you can drop into your skill folder as well as some snippets to augment into [AGENTS.md](http://AGENTS.md) / [CLAUDE.md](http://CLAUDE.md) and provides some guidance on invoking one or more of these providers in a prompt.   

So some upfront investment in creating accounts with these providers, crediting balances beyond what they provide initially  and integrating them in your harness but subsequently the skill file as well as some lines in AGENTS.md/CLAUDE.md should make this very easy to incorporate external search providers in a prompt

Nvidia says its inference accelerator Groq 3 LPX has entered full production, with Nebius signing on as the first customer, and SpaceX will deploy Vera CPUs

[![User Avatar](https://storage.googleapis.com/papyrus_images/02c2ae394b22796768f2395654a30ac3daabac86e3488f2ee586110021cb296d.jpg)](https://twitter.com/GroqLLC)

[Groq](https://twitter.com/GroqLLC)

[@GroqLLC](https://twitter.com/GroqLLC)

[](https://twitter.com/GroqLLC/status/2091908837305663688)

We are thrilled to announce that Groq will be among the first adopters of NVIDIA Groq 3 LPX, deploying it alongside NVIDIA Vera Rubin NVL72 in our purpose-built AI inference Cloud. Groq is working with Dell Technologies to deploy NVIDIA Groq 3 LPX.  
  
When Groq brings NVIDIA Groq 3[

![](https://storage.googleapis.com/papyrus_images/52fcdbe581e5de29bd41dd75bfb7b82a175927616e5a884171a0cc691d6029e6.png)

groq.com

Groq Among the First to Bring NVIDIA Groq 3 LPX and Vera Rubin NVL72 to Market
------------------------------------------------------------------------------

Groq is the premier neocloud for fast inference. One fully integrated platform for infrastructure, inference, and control. Millions of developers run trillions of tokens on Groq every week.





](https://t.co/YYY19qYC58)

[303](https://twitter.com/GroqLLC/status/2091908837305663688)[

3:21 PM • Aug 24, 2026

](https://twitter.com/GroqLLC/status/2091908837305663688)

[![User Avatar](https://storage.googleapis.com/papyrus_images/5848bccca58719b811c7fb02b5f45961d7b1ebf8c07f3b15d4288a08097885c5.jpg)](https://twitter.com/elonmusk)

[Elon Musk](https://twitter.com/elonmusk)

[@elonmusk](https://twitter.com/elonmusk)

[](https://twitter.com/elonmusk/status/2091939113008238838)

SpaceX, in partnership with Nvidia, has designed a space-optimized Vera Rubin NVL72 system for launch to orbit in Q4 next year, with significant scale in 2028

[![User Avatar](https://storage.googleapis.com/papyrus_images/4667338bf7dfb2cf9a7c5a37d540656fb58469885ae88bf579e0a39acd93f600.jpg)](https://twitter.com/nvidia)

[NVIDIA](https://twitter.com/nvidia)

[@nvidia](https://twitter.com/nvidia)

[](https://twitter.com/nvidia/status/2091920680317046847)

The first CPU built for agents is going to work at scale.  
  
[@SpaceX](https://twitter.com/SpaceX) is deploying NVIDIA Vera to accelerate the orchestration, code execution, and data processing that powers its next generation of agentic AI — keeping GPUs fed and agents acting fast.  
  
From gigawatt AI factories

![](https://storage.googleapis.com/papyrus_images/d4500e4ced9abc65d533e2665d2957c1283c3a091de9d7ec45ba42699dd72d23.jpg)

[34.3K](https://twitter.com/elonmusk/status/2091939113008238838)[

5:22 PM • Aug 24, 2026

](https://twitter.com/elonmusk/status/2091939113008238838)

Shopify CEO Tobi Lütke criticizes Anthropic's Claude Code for ignoring the emerging open standard [AGENTS.md](http://AGENTS.md) and .agents/skills directory, sticking only to its proprietary [CLAUDE.md](http://CLAUDE.md) file. 

This creates "split brain" inconsistencies in large teams where developers use multiple AI coding tools, requiring workarounds like symlinks or @imports 

 that add unnecessary friction. 

The post highlights growing industry pressure for AI agents to adopt shared, vendor-neutral conventions to streamline codebase instructions across tools like Cursor, Codex, and Copilot.

[![User Avatar](https://storage.googleapis.com/papyrus_images/09f2bec19b5ce99a506e82755d984a92ab82ce9a76887a9e9e09550c12b2502d.jpg)](https://twitter.com/tobi)

[tobi lutke](https://twitter.com/tobi)

[@tobi](https://twitter.com/tobi)

[](https://twitter.com/tobi/status/2092259436538495186)

I’m thinking about banning Claude code at Shopify until they change their mind and read AGENTS.md and .agents/skills etc.  
  
Insisting on only reading CLAUDE.md sometimes leads to split brain problems when different team members use different tools. Just unnecessary.

[19.3K](https://twitter.com/tobi/status/2092259436538495186)[

2:34 PM • Aug 25, 2026

](https://twitter.com/tobi/status/2092259436538495186)

PS:   This is also my pet peeve with Anthropic that their harness will read only in their directories and filename.

Wednesday 26th August 2026
--------------------------

OpenAI's post shares initial test results for Jalapeño, its first custom AI inference chip, highlighting gains in power efficiency and response speed by delivering higher throughput and lower latency simultaneously. 

Performance benchmarks show Jalapeño achieving 1.5-1.9x more AI work per watt and 1.7-3.6x lower end-to-end latency compared to prior systems across large models like 120B, 670B, and 1T parameter sizes. 

Deployment in OpenAI's infrastructure is planned by the end of 2026, as the first step in a multi-generation hardware roadmap with Gen 2 already in development to further boost efficiency for products like ChatGPT.

[![User Avatar](https://storage.googleapis.com/papyrus_images/029dab812e4c272f2eda74f86a3f50a4047d80160e62cb1c72bc4536a142b9b9.jpg)](https://twitter.com/OpenAI)

[OpenAI](https://twitter.com/OpenAI)

[@OpenAI](https://twitter.com/OpenAI)

[](https://twitter.com/OpenAI/status/2092300846675505602)

Since announcing Jalapeño, our first custom inference chip, we’ve been testing it and the system around it.  
  
The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without

![](https://pbs.twimg.com/amplify_video_thumb/2092299952433061888/img/w27YXIvgGWvo4X_E.jpg)

[14K](https://twitter.com/OpenAI/status/2092300846675505602)[

5:19 PM • Aug 25, 2026

](https://twitter.com/OpenAI/status/2092300846675505602)

[![User Avatar](https://storage.googleapis.com/papyrus_images/3332819c227238b93a75c1ee04cf4aed8460a455a8cc8a1d9ed33361d03db012.jpg)](https://twitter.com/TokenGremlin)

[Token Gremlin](https://twitter.com/TokenGremlin)

[@TokenGremlin](https://twitter.com/TokenGremlin)

[](https://twitter.com/TokenGremlin/status/2092342432679461125)

OpenAI casually went from:  
  
“we make AI models”  
  
to:  
  
“we design the models, the serving stack, the kernels, the memory system, the network... and now the chip too.”  
  
OpenAI’s first custom inference chip, Jalapeño, is showing:  
  
→ 3.6× lower latency  
→ 4.1× higher interactive

[![User Avatar](https://storage.googleapis.com/papyrus_images/029dab812e4c272f2eda74f86a3f50a4047d80160e62cb1c72bc4536a142b9b9.jpg)](https://twitter.com/OpenAI)

[OpenAI](https://twitter.com/OpenAI)

[@OpenAI](https://twitter.com/OpenAI)

[](https://twitter.com/OpenAI/status/2092300846675505602)

Since announcing Jalapeño, our first custom inference chip, we’ve been testing it and the system around it.  
  
The results show a major advance: more intelligence from every watt and faster responses, delivering both higher throughput and lower latency in one architecture without

![](https://pbs.twimg.com/amplify_video_thumb/2092299952433061888/img/w27YXIvgGWvo4X_E.jpg)

[80](https://twitter.com/TokenGremlin/status/2092342432679461125)[

8:04 PM • Aug 25, 2026

](https://twitter.com/TokenGremlin/status/2092342432679461125)

Some snippets on SemiAnalysis longish article on Jalapeno

> In general first generation chips are not competitive, but OpenAI bucks the trend by being industry leading and beating every Nvidia, AMD, and Google chip we have been able to test on multiple top open source models. OpenAI does this with extreme hardware software codesign. Surprisingly, OpenAI is not over specialization on any specific part of model inference, but instead by focusing on being a general chip that delivers high performance in all scenarios.
> 
> Everyone says that OpenAI’s chip is specialized for OpenAI models, but that’s wrong, OpenAI made a generalized chip for AI inference.

Dwarkesh Patel shares his latest podcast with Dylan Patel of SemiAnalysis, analyzing AI lab economics including the projected shift of compute from inference to training as recursive self-improvement nears. 

The discussion forecasts Anthropic and OpenAI controlling most global usable FLOPs within years by outbidding others through superior monetization, driven by economies of scale, compute scarcity, and continual learning. 

They examine risks from over $10T in AI capex by decade's end potentially triggering sovereign debt crises via elevated interest rates, hyperscaler borrowing, and economic fallout for non-AI sectors and countries.

[![User Avatar](https://storage.googleapis.com/papyrus_images/ab12e74ab4a4d2c8eb0522f30070c8e8588a54bb64c3c98e7166ca33336b8036.jpg)](https://twitter.com/dwarkesh_sp)

[Dwarkesh Patel](https://twitter.com/dwarkesh_sp)

[@dwarkesh\_sp](https://twitter.com/dwarkesh_sp)

[](https://twitter.com/dwarkesh_sp/status/2092280377255551320)

Had a lot of fun chatting again with my twin brother [@dylan522p](https://twitter.com/dylan522p)  
  
We went through lab economics over the next few years - the shift from inference to training as RSI draws near; and how Anthropic and OpenAI are on track to control most of the world’s usable FLOPs within the next

![](https://pbs.twimg.com/media/HQlF9bCbYAAUQMo.jpg)

[1,928](https://twitter.com/dwarkesh_sp/status/2092280377255551320)[

3:58 PM • Aug 25, 2026

](https://twitter.com/dwarkesh_sp/status/2092280377255551320)

[![](https://paragraph.com/editor/youtube/play.png)](https://www.youtube.com/watch?v=aV26V1UvkJw)

Perplexity Launches Fully Local AI Agent on NVIDIA DGX Spark

Perplexity released Portable Computer, a local version of its AI agent platform that handles orchestrator, subagents, and tasks on NVIDIA DGX Spark without cloud reliance for core functions. It starts tasks locally by default and only seeks user approval for cloud use on complex reasoning, ensuring privacy by flagging personal data and avoiding direct file access. Available now for Pro and Max subscribers with compatible hardware, it excels in benchmarks like 85.4% on real-world tasks and supports app connectors for Gmail, Slack, and GitHub. NVIDIA's Jensen Huang even gifted a DGX Station to the team after an early demo, fueling excitement for private, affordable AI.

[![User Avatar](https://storage.googleapis.com/papyrus_images/e08d2690b4212e9790d6fa3fdc9353c01ef6436bab5b299cf996d52c0b685bfc.jpg)](https://twitter.com/perplexity_ai)

[Perplexity](https://twitter.com/perplexity_ai)

[@perplexity\_ai](https://twitter.com/perplexity_ai)

[](https://twitter.com/perplexity_ai/status/2092268362386780270)

Today we’re launching Portable Computer on @NVIDIA DGX Spark.  
  
Portable Computer is a fully local version of Perplexity Computer, where the entire runtime: orchestrator LLM, subagent LLM, agent harness all run on your local hardware. No cloud dependency.

![](https://pbs.twimg.com/amplify_video_thumb/2092268294791393281/img/6XoQIgKV2-kSgEmw.jpg)

[4,846](https://twitter.com/perplexity_ai/status/2092268362386780270)[

3:10 PM • Aug 25, 2026

](https://twitter.com/perplexity_ai/status/2092268362386780270)

"Even under the best circumstances, the transition to this new AI era will be one of the most turbulent times in human history," Bill Gates writes in a new, 12-page 5,784 word essay.

He says that “we are not preparing adequately”, calling for a regulatory framework

He also says tech executives are privately “very worried” about AI disruption but publicly downplay the risks to protect fundraising and planned IPOs

Looks like it's going to be a China open-weights night/US early morning with Alibaba's Qwen3.8-Flash-Next as well as GLM-5.3-Flash which was unmasked as the stealth model 0xalpha

[![User Avatar](https://storage.googleapis.com/papyrus_images/63cac069817eb46d92a6879de6375c225392fa290e86a5e49c9db88333c4e7f4.jpg)](https://twitter.com/Alibaba_Qwen)

[Qwen](https://twitter.com/Alibaba_Qwen)

[@Alibaba\_Qwen](https://twitter.com/Alibaba_Qwen)

[](https://twitter.com/Alibaba_Qwen/status/2092591393424515114)

![⚡](https://abs-0.twimg.com/emoji/v2/72x72/26a1.png)Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!  
  
The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens.  
  
125B parameters + 51B N-gram

![](https://storage.googleapis.com/papyrus_images/05209ac84047d77e611d4dde5ddbd2eecb7e22bd7db871365220ee80f2a1235a.jpg)

[7,662](https://twitter.com/Alibaba_Qwen/status/2092591393424515114)[

12:34 PM • Aug 26, 2026

](https://twitter.com/Alibaba_Qwen/status/2092591393424515114)

[https://x.com/i/trending/2092526498800718248](https://x.com/i/trending/2092526498800718248)

[Z.ai](http://Z.ai) announces GLM-5.3-Flash, a 320B total / 18B active parameter natively multimodal model with 1M-token context, released under MIT license and optimized to run on Chinese AI chips after preview as Ox Alpha. 

The model offers strong coding and agentic performance at low cost, with API pricing of $0.15 input / $0.50 output per million tokens, outperforming GLM-5.2 across benchmarks like DeepSWE and AutomationBench.  **50% discount till Sep 9th via OpenRouter** 

Architectural changes including hybrid sparse-linear attention and specialized pre-training enable efficient scaling, achieving results near Claude Opus 4.8 on select tasks while reducing compute demands significantly.

From OpenRouter

> We expect many providers to onboard this model throughout the week

[![User Avatar](https://storage.googleapis.com/papyrus_images/0da4c56e3f4d747a9cbe7764979e4bf5ba04aae54f1a5d7152070023bfd0a598.jpg)](https://twitter.com/Zai_org)

[Z.ai](https://twitter.com/Zai_org)

[@Zai\_org](https://twitter.com/Zai_org)

[](https://twitter.com/Zai_org/status/2092616204787626030)

Introducing GLM-5.3-Flash  
  
\- Leading capabilities at a highly competitive price  
\- Natively multimodal with a 1M-token context window  
\- A 320B-A18B model released under the MIT License  
\- Previously previewed as Ox Alpha, running entirely on Chinese AI chips  
  
Blog:

![](https://storage.googleapis.com/papyrus_images/3d3a079692ebdfa67cf101e20dc93a8f0c6738be45d9373fa0da9322a165eac2.jpg)

[23.5K](https://twitter.com/Zai_org/status/2092616204787626030)[

2:12 PM • Aug 26, 2026

](https://twitter.com/Zai_org/status/2092616204787626030)

[![User Avatar](https://storage.googleapis.com/papyrus_images/fcb57ac531b7bcea96bede3d02b1c04020bcbc6e11aa5582f8cc8c19d7077098.jpg)](https://twitter.com/OpenRouter)

[OpenRouter](https://twitter.com/OpenRouter)

[@OpenRouter](https://twitter.com/OpenRouter)

[](https://twitter.com/OpenRouter/status/2092616792053338229)

1M-token context. 131K max output. \`max\` reasoning by default.  
  
Launch pricing via [@Zai\_org](https://twitter.com/Zai_org) is 50% off through Sep 9 at 16:00 UTC:  
  
$0.075/M input  
$0.25/M output  
$0.015/M cached input  
  
Then $0.15/M, $0.50/M, and $0.03/M.

[116](https://twitter.com/OpenRouter/status/2092616792053338229)[

2:14 PM • Aug 26, 2026

](https://twitter.com/OpenRouter/status/2092616792053338229)

[![User Avatar](https://storage.googleapis.com/papyrus_images/17ec845d08675582e04ff0c0ad3d516408a88c5ca5c089c9d4fa90784a1c8500.jpg)](https://twitter.com/zcode_ai)

[ZCode](https://twitter.com/zcode_ai)

[@zcode\_ai](https://twitter.com/zcode_ai)

[](https://twitter.com/zcode_ai/status/2092635718766215590)

GLM-5.3-Flash is live in ZCode today ![🎉](https://abs-0.twimg.com/emoji/v2/72x72/1f389.png)  
  
The first natively multimodal GLM-5 model: it outperforms GLM-5.2 and approaches Claude Opus 4.8 on coding and agentic tasks, at 1/10 the cost.  
  
Why use it in ZCode:  
• 3× the quota of GLM-5.3 on Coding Plan, plus an extra 1.5× only in[

![](https://storage.googleapis.com/papyrus_images/bb067fc04f8d512ab8ea6b93d18338aae8ce00cda466ffad36daf01d6c7e108f.png)

zcode.z.ai

ZCode | GLM-5.3 官方 Harness
--------------------------

ZCode 将 GLM-5.3 与领先的 AI 编程 Agent 和你的现有工具链结合，帮助你顺畅完成规划、编码、评审与上线。





](https://t.co/91cnk4pSBh)

[1,025](https://twitter.com/zcode_ai/status/2092635718766215590)[

3:30 PM • Aug 26, 2026

](https://twitter.com/zcode_ai/status/2092635718766215590)

[![User Avatar](https://storage.googleapis.com/papyrus_images/51db31a2ffc8cf9397bcc6fe71f8af7338664b4c3dced759d79bc3bcc2d9ec3a.jpg)](https://twitter.com/ArtificialAnlys)

[Artificial Analysis](https://twitter.com/ArtificialAnlys)

[@ArtificialAnlys](https://twitter.com/ArtificialAnlys)

[](https://twitter.com/ArtificialAnlys/status/2092663573021606119)

GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index. At $0.09 Cost per Task, it sits comfortably on the Intelligence vs. Cost per Task Pareto frontier  
  
[@Zai\_org](https://twitter.com/Zai_org) has released GLM-5.3-Flash, a smaller and cheaper sibling to GLM-5.3 at 320B total parameters and

![](https://storage.googleapis.com/papyrus_images/8f0566a6d86ddf23833e46f78d52c63d9685393843e2043457e073ddbf0abe34.jpg)

[1,941](https://twitter.com/ArtificialAnlys/status/2092663573021606119)[

5:20 PM • Aug 26, 2026

](https://twitter.com/ArtificialAnlys/status/2092663573021606119)

via [Alba Chung](https://t.me/Cha_chung)

Hi this is Alba with Alibaba ☁ 

Wan 3.0 is now on ArtArch AI

→ Reality-grade rendering

→ Cinematic motion

→ Stronger scene consistency

Limited time: Aug 25–Sep 7: 100 FREE generations for every user. 

No duration limit.

Start creating:

[![User Avatar](https://storage.googleapis.com/papyrus_images/7012c996b113ec131ab221feb62e95f28742cd1157562fed5891fe5357d49a14.jpg)](https://twitter.com/ArtArchAI)

[ArtArch.Ai](https://twitter.com/ArtArchAI)

[@ArtArchAI](https://twitter.com/ArtArchAI)

[](https://twitter.com/ArtArchAI/status/2092265835998089243)

Wan 3.0 is now on ArtArch.  
→ Reality-grade rendering  
→ Cinematic motion  
→ Stronger scene consistency  
  
Limited time, Aug 25–Sep 7: 100 FREE generations for every user. No duration limit.  
Start creating: [artarch.ai/models/wan-3.0](https://t.co/eGyDCnXKgv)  
  
[#GenerativeAI](https://twitter.com/hashtag/GenerativeAI) [#Wan3](https://twitter.com/hashtag/Wan3) [#AIFilmmaking](https://twitter.com/hashtag/AIFilmmaking) [#ArtArch](https://twitter.com/hashtag/ArtArch)

![](https://pbs.twimg.com/amplify_video_thumb/2092262067965485057/img/rMVwCIYUne42S5JP.jpg)

[10](https://twitter.com/ArtArchAI/status/2092265835998089243)[

3:00 PM • Aug 25, 2026

](https://twitter.com/ArtArchAI/status/2092265835998089243)

Thursday 27th August 2026
-------------------------

Curating some links about Qwen-3.8-Flash-Next which is now available via OpenRouter at 0.16/0.47/0.016 so priced competitively with GLM-5.3-Flash.   

Shall be interesting when other providers join the fray in serving this model.

[![User Avatar](https://storage.googleapis.com/papyrus_images/f1a681bf10b398bb95bb0ebb1d1247d70771f99b045998bb68f0bd1fae6cdc9c.jpg)](https://twitter.com/SemiAnalysis_)

[SemiAnalysis](https://twitter.com/SemiAnalysis_)

[@SemiAnalysis\_](https://twitter.com/SemiAnalysis_)

[](https://twitter.com/SemiAnalysis_/status/2092688580111974648)

Congrats to [@Alibaba\_Qwen](https://twitter.com/Alibaba_Qwen) on the release of Qwen3.8-Flash-Next, using the same architecture innovations as their upcoming Qwen4 model! Such innovations include:  
  
![🟠](https://abs-0.twimg.com/emoji/v2/72x72/1f7e0.png) 51-billion-param N-gram Embedding to look up a table with very little extra computation, which means the embedding

![](https://storage.googleapis.com/papyrus_images/23d2e329e557525d6aea562746568baa5d0741dfd8f3e0632f031d40610929ad.jpg)

[326](https://twitter.com/SemiAnalysis_/status/2092688580111974648)[

7:00 PM • Aug 26, 2026

](https://twitter.com/SemiAnalysis_/status/2092688580111974648)

[![User Avatar](https://storage.googleapis.com/papyrus_images/49a5244b32b42d338bfc35b137cb7183e46f74dd4dbf225d258c13eb31cc8ee1.jpg)](https://twitter.com/sushsrinivasan)

[Sushaanth Srinivasan](https://twitter.com/sushsrinivasan)

[@sushsrinivasan](https://twitter.com/sushsrinivasan)

[](https://twitter.com/sushsrinivasan/status/2092724340295164287)

Qwen4 spends 51.2B of its parameter budget on memory instead of computation.  
  
Qwen3.8-Flash-Next shipped this morning, and its config calls it Qwen4. Of its 180B parameters, 51.2B are a hash table. That's Engram, from DeepSeek's paper in January, implemented down to the hash

![](https://storage.googleapis.com/papyrus_images/3aaf1e9d26e61e160c6be092600723aee74f351a9f723f705e3fbb61ff32ca40.jpg)

[20](https://twitter.com/sushsrinivasan/status/2092724340295164287)[

9:22 PM • Aug 26, 2026

](https://twitter.com/sushsrinivasan/status/2092724340295164287)

[![User Avatar](https://storage.googleapis.com/papyrus_images/b9faad294c7189f84591290d6c459ac9b78dd18dbd99db88d44fb12c2535c944.jpg)](https://twitter.com/ModelScope2022)

[ModelScope](https://twitter.com/ModelScope2022)

[@ModelScope2022](https://twitter.com/ModelScope2022)

[](https://twitter.com/ModelScope2022/status/2092590458711232757)

Qwen3.8-Flash-Next is here! An open-weight multimodal MoE built on a brand new architecture, with native 256K context extendable to 1M via YaRN. ![🤖](https://abs-0.twimg.com/emoji/v2/72x72/1f916.png)[modelscope.ai/collections/Qw…](https://t.co/bzT3MaDDwe)  
  
![🏆](https://abs-0.twimg.com/emoji/v2/72x72/1f3c6.png) Leads every compared model on SWE-bench Pro (62.5 vs 53.4 for Claude-Opus-4.6 Max), SWE-bench

![](https://storage.googleapis.com/papyrus_images/34ef44450f1ab6489deb6a515492a8a8e4e19933a11585f69f2b39efbc692c49.jpg)

[1,063](https://twitter.com/ModelScope2022/status/2092590458711232757)[

12:30 PM • Aug 26, 2026

](https://twitter.com/ModelScope2022/status/2092590458711232757)

An impressive technical breakdown by [Ryan Smith](https://x.com/RyanSmithAT) on Sambanova's SN50 RDU

Nvidia Acquires Hugging Face for $12.9 Billion

The agreement gives Nvidia control of Hugging Face, often called the GitHub for AI, where millions of models and datasets are shared by researchers and companies worldwide. Reports from The Information, Bloomberg, and Business Insider pegged the valuation around $13 billion, down slightly from earlier talks, following Nvidia's prior investment in the startup's 2023 funding round at $4.5 billion. The move expands Nvidia beyond chips into the open-source ecosystem that powers AI on its GPUs, even as Hugging Face's revenue sits around $150 million annually. It highlights hardware giants racing to own more of the AI stack amid strong demand.

[https://x.com/i/trending/2092825795001790544](https://x.com/i/trending/2092825795001790544)

Friday 28th August 2026
-----------------------

For the video editors in da house via Shamir Allibhai of [Hey Eddie](https://www.heyeddie.ai/) fame who is tentatively my upcoming guest in my [conversation series.](https://yusuf-goolamabbas-53.notion.site/Recordings-of-various-talks-I-ve-organised-99de8a60fe1c4150968ba8fdc3764321?source=copy_link)    Luma link for folks to register for the virtual event shall be posted once dates/times are finalised with Shamir

The iOS version of Eddie AI is now available

[https://apps.apple.com/app/eddie-ai-video-editor/id6792122289](https://apps.apple.com/app/eddie-ai-video-editor/id6792122289)

Shamir also pointed me towards this completely free product he built  [https://99.life/](https://99.life/)

Here's a brief blurb [about the above product](https://99.life/about)

99 is a free arena for AI video editing. You bring the clips. Two AI models each cut their own edit. You watch both without knowing which model made which, and pick the winner. Every vote moves the public scoreboard.

As to why the above product is Free

![](https://paragraph.com/editor/callout/information-icon.png)

Why it's free

No trick, just a trade. Blind votes from real people on real footage are the honest measure of an AI editor, and model makers pay for that measure. 99 owns the edits the models generate, and the evaluation data the arena produces — votes, watch behavior, and the content they judge — is used to train AI models and is licensed to the companies that build them. That deal funds the arena. You get free edits and a real say in which models win; the model makers get judged. We think that is a fair trade, and we would rather say it in bold than bury it in fine print. The details live in our [Terms](https://99.life/terms) and P[rivacy Policy.](https://99.life/privacy)

Tencent Hunyuan released Hy4 preview, a 770B parameter MoE model with 49B active parameters and 1M context window optimized for productivity applications. 

The model is fully open-sourced with weights on Hugging Face and GitHub, offered at consistent affordable pricing as a frontier option. 

Benchmark visuals show Hy4 preview leading or matching top models like Claude Opus 5 across coding, agentic tasks, tool use, and math evaluations including Terminal Bench, SWE tasks, and OneMillionBench.

[![User Avatar](https://storage.googleapis.com/papyrus_images/47df04b66d6c8f8851f8bf4676b670017ecbfff576525418fc2c5d612be7832a.jpg)](https://twitter.com/TencentHunyuan)

[Tencent Hy](https://twitter.com/TencentHunyuan)

[@TencentHunyuan](https://twitter.com/TencentHunyuan)

[](https://twitter.com/TencentHunyuan/status/2093222928720761009)

![🚀](https://abs-0.twimg.com/emoji/v2/72x72/1f680.png) Hy4 preview is here.  
  
770B, 49B active, 1M context.  
Built for productivity.  
Open source frontier.  
Consistent affordable price.  
Use it. Tell us what breaks.  
  
More on Hy blog：[hy.tencent.ai/research/hy4-p…](https://t.co/rbl1IWRk3C)  
HuggingFace：[huggingface.co/tencent/Hy4-pr…](https://t.co/mE9wevH5XR)  
Github：[github.com/Tencent-Hunyua…](https://t.co/pyl9zckpoL)

![](https://storage.googleapis.com/papyrus_images/5fb89abaeae75b53164acc445306a639e66f6167726ed728599adef0668d75e9.jpg)

[1,586](https://twitter.com/TencentHunyuan/status/2093222928720761009)[

6:23 AM • Aug 28, 2026

](https://twitter.com/TencentHunyuan/status/2093222928720761009)

[https://hy.tencent.ai/research/hy4-preview](https://hy.tencent.ai/research/hy4-preview)

via Diamond Hands Dig

Your agent can now search & fetch any webpage for 100% FREE.

Them: $7 per 1,000 searches.

Us: $0. No subscriptions, no quotas.

Humans search Google for free. Agents shouldn't have to pay either.

Made possible by Tiny\_Fish and MonidHQ  

[![User Avatar](https://storage.googleapis.com/papyrus_images/8d4b77d7ddcd53f961991d6692d3a2c8f6adceff073c8e0f74b082716f62a062.jpg)](https://twitter.com/shengkunye)

[Shengkun Ye](https://twitter.com/shengkunye)

[@shengkunye](https://twitter.com/shengkunye)

[](https://twitter.com/shengkunye/status/2093050916953903451)

We just killed Exa, Tavily, SerpAPI, and Brave.  
  
Your agent can now search & fetch any webpage for 100% FREE.  
  
Them: $7 per 1,000 searches.  
Us: $0. No subscriptions, no quotas.  
  
Humans search Google for free. Agents shouldn't have to pay either.  
  
Made possible by [@Tiny\_Fish](https://twitter.com/Tiny_Fish) and

![](https://pbs.twimg.com/amplify_video_thumb/2093050017145864192/img/4De8-Lzom2ODsbpA.jpg)

[4,540](https://twitter.com/shengkunye/status/2093050916953903451)[

7:00 PM • Aug 27, 2026

](https://twitter.com/shengkunye/status/2093050916953903451)

Google Launches Expert Intelligence for Ebooks in Gemini Notebook

Google's new Expert Intelligence feature lets users add eligible Google Play Books ebooks to Gemini Notebook for AI-powered summaries, infographics, audio overviews, and quizzes all grounded in the book's text. People can mix these insights with personal notes, like applying management lessons to work scenarios, with over 100,000 titles from publishers such as Penguin Random House and Bloomsbury available at launch. Featured Notebooks pair bestselling authors like Michael Pollan and Steven Pinker with extra interactive elements, and Google plans expansions to subscriptions and textbooks soon.

[![User Avatar](https://storage.googleapis.com/papyrus_images/59968cdd0dfa43fa749bf72bb022e3c9131016d4922199b369c8ccb733630071.png)](https://twitter.com/Gemini_Notebook)

[Gemini Notebook](https://twitter.com/Gemini_Notebook)

[@Gemini\_Notebook](https://twitter.com/Gemini_Notebook)

[](https://twitter.com/Gemini_Notebook/status/2093059543362306369)

Introducing Expert Intelligence ![✨](https://abs-0.twimg.com/emoji/v2/72x72/2728.png)  
  
A cross-[@Google](https://twitter.com/Google) initiative that helps you engage with your trusted sources, starting with eligible [@GooglePlay](https://twitter.com/GooglePlay) ebooks in Gemini Notebook. You can now combine expertise from your favorite authors with other sources and engage with your books in

![](https://pbs.twimg.com/amplify_video_thumb/2093059496046391296/img/P6WpLrOP-2KKxf3b.jpg)

[2,909](https://twitter.com/Gemini_Notebook/status/2093059543362306369)[

7:34 PM • Aug 27, 2026

](https://twitter.com/Gemini_Notebook/status/2093059543362306369)

[![User Avatar](https://storage.googleapis.com/papyrus_images/d1526f78e94a4647681d9c2204882d67e67c6298acf59d8614f577a8527fa4ce.jpg)](https://twitter.com/NewsFromGoogle)

[News from Google](https://twitter.com/NewsFromGoogle)

[@NewsFromGoogle](https://twitter.com/NewsFromGoogle)

[](https://twitter.com/NewsFromGoogle/status/2093063328989856092)

Over the past several months, we’ve been meeting with authors and publishers to better understand how AI can help people get the most out of their books ![📚](https://abs-0.twimg.com/emoji/v2/72x72/1f4da.png).  
  
Today we’re launching Expert Intelligence, a new way to get insights about your favorite titles from the authors you love,

![](https://pbs.twimg.com/amplify_video_thumb/2093063295313780736/img/eSLT_wjfJKBvUx7i.jpg)

[248](https://twitter.com/NewsFromGoogle/status/2093063328989856092)[

7:49 PM • Aug 27, 2026

](https://twitter.com/NewsFromGoogle/status/2093063328989856092)

Saturday 29th August 2026
-------------------------

OpenAI is terminating its partnership with the AI code editor Cursor following SpaceX's acquisition of the tool, with direct model access ending on November 12 to allow transition time, 

The decision centers on trust issues, while OpenAI offers workarounds like using personal API keys or its own IDE extensions and commits to supporting diverse developer tools and open-source initiatives. 

This move impacts developers who integrated GPT models into Cursor workflows, reflecting competitive dynamics in the AI tooling space between major players.

[![User Avatar](https://storage.googleapis.com/papyrus_images/b5d8d307939217b13930a49ab51622dab7841b0a05e33b3b01643b1f30fed85a.jpg)](https://twitter.com/thsottiaux)

[Tibo](https://twitter.com/thsottiaux)

[@thsottiaux](https://twitter.com/thsottiaux)

[](https://twitter.com/thsottiaux/status/2093515916076343774)

We unfortunately have decided that we cannot continue providing access to our models through Cursor and are ending our partnership. It boils down to trust and we’ve asked that this takes effect on November 12 to give you some time to plan.  
  
Many have used the GPT models through[

![](https://storage.googleapis.com/papyrus_images/ac856e45f8ea06eae619d6b50134d86e5a75ae38e35884334e8c1e5a04a1e0bd.jpg)

openai.com

Our decision on Cursor following its acquisition by SpaceX
----------------------------------------------------------

Our decision to wind down our contract providing OpenAI models to Cursor following its acquisition by SpaceX.





](https://t.co/Oj76Bkc2KX)

[![User Avatar](https://storage.googleapis.com/papyrus_images/029dab812e4c272f2eda74f86a3f50a4047d80160e62cb1c72bc4536a142b9b9.jpg)](https://twitter.com/OpenAI)

[OpenAI](https://twitter.com/OpenAI)

[@OpenAI](https://twitter.com/OpenAI)

[](https://twitter.com/OpenAI/status/2093515564786540695)

We’re ending our partnership with Cursor following its acquisition by SpaceX. Under our proposal, Cursor’s direct access to our models would end on November 12.  
  
We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care

[5,600](https://twitter.com/thsottiaux/status/2093515916076343774)[

1:47 AM • Aug 29, 2026

](https://twitter.com/thsottiaux/status/2093515916076343774)

* * *

Below is my personal website which aggregates links to many of my socials as well as the various content and community that I curate. Feel free to share this link to others who you think may find this content/community useful to them

[**https://linktr.ee/goolamabbas**](https://linktr.ee/goolamabbas)

The cover image of this newsletter via generated via the Seedream 5 Pro model within the [**Krea**](https://www.krea.ai/refer/EJQQP9FJ) tool via the following prompt

![](https://paragraph.com/editor/callout/information-icon.png)

A futuristic city on an alien planet, surrounded by sand dunes and strange spires, painted in the style of Frank Frazetta. The buildings have domes with intricate details, and there's a sense that something mysterious is waiting to be discovered within them. In the background, there's a vast desert landscape with orange hues under a blue sky. A small red light shines from one building, adding contrast against the dark tones of the scene. It feels like you could find treasures or secrets here.

---

*Originally published on [This Week in All Things AI](https://paragraph.com/@twiata/this-week-in-all-things-ai-week-35-2026)*
