Cover photo

This Week in All Things AI - Week 41-2026

Sunday 4th October 2026 to Saturday 10th October 2026

This Week in All Things AI covers key developments in models, agents, tools, infrastructure, and policy curated from discussions in the All Things AI Telegram group.

If you follow AI for work, research, investing, or just to understand where the technology is heading, this weekly brief is a concise way to scan the most important launches, risks, and resources in a few focused minutes.

The week of 4th to 10th October 2026 highlighted a growing emphasis on making AI faster and more economical through both broader capabilities and specialised models. Mistral Large 4 entered public preview with multimodal support and a million-token context window, while its weights remained forthcoming. Anthropic released Haiku 5.5 for high-volume work, with lower pricing aimed at tasks such as summarisation, classification and supporting larger agents. Microsoft-Decision-1, Perplexity’s pplx-decider-v1.1 and Celeris-1 Decision expanded the options for applications that need scored choices rather than generated prose. Google’s Nano Banana 2.1 brought reported improvements in image editing and subject consistency, while discussions of Celeris’s diffusion models and MachGen’s inference stack explored different approaches to reducing generation time and cost.

The supporting infrastructure and everyday workflows showed why model releases are only part of that progress. Richard Ho’s interview on OpenAI’s Jalapeño chip described AI-assisted design and the challenge of keeping hardware useful as workloads change, alongside the need to validate optimisations engineers may not immediately understand. Perplexity’s embedding models addressed another practical issue: finding information in text and images while retaining the context needed to interpret an answer. Google’s rollout of native Markdown editing and collaboration in Drive and Docs made a format widely used in AI workflows easier for people to work on together. Discussion of deployers’ responsibilities under the EU AI Act added a governance dimension to the week’s focus on putting these capabilities to use.

The sections that follow walk through these items day by day, with short context and links so you can dive deeper into the pieces most relevant to your work or interest


Sunday 4th October 2026

Monday 5th October 2026

🌶 AI helping design the chips that run AI—there’s plenty to unpack in Ian Cutress’s interview with Richard Ho, who leads OpenAI’s Jalapeño hardware team.

A few hooks for the AI enthusiasts here:

• Ho says models without task-specific fine-tuning helped save over 13% of die area in one optimisation example. He also describes accepting optimisations the engineers didn’t immediately understand—with rigorous validation still essential.

• Why build one accelerator for the different stages of inference, while others split them across specialised chips? His answer is about keeping an entire fleet useful as workloads change.

• Why share architectural ideas competitors might copy? OpenAI still needs GPUs—and wants those GPUs to get better.

• His provocative take on the biggest bottleneck: the supply chain still underestimates how large demand could become.

The conversation connects chip design, AI-assisted engineering and the economics of making intelligence cheaper. Worth reading—and setting aside time for the full interview.

Article and edited transcript: More Than Moore

▶️ Full video: The Man Who Built OpenAI’s First Chip

via David An

In our new Provocation Talk, we discuss some major impact of "Deployers" according to the EU AI Act. This is also relevant for any AI Deployers who are serving EU customers https://www.linkedin.com/feed/update/urn:li:activity:7512819988513132544/

Tuesday 6th October 2026

Google Workspace now supports native Markdown (.md) files in Drive and Docs, enabling preview, editing, commenting, and real-time collaboration without converting to Google Doc format.

The update addresses Markdown's growing use as a standard for AI-human workflows including specs, READMEs, task lists, and generated documentation.

Rollout has begun for Workspace customers and personal Google accounts, as shown in a screenshot of a collaborative tech spec document with comments preserved in .md.

https://workspaceupdates.googleblog.com/2026/10/preview-edit-and-collaborate-on-Markdown-files-natively-across-Drive-and-Docs.html

Wednesday 7th October 2026

Mistral Large 4 nicknamed le Chonk is in public preview.  It's an MoE model with 1.05 trillion parameters and 49 billion active on each token, a 1.6 billion-parameter vision encoder, and a 1 million-token context.  Mistral says it was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters.  Available via OpenRouter at 0.68/Mtok input and $2.09/Mtok output.  The weights are still a few weeks away.

Google just dropped Nano Banana 2.1, the next version of its image model. It runs on Gemini 3.6 Flash and is meant to be the faster, everyday option rather than the heavy Pro model.

Google reports gains over prior versions in visual design, mask-based editing, subject consistency, and overall naturalness. On DeepMind’s internal preference evaluations, the thinking variant scores ahead of both Nano Banana 2 and Nano Banana Pro, with the largest improvements in infographic factuality and multi-character consistency.

The model accepts text and image inputs, supports output up to 4K, and can ground generations with Google Search. It is available in AI Studio and the Gemini API, and is rolling out in AI Mode in Search.

Thursday 8th October 2026

via Justin Pinter

Anthropic has released Haiku 5.5 which Cheapest, fastest, and most capable small model Anthropic has shipped. Built for high-volume, cost-sensitive work — summaries, compaction, classification, database queries, subagents under Opus/Sonnet, live support, and browser use. 1M context length

Pricing

  • Prompts ≤100k ( covers ~90% of prior Haiku usage): $0.10 input / $0.50 output — 90% below Haiku 4.5.

  • Longer prompts: $0.50 / $2.50.

  • Cache reads: $0.01 (or $0.05 over 100k).

via Matthew Weigand

Perplexity has open-sourced a new collection of decision and embedding models with SOTA or near-SOTA performance in the last few weeks: 

𝗗𝗲𝗰𝗶𝘀𝗶𝗼𝗻 𝗺𝗼𝗱𝗲𝗹 

pplx-decider-v1.1 is a general-purpose 27B classifier fine-tuned from Qwen3.8-27B. It returns probabilities and confidence over a pre-defined list of options. v1.1 raises the best Hugging Face Decision Index 0.3 score to 62.8. It's served through our Decisions API and OpenRouter. 

- Model: https://lnkd.in/g8jqA4DM

- API: https://lnkd.in/g3fDbpGP

𝗠𝘂𝗹𝘁𝗶𝗺𝗼𝗱𝗮𝗹 𝗲𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴𝘀

pplx-embed-v2-late is a family of 0.6B and 9B models that create one embedding per token, for both text and images such as rendered document pages. The two sizes share one embedding space, so you can choose which is used for query vs document encoding. E.g., index documents offline with 9B and encode live queries with the cheaper and faster 0.6B. The 0.6B model matches models with 5x its active parameters on ViDoRe V3, and the 9B model scores 92.4% on MADQA. On Q2D-Web, both lead the evaluated baselines.

- Research blog: https://lnkd.in/gAMfC3kY

- Model: https://lnkd.in/gufZmNgz, https://lnkd.in/gS75krcJ

𝗖𝗼𝗻𝘁𝗲𝘅𝘁𝘂𝗮𝗹 𝗲𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴𝘀

pplx-embed-v2-context-9b-preview embeds each passage together with the document it comes from, so search can find both the answer and the context that supports it. It still produces one vector per chunk, at no extra inference cost, in 1024 or 2048 dimensions with native int8. It has the highest average nDCG@10 on ConTEB, and it leads on turbopuffer's private context-bench for answer, evidence and document retrieval. 

- Research blog: https://lnkd.in/gGzqF2d2

- Model: https://lnkd.in/g_7PGXgJ

Friday 9th October 2026

Going to share a couple of posts about some providers who I feel are pushing the envelope in diffusion models which is a different approach since most models we interact with regularly are auto-regressive [one-token ouput at a time].   

I encourage group members to try these providers and models on your own  workloads

Celeris claims to be Speedy Gonzales for both their diffusion models [Celeris-1 Magnus and Celeris-1] as well as their decision model Celeris-1 Decision

MachGen — high-performance inference for image, video, and world models.

Founded by Kismat Singh (ex-Intel AI frameworks, NVIDIA) and Manoj Krishnan. Team includes people who worked on TensorRT at NVIDIA, vLLM at Google, and large-model training at Meta. Out of stealth mid-2026.

They build a diffusion-specific stack (kernels, attention, caching, scheduling) so open models run at the same quality, much faster and cheaper — not distilled, same weights. 

Claimed numbers: HiDream ~1s vs 6s, Flux ~1.5s vs 6.2s, LTX 2.3 ~11s vs 67s, MiniMax H3 up to ~15x. 

Roughly 4–6x lower latency and 2–4x lower cost vs standard deployments.

Time to plonk in a few dollars of credits at both these providers and kick their respective tires and then figure out a way to track the cumulative spend with these 'few dollars' added at a gazillion providers 😢

Saturday 10th October 2026

Microsoft-Decision-1 is out: a decision-scoring model, not a generative LLM.

Post-trained on Qwen3.5-9B, 

Microsoft claims top accuracy on a 36-benchmark suite (~150k held-out questions), and the fastest latency measured: 4.5× quicker than Quyet-1.0-Large and 35× quicker than GPT-6 Sol. Pricing is $0.042 per million input tokens, output free.

Available now in Microsoft Foundry; on OpenRouter as well. First in a planned family; later versions will be rebased on MAI and OpenAI models.


Below is my personal website which aggregates links to many of my socials as well as the various content and community that I curate. Feel free to share this link to others who you think may find this content/community useful to them

https://linktr.ee/goolamabbas

The cover image of this newsletter via generated via the Nano Banana 2.1 model within the Krea tool via the following prompt

"Spirits of the Land": Envision a sweeping landscape where nature spirits come alive. Trees bend and twist into fantastical shapes, their leaves rustling with whispers. Wild animals roam free, their pelts emblazoned with shimmering patterns that reflect the colors of the land. And in the distance, a rainbow arcs across the sky, connecting the mortal realm to the Otherworld. insanely detailed, dynamic lighting, clean reflections, Vivid colors, Ultra Precision, Sharp Edges, maximum Quality