A large language model that can "reason".

In December, Google's DeepMind released an experimental new web browsing proxy called Mariner. In the preview demo provided by the company, Mariner seems to be running into a problem. Megha Gore, the company's product manager, had asked the agent to find her a recipe for Christmas cookies that were identical to the ones in a photo she had gried. Mariner found a recipe online and began adding ingredients to Gore's online shopping cart.

Then it stopped because it didn't know what kind of flour to choose. Gore watched as Mariner explained its steps in the chat window: "It says, 'I'm going to use the browser's 'back' button to go back to the recipe. ’”

It was an extraordinary moment. Instead of hitting a wall, the agent breaks down the task into different actions and chooses an action that might solve the problem. Figuring out that you need to hit the "back" button may sound simple, but for a mindless robot, it's nothing short of rocket science. And it worked: Mariner went back to the recipe, confirmed the type of flour, and continued to fill Gore's shopping cart with flour.

Google DeepMind is also building an experimental version of its latest large language model, Gemini 2.0, which takes this step-by-step approach to solving problems, called Gemini 2.0 Flash Thinking.

But OpenAI and Google are just the tip of the iceberg. Many companies are building large language models that use similar techniques, allowing them to better perform a range of tasks, from cooking to programming. More discussion on reasoning is expected this year (we know, we know).