Introduction

Welcome to today's Daily Pulse from Nicolas's AI Lab - the AI briefing for busy professionals, founders, and business owners. Around 6 minutes. Straight to what matters.

The frontier AI labs are shipping faster and cheaper. Anthropic's Claude Sonnet 5.5 lands near its top-tier model at half the price, OpenAI's DevDay teases over 20 launches, and AMD makes an $8.2B bet on Fei-Fei Li's 3D world models. At the same time, researchers and economists are warning that self-improving AI and mass adoption of personal agents carry real risks.

Today at a Glance

  • ⚡ Claude Sonnet 5.5 matches near-Opus performance at roughly half the cost and 30% faster.

  • 🌐 AMD buys World Labs for $8.2B, with Fei-Fei Li joining as chief scientist.

  • 🚨 A new paper warns of an intelligence explosion as AI takes on more of its own R&D.

Major AI News

Major AI News

Claude Sonnet 5.5 nears top-tier performance at half the price

Anthropic released Claude Sonnet 5.5, a mid-tier model that uses fewer tokens to do the same work. It claims up to 30% lower cost and 30% faster responses than the prior version, with a jump on Terminal-Bench 4.0 from 10.3% to 70.6%. It also improves image and chart interpretation and handles up to 1 million tokens of context. It arrives as a drop-in replacement for the earlier model.

Why it matters: It shows models getting cheaper and smarter at once, which lowers the cost of running agents and coding workflows at scale.

  • Same capability for roughly 30% less cost per task.

  • Strong gains on coding benchmarks, rivalling higher-tier models.

AMD acquires World Labs for $8.2B

AMD agreed to buy World Labs, founded by Fei-Fei Li, in an all-stock deal worth $8.2B. The startup builds models that generate editable 3D worlds and has worked to optimise them on AMD hardware. Li will become chief scientist and report to CEO Lisa Su. Its first product, Marble, turns text, images or video into 3D spaces, with a newer tool called Atlas in testing.

Why it matters: It gives AMD direct access to world-generating AI expertise and tighter hardware-software integration as it competes for AI market share.

  • AMD gains 3D world-model research and a marquee AI scientist.

  • Signals hardware makers buying software talent to differentiate.

Source: AMD

AI leaders warn of a self-improving intelligence explosion

A new paper co-authored by industry experts warns that AI systems iteratively improving themselves could compress years of progress into months or weeks. Anthropic reports AI now handles 26% of its R&D tasks, up from 1% in March. The authors urge slowing the pace of capability gains, embedding external auditors and monitoring for runaway growth. They raise extreme scenarios including human marginalisation if self-improvement goes unchecked.

Why it matters: Labs appear more aligned on safety even as nations race to accelerate AI, setting up a tension that will shape policy.

  • AI's share of Anthropic R&D rose from 1% to 26% since March.

  • Authors call for external audits and slower capability advances.

Source: Anthropic

Fun AI Topics

Fun AI News

OpenAI launches dots, always-on agents inside ChatGPT

OpenAI launched dots at DevDay on 29 September. They are always-on agents that live inside ChatGPT and keep working after you close the chat. Each dot runs on GPT-6 Astra, gets its own cloud computer and connects to over 4,000 apps through plugins. You can name yours, give it a look, and reach it in ChatGPT, Slack or Teams. It is rolling out to Pro and Business Premium users, one day after xAI launched Grok Team Bots.

Why it is interesting: The two biggest agent launches of the week landed a day apart, and both are selling the same thing: an AI coworker that runs in the background and checks in before anything important goes out.

  • Dots bring finished work back for your approval before it goes live.

  • Included for Pro and Business Premium users at no extra charge.

Source: OpenAI

Personal agents book too many dinners

As people deploy agents to negotiate bills, cancel subscriptions and make reservations, some agents overshoot. One example had an agent repeatedly pinging reservation systems far beyond normal human usage. Minor annoyances now, but concerning if millions of agents act at once.

Why it is interesting: It shows how well-intentioned automation can create side effects that grow with scale.

  • Agents do exactly what they are told, sometimes too literally.

  • Volume at scale is the real worry, not single mishaps.

An economist warns agents could trigger bank runs

Torsten Slok of Apollo cautions that automated agents acting in unison could create systemic risk. If every household agent chases higher yields at the same time, deposits could drain rapidly from lower-yield banks. Individually useful tools could overwhelm systems not built for sudden mass migration of funds.

Why it is interesting: It reframes personal finance agents as a potential source of coordinated market instability.

  • Synchronised agent behaviour could destabilise institutions.

  • Scale turns convenience into a systemic concern.

AI Tools

  • CodeAF: A coding harness for open models that solved nearly four times as many GitHub issues as its rival at about half the cost. github.com

  • TabPFN-3.5: A prediction model accessible via a single API call, ranking top on multiple benchmarks ahead of traditional algorithms. priorlabs.ai

  • Grok Team Bots: Shared Slack bots that accumulate team knowledge while keeping private conversations separate, for Teams and Enterprise plans. x.ai

  • Manus 2.0: An expanded toolset with a video editor, game builder and automated workflows across email, Slack and calendar. manus.im

  • VerifyAX: Simulates real-world usage to detect AI agent failures before they happen, supporting quality and compliance. conscium.com

Expert Prompt

Context: Use this to pick the right AI model for a task rather than defaulting to the most expensive one. It works by planting hidden errors in a brief and seeing which model catches them with the least manual fixing.

Prompt: Here is a launch email brief. Draft the email, then flag any factual errors, unverified claims or outdated details you find and explain your reasoning. Do not fill gaps with assumptions. [paste brief]

Example use case: Run the same brief through two or three models, track which one spots an outdated launch date and an unproven time-saving claim, and default to the model that corrects mistakes fastest.

Trending AI News

ElevenLabs debuts Eleven v4 voice model

ElevenLabs launched Eleven v4, an emotionally expressive voice model built on a new architecture. It offers improved tone, pacing and emotion for conversational and generative applications.

Why it is important: More natural voice output widens the practical uses for conversational AI in products and services.

Source: ElevenLabs

Google retires Gems for Skills

Google is discontinuing its Gems feature on 17 November 2026 and replacing it with Skills. Existing Gems convert automatically, custom instructions remain, and Skills can be invoked across Gemini tasks by typing a slash in a thread.

Why it is important: It reflects a broader push toward flexible, multi-purpose AI agents with simpler interfaces.

Human-backed agents challenge on trust

A startup opened sign-ups for Fo, a personal AI agent that loops in human assistants when AI alone cannot finish a task. The company claims better trust and completion rates, though its benchmark data has not been independently verified.

Why it is important: Hybrid human-AI agents test whether reliability, not just autonomy, wins user trust.

That's it for today's Daily Pulse. Forward this to one person who wants to stay ahead of AI. See you in the next one. - Nicolas

Get the next issue