• AI Horizons
  • Posts
  • OpenAI Says 10,000 AI Agents Cracked a 90-Year-Old Math Problem

OpenAI Says 10,000 AI Agents Cracked a 90-Year-Old Math Problem

PLUS: AlphaGenome maps nine billion DNA variants, Meta launches Muse, and UMG plans licensed AI remixes with ElevenLabs

In partnership with

Welcome back to AI Horizons, your weekly guide to the latest in AI and tech for builders, leaders, and curious minds everywhere. Here’s what’s on deck:

  • OpenAI proposes Navier–Stokes proof

  • AlphaGenome maps DNA variants

  • Meta launches personal Muse agents

  • OpenAI opens Agents API

  • UMG plans licensed AI remixes

  • SWE-2 targets cheaper coding

FEATURED INSIGHT💡

OpenAI’s 88-hour math sprint sparks a fight over credit

OpenAI says an internal AI system has produced a solution to the Navier–Stokes existence and smoothness problem, one of mathematics’ Millennium Prize Problems. Announced September 8, the work used roughly 10,000 concurrent agents powered by an unreleased model. The company reports about 88 hours of search followed by 17 hours of formalization and verification with GPT-6 Astra. The experiment combines an unreleased model with extensive parallel computation and human direction; it does not establish what a public chatbot can do on its own.

The released Lean formalizations describe cases where initially smooth fluid motion, subject to a smooth external force, cannot remain smooth indefinitely. In ordinary terms, the equations allow a breakdown even when the starting conditions and applied force are well behaved. The repository provides instructions for independent checking, giving mathematicians a concrete object to inspect. A useful next step is to establish that the formal statement, its assumptions, and the written argument match the problem being claimed.

The announcement also exposed a dispute over scientific credit. NYU mathematician Tristan Buckmaster told ABC that OpenAI pursued the problem after learning of his work with Levent Alpöge and questioned whether their unpublished Codex sessions had influenced it. OpenAI’s September 10 update says those prompts could not have influenced the system, including through training. For research teams, the episode makes agreements about unpublished work, attribution, and competing projects part of choosing an AI partner. The claimed mathematical advance and the disputed conduct deserve separate scrutiny.

How Jennifer Aniston’s LolaVie brand grew sales 40% with CTV ads

For its first CTV campaign, Jennifer Aniston’s DTC haircare brand LolaVie had a few non-negotiables. The campaign had to be simple. It had to demonstrate measurable impact. And it had to be full-funnel.

LolaVie used Roku Ads Manager to test and optimize creatives — reaching millions of potential customers at all stages of their purchase journeys. Roku Ads Manager helped the brand convey LolaVie’s playful voice while helping drive omnichannel sales across both ecommerce and retail touchpoints.

The campaign included an Action Ad overlay that let viewers shop directly from their TVs by clicking OK on their Roku remote. This guided them to the website to buy LolaVie products.

Discover how Roku Ads Manager helped LolaVie drive big sales and customer growth with self-serve TV ads.

The DTC beauty category is crowded. To break through, Jennifer Aniston’s brand LolaVie, worked with Roku Ads Manager to easily set up, test, and optimize CTV ad creatives. The campaign helped drive a big lift in sales and customer growth, helping LolaVie break through in the crowded beauty category.

ON THE HORIZON 🌅

AlphaGenome turns genetic variation into a searchable atlas

Google DeepMind released AlphaGenome Atlas on September 8, making precomputed predictions for roughly nine billion possible single-letter changes in human DNA accessible through a research portal. Its AlphaGenome Variant Impact score combines predictions about gene regulation with AlphaMissense’s estimates of protein effects. Researchers can rank candidate variants and inspect which predicted molecular changes contribute to a score, including disruptions to gene expression or RNA splicing.

That could help labs decide which experiments to run first. DeepMind reports that collaborators used the scores to identify a previously overlooked variant affecting DNM1, then experimentally confirmed the predicted splicing problem. This example supports a focused use: narrowing a large candidate set and proposing a mechanism worth testing. Atlas does not establish the biological consequences of every mutation it scores. Its free academic portal makes those hypotheses easier to explore without building a prediction pipeline, while broader value will depend on how reliably the rankings guide experiments across different research questions. Commercial access to Atlas is described as coming to Google Cloud soon.

LATEST IMPORTANT NEWS 📰

Meta gives Muse its own cloud computer

Meta introduced Muse on September 8 as a personal agent that can carry out tasks such as travel booking and email. It is rolling out in the United States on iOS, Android, and the web. Meta says each agent operates inside a dedicated virtual machine, with a separate Sentinel agent approving outbound actions and requesting permission when needed. A version designed to keep even Meta from accessing the VM’s contents is planned for later this year. That distinction matters when assessing the privacy protections available at launch.

OpenAI opens the Codex harness to developers

The Agents API entered public beta on September 10, giving developers OpenAI-managed orchestration for long-running agents. The service handles context compaction, tool use, and coordinated subagents, while developers choose OpenAI-hosted sandboxes, their own infrastructure, or a supported sandbox provider. OpenAI says there is no additional API fee beyond the tokens and tools used. For teams building agents, this moves part of the maintenance burden to OpenAI; application-specific tools, permissions, and success criteria still need deliberate design. Sandbox choice remains a separate decision from which model runs the work.

UMG and ElevenLabs plan licensed fan remixes

Universal Music Group announced a multiyear agreement with ElevenLabs on September 10 covering licensing and joint product development. Their planned fan platform will support remixes, mashups, and new interpretations of music from participating artists and songwriters. ElevenLabs confirms the platform is still in development and will be separate from its existing music products. For creators and music businesses, the participation requirement is central: the announcement does not make UMG’s entire catalog available for unrestricted AI use, and it does not announce a live service that fans can use today.

Your AI budget tripled. See real usage patterns with Harmonic.

AI spend is now a major P&L line item—but most teams can't show what it's producing.

Harmonic Security maps AI activity to use cases and teams, revealing real productivity, shelfware, data risk, and adoption trends across approved and unapproved tools.

Give your board the data behind the return.

FOR THE TECHNICALLY INCLINED 🛠️

SWE-2 trains reasoning effort around task cost

Cognition’s SWE-2, released September 10, builds on Kimi K3 with reinforcement learning that trains multiple reasoning-effort levels in one run. Each level applies a different cost penalty, so training rewards successful work while accounting for the resources it consumes. Cognition reports 50.0% on FrontierCode 1.1 Main, versus 50.9% for Fable 5.1, at 64% lower cost at that comparison point. FrontierCode evaluates tasks written by open-source maintainers, with mergeability as its standard. The launch table also shows a substantial weakness: SWE-2 scores 27.3% on Terminal-Bench 4, against Fable 5.1’s 55.8%. These are provider-reported comparisons, with results dependent on the benchmark and agent setup. Teams evaluating the model should compare cost per accepted change on their own repositories, including review and rework. SWE-2 is available in Devin Desktop and CLI, with Web and Fusion rolling out.

AI TOOL OF THE DAY 🚀

The Google Cloud Developer Plugin, released September 10, bundles official documentation access, Cloud skills, and guidance for authentication and gcloud operations into an installable toolkit for supported AI coding agents.

Unify Your Teams and Tech Stack With HubSpot

Connect your customer data, teams, and tools without the hassle of complex integrations or lengthy setups. One easy platform gives marketing, sales, and service teams a unified customer view and the tools to turn it into growth.

Why HubSpot and what's new

  • Generate leads and automate marketing with Marketing Hub

  • Build your pipeline and close more deals with Sales Hub

  • Scale customer support and drive retention with Service Hub

  • Keep customer data clean, connected, and actionable with one, unified platform

Join 306,000+ in over 135 countries using HubSpot to grow their businesses.

See what a more connected approach to growth can do for you and your team. Get setup quickly and start checking off your hardest tasks. 

That's all for now!

We'll catch you in the next one.

Cheers,

The AI Horizons Team

P.S. If you missed our last issue, no worries, you can check out all previous issues here!

P.P.S We value your thoughts, feedback, and questions - feel free to respond directly to this email!

... and if you enjoyed this email and would like to support our work and help us keep bringing you cutting-edge AI insights, you can donate here. Every bit makes a difference—thank you for your support!

What did you think about today's email?

Login or Subscribe to participate in polls.