- AI Horizons
- Posts
- Gemini 4 Hunts Software Flaws, Runway Trains Robots, ChatGPT Gets Dots
Gemini 4 Hunts Software Flaws, Runway Trains Robots, ChatGPT Gets Dots
PLUS: Claude Sonnet 5.5 promises cheaper tasks, Microsoft Quine prioritizes cancer experiments, and LIFT gives AI a richer memory between tokens
Welcome back to AI Horizons, your weekly guide to the latest in AI and tech for builders, leaders, and curious minds everywhere. Here’s what’s on deck:
Gemini 4 hunts vulnerabilities
Runway trains robot controllers
Sonnet promises cheaper tasks
Dots keep projects moving
Quine prioritizes cancer experiments
LIFT carries hidden state
FEATURED INSIGHT💡
Google's Gemini 4 hunts software flaws before its public debut

Google announced Gemini 4 Argon on September 30, with an output limit of one million tokens, up from 64,000. That gives a single run more room for extended reasoning and work. Access initially goes to trusted cybersecurity defenders; Axios confirms the limited rollout. Broader access for developers and consumers remains planned.
Google says Argon agents have already helped free more than 300 TiB of memory across its data centers. They are also working on large C/C++ migrations to Rust, including the Fuchsia operating system's kernel, with automated testing and human review before deployment. These examples suggest a useful target for agents: expensive engineering maintenance with measurable results. They are company-reported outcomes, so they don't establish what another team should expect.
The cybersecurity rollout makes control part of the product. Google says Argon can find, validate, and patch critical vulnerabilities, while it strengthens safeguards before wider release. For a team evaluating the model, the useful question is whether proposed fixes survive tests and review in its own codebase. A longer output allowance gives an agent more room to work; checking that work still takes engineering effort.
Blu Dot surpasses 2,000% ROAS with self-serve CTV ads
Home furniture brand Blu Dot blew up on CTV with help from Roku Ads Manager. Here’s how:
After a test campaign reached 211,000 households and achieved 1,010% ROAS, the brand went all in to promote its annual sales event. It removed age and income constraints to expand reach and shifted budget to custom audiences and retargeting, where intent was strongest.
The results speak for themselves. As Blu Dot increased their investment by 10x, ROAS jumped to 2,308% and more page-view conversions surpassed 50,000.
“For CTV campaigns, Roku has been a top performer,” said Claire Folkestad, Paid Media Strategist, Blu Dot. “Comping to our other platforms, we have seen really strong ROAS… and highly efficient CPMs, lower than any other CTV partner we've worked with.”
Using Roku Ads Manager, the campaign moved from a pilot to a permanent performance engine for the brand.
ON THE HORIZON 🌅
Runway wants video-trained AI to control real robots

Runway's Praxis-1 extends the company's video pretraining into a model that chooses actions for real robots. Early partners are testing it on their own hardware, including a robotic arm and a mobile platform. Runway plans to release the weights publicly in the coming months; they are not yet generally available.
The approach addresses a costly training problem. Robot demonstrations require physical equipment and repeated data collection, while ordinary video contains examples of hands moving, objects changing position, and tasks unfolding. Runway argues that learning those patterns gives robot controllers a useful starting point. Its announcement shows manipulation tasks and reports a policy moving between studio and kitchen environments without retraining.
The opportunity is to reuse what video models learn across different machines. That could reduce the work needed to adapt a controller to new surroundings, if the results hold beyond demonstrations. Developers should look for task-success rates, intervention requirements, and tests on unfamiliar objects when the release arrives. Partner trials can expose failures that a successful video clip leaves out.
LATEST IMPORTANT NEWS 📰
Claude Sonnet 5.5 promises cheaper tasks at the same token price
Anthropic released Claude Sonnet 5.5 on September 28, saying it generates output more than 30% faster and costs up to 30% less per task than Sonnet 5 in its tests. Standard input and output rates remain $2 and $10 per million tokens; the claimed saving comes from using fewer tokens. AWS lists the model as active in Bedrock. For frequent, well-defined jobs such as bug fixes and document production, test whether that efficiency survives your quality threshold. Track retries and review time alongside the token bill.
OpenAI's dots keep projects moving between conversations
OpenAI introduced dots on September 29: persistent agents powered by GPT-6 Astra, with their own cloud computer and access to connected apps. They can carry ongoing projects forward between conversations. Rollout begins with Pro and Business Premium users in eligible markets, while Enterprise access requires an administrator-enabled beta. OpenAI says background “proactive research” uses read-only tools; actions that change accounts or share information undergo a separate review process. For a small team, a useful first assignment could be preparing changes for review as new information arrives, with clear limits on what the agent may do.
Microsoft's Quine helps choose which cancer experiments to run
Microsoft Research introduced Quine on September 29 to connect biological models, scientific tools, literature, and laboratory feedback. In work with Broad Institute researchers, it prioritized compounds intended to change pancreatic cancer cell states; Microsoft reports that several leading candidates produced the predicted changes in laboratory assays. The potential benefit is a shorter search for experiments worth running. This is early research technology, with initial access limited to the Quine Fellows program and selected collaborations. The results do not establish a treatment benefit for patients, and the system is not intended for clinical use.
Stop typing AI prompts. Start talking.
You think 4x faster than you type. So why are you typing prompts? Wispr Flow turns your voice into ready-to-paste text inside any AI tool. Speak naturally, tangents and all, and Flow cleans it up. Available on Mac, Windows, iPhone, and Android.
FOR THE TECHNICALLY INCLINED 🛠️
LIFT gives AI a richer memory between tokens
A September 29 preprint introducing LIFT tests a way to pass internal information between generation steps. In a standard transformer, the decoded token is the route by which deep-layer information can return to earlier layers on the next step. LIFT adds a predicted state alongside that token, preserving a richer signal. During training, the target states come from an existing language model's next-token distributions and are computed in advance, so training can still run in parallel across positions.
The authors tested models with 135 million to one billion parameters and report stronger results than standard transformers at matched token budgets, while matching or beating comparisons with matched compute. That makes the training design worth watching: recurrence can help without forcing every training position to wait for the preceding one. These remain small-model research results. Tests at larger scales will need to establish whether the gains persist and whether the added inference work remains worthwhile.
AI TOOL OF THE DAY 🚀
Ideogram 4.5 offers targeted image editing for product photos, design mockups, and other visuals, with tools for changing a detail or color while limiting unwanted changes elsewhere in the image.
Some teams never seem to stop moving. They're on Attio, the agentic CRM.
It’s your always-on revenue engine: agents and workflows build pipeline, chase every buying signal, and move deals forward alongside your team.
Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?
That's all for now!
We'll catch you in the next one.
Cheers,
The AI Horizons Team
P.S. If you missed our last issue, no worries, you can check out all previous issues here!
P.P.S We value your thoughts, feedback, and questions - feel free to respond directly to this email!
... and if you enjoyed this email and would like to support our work and help us keep bringing you cutting-edge AI insights, you can donate here. Every bit makes a difference—thank you for your support!
What did you think about today's email? |



