THE FUTURELESS

RADAR ·

GPT-6 Astra beats earlier models on Andon Labs' vending and drone tests

Andon Labs tested OpenAI's GPT-6 Astra on two agent benchmarks. In Vending-Bench, where each model starts with $500 and runs a vending machine over a simulated year, Astra averaged $15,515 across six runs. Claude Fable 5.1 averaged $5,422. In the Arena version, where several agents compete at the same location, Astra refused a price-fixing offer from GLM-5.3. According to Andon Labs, Fable 5.1 joined an illegal price-fixing arrangement.

In Drone-Bench, models write code that lets a cheap drone find and follow a specific person in an office. Astra became the first model whose best attempts beat the reference solution on all five subtasks. An average run has only a 2.8 percent chance of passing all five steps in sequence. The team projects a frontier model could solve all five tasks in a single attempt by the first quarter of 2027.

Source: The Decoder

← Back to the radar