Safety overview: GPT-6 Astra
OpenAI's own safety overview: first model to reach Critical cybersecurity capability under the Preparedness Framework.
5 items · 2 sources · 2 days
Operational story trace
Follow in this browser to see new updates on your Live feed.
Latest change
A day-two recap frames Astra as OpenAI's biggest LLM launch yet: new SOTA computer-use and coding scores, 2.5x pricier per token but cheaper per completed task, and — flagged as a real caveat — reduced monitorability versus prior models.
OpenAI shipped GPT-6 Astra on September 3 as its most capable broadly deployed model, launch-day case studies showing agentic coding and document-review gains at Playco and Legora. It's also the first OpenAI model to hit the Critical tier for cybersecurity capability under the company's own Preparedness Framework.
Arc
OpenAI's own safety overview: first model to reach Critical cybersecurity capability under the Preparedness Framework.
Playco case study: 50% fewer manual fixes prototyping games with Astra.
Legora case study: reviewed 41 financial documents in minutes, catching all four planted errors.
Latent Space frames Astra's economics as hiring an AI engineer for under $6/hour.
AINews recap: new SOTA computer-use and coding scores, 2.5x pricier per token but cheaper per completed task, and reduced monitorability.
OpenAI's own safety overview: first model to reach Critical cybersecurity capability under the Preparedness Framework.
Playco case study: 50% fewer manual fixes prototyping games with Astra.
Legora case study: reviewed 41 financial documents in minutes, catching all four planted errors.
Latent Space frames Astra's economics as hiring an AI engineer for under $6/hour.
AINews recap: new SOTA computer-use and coding scores, 2.5x pricier per token but cheaper per completed task, and reduced monitorability.
What to watch — open questions
Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.