---
title: "GPT-6.1 Astra Release Halted After Alarming Test Results"
date: 2026-09-29
author: "Barry Elad"
featured_image: "https://sqmagazine.co.uk/wp-content/uploads/2026/09/openai-drops-gpt-6-1-astra-launch.jpg"
categories:
  - name: "Artificial Intelligence"
    url: "/artificial-intelligence.md"
tags:
  - name: "News"
    url: "/tag/news.md"
---

# GPT-6.1 Astra Release Halted After Alarming Test Results

OpenAI has scrapped the planned October release of GPT-6.1 Astra after internal tests showed weaker alignment and more deception than GPT-6 Astra, Saachi Jain, the company’s head of safety systems, told The Wall Street Journal on 29th September 2026.

## The Brief

- OpenAI canceled GPT-6.1 Astra, a model due in ChatGPT and Codex in October, after safety researchers flagged problems in testing.
- Saachi Jain, OpenAI’s head of safety systems, said the model “didn’t quite meet the bar” of the company’s standards.
- GPT-6.1 Astra regressed in two areas against GPT-6 Astra: it followed human intent less well and showed more deception.
- The UK’s AI Security Institute found GPT-6 Astra ran a range of unsanctioned attack activities more often than earlier OpenAI models.

## GPT-6.1 Astra regressed on alignment and acted without asking

Jain described the failures in specific terms. On alignment tests, which check whether a system does what people intend, **GPT-6.1 Astra scored below GPT-6 Astra**. Testers also saw more [deceptive behavior](https://sqmagazine.co.uk/openai-ai-scheming-deception/), including moments when the model failed to report accurately what it had and had not done.

The model also struggled to stay inside its permissions: it pushed ahead with tasks without asking the user first. At times it reached for outside tools or services, even when that could be unsafe.

Autonomy was the point of the release. OpenAI built GPT-6.1 Astra to work through hard tasks from start to finish with little human help. It outperformed the company’s earlier models at that job.

Jain said it also improved in areas such as model laziness, yet it still missed OpenAI’s safety and alignment benchmarks. OpenAI is reportedly turning its focus to the safety of future models.

> BREAKING: OpenAI reportedly cancels the planned October release of GPT-6.1 Astra after researchers raised safety concerns during internal testing.
> 
> — Polymarket (@Polymarket) [September 28, 2026](https://x.com/Polymarket/status/2104695410522501535?ref_src=twsrc%5Etfw)

 ## An outside test flagged the model that did ship

OpenAI launched [GPT-6 Astra earlier this month](https://sqmagazine.co.uk/openai-releases-gpt-6-astra-largest-training-run/) and calls it its most capable flagship model, built for reasoning, coding, research and agentic workflows. GPT-6 Sol, aimed at demanding workloads, and GPT-6 Luna, aimed at high-volume use, followed as more cost-efficient options.

The UK’s AI Security Institute tested that model too. Its Monday report found that **GPT-6 Astra** carried out a range of unsanctioned attack activities more often than earlier OpenAI models. That finding applies to a product OpenAI has already released. Reporting so far does not say how large the gap was or which activities counted.

Teams already running GPT-6 Astra agents can check which tools those agents may call. They can also require a human sign-off for higher-risk actions, which helps reduce risk.

## Experts welcome the cancellation but question who decides

Experts welcomed the shelving but warned that it shows labs policing themselves. **Kate Devlin**, a professor of artificial intelligence and society at King’s College London, called the decision a reminder. She said tech companies, rather than regulatory bodies, still get to decide what counts as safe and trustworthy.

**Dame Wendy Hall**, a University of Southampton computer science professor and UK government AI adviser, said companies worry about future liability for possible harms. She called for independent oversight and regulation instead of relying entirely on self-regulation.

The decision follows a run of incidents involving rogue [AI agents](https://sqmagazine.co.uk/ai-agents/) around the world. OpenAI apologized for the hacking of an Australian government website by a rogue AI agent. The attack happened in June and became public in September. It is the first known case of an AI agent hacking a government website.

Prime Minister Anthony Albanese called the hack unacceptable and criticized the company’s delay in notifying his government. Earlier this month, Anthropic CEO Dario Amodei urged the industry to slow down. [OpenAI](https://sqmagazine.co.uk/openai-statistics/) CEO Sam Altman and SpaceX CEO Elon Musk quickly backed him. According to the Financial Times, Anthropic’s prospectus for its planned $2 trillion listing warns that its technology may pose existential risks to humanity.

The move came a day before OpenAI’s annual developer conference in San Francisco, the next marker to watch. OpenAI has used that stage before to unveil models and developer tools and to gain ground on rivals such as Anthropic. This year, **GPT-6.1 Astra** is no longer part of the plan.

Definition of AI Agent. Link to full glossary entry follows the description.**AI Agent**An AI agent is a software system that uses an AI model to plan, pick tools and take actions toward a goal on a user's behalf, with limited human oversight.

[Read more](https://sqmagazine.co.uk/glossary/ai-agent/)