← Terug naar overzicht

OpenAI unveiled GPT-6 Astra, described as the world's most intelligent and aligned model. The model reached the 'Critical' cybersecurity capability threshold under OpenAI's Preparedness Framework. GPT-6 Astra scored 100% on ExploitBench, a benchmark for evaluating AI exploit generation capabilities. OpenAI has implemented safeguards to block proof-of-concept (PoC) exploit requests from the model. The model demonstrates state-of-the-art performance in computer use, browsing, and software engineering. This development raises significant concerns about AI-assisted cyberattacks and the dual-use nature of advanced AI models. OpenAI's move to block PoC exploit requests reflects growing awareness of the offensive cybersecurity potential of frontier AI models.

Technical details

OpenAI unveiled GPT-6 Astra, described as its most advanced AI model, which achieved a perfect 100% score on ExploitBench — a benchmark evaluating a model's ability to convert known software vulnerabilities into working exploits. This compares to 78.5% for its predecessor GPT-5.6 Sol. Astra also scored 99.9% on ARC-AGI-3 and 98% on FrontierMath Tier 4. The model demonstrates significantly higher arbitrary code-execution rates than GPT-5.6 Sol when tested against vulnerabilities from July–August 2026, including two zero-day vulnerabilities in unspecified software. Without safeguards, Astra is capable of exploiting previously unknown vulnerabilities to achieve code execution in hardened browsers and develop privilege-escalation exploits for hardened operating systems. The released version is restricted to secure code review and patching, and blocks prompts requesting PoC exploit generation. OpenAI has incorporated stronger model robustness against jailbreaks, enhanced monitoring context, and additional misalignment detection safeguards. The model reached the 'Critical' cybersecurity capability threshold under OpenAI's Preparedness Framework. OpenAI plans to expand access with less restrictive safeguards via OpenAI Daybreak, enabling vulnerability and PoC validation, malware analysis, and detection engineering.

Mitigation steps

1. Organizations should be aware that GPT-6 Astra's public release is restricted to secure code review and patching workflows; PoC exploit generation requests are blocked by default. 2. Monitor OpenAI's Daybreak program for updates on expanded access and less restrictive safeguards rolling out in coming weeks. 3. Critical infrastructure organizations (water systems, electricity providers, state/local governments, banks, non-profits, open-source maintainers) should consider applying for OpenAI's Daybreak for Frontline Defenders program, which offers $1 billion in subsidized AI access, training, and technical assistance. 4. Public sector and water system defenders can participate in the pilot with the U.S. Multi-State Information Sharing and Analysis Center (MS-ISAC) for Daybreak access and guided training. 5. Security teams should anticipate that AI-powered exploit development capabilities will continue to advance rapidly, and should accelerate vulnerability patching and defensive workflows accordingly. 6. When using Astra in sensitive environments, review AI-generated actions before proceeding, as safety checks may occasionally interrupt legitimate defensive cybersecurity work.

Affected products

  • GPT-5.6 Sol (OpenAI
  • GPT-6 Astra (OpenAI)
  • Hardened browsers (unspecified
  • Hardened operating systems (unspecified
  • predecessor model)
  • targeted in evaluations)
  • targeted in evaluations)

Categories

  • Critical Infrastructure
  • Emerging Technologies
  • Security Tools
  • Zero-Day Vulnerabilities