Kurs
GPT-6 Astra's safety pitch is about alignment rather than refusals alone. OpenAI calls it "our most aligned model," pointing to a 0.00% score on an internal circumvention benchmark (versus 0.29% for GPT-5.6 Sol) and a 4.2% rate on an internal hallucination benchmark about its own capabilities (versus 12.2% for Sol). It also says Astra never tried to bypass an auto-review safeguard in testing, even when that safeguard was made easy to evade.
- The Takeaway: Grok 4.7 focuses on stopping bad requests upfront. GPT-6 Astra focuses on trusting its judgment once it's already acting autonomously.



