Gemini 4 Is Here, and Google’s Flagship Tops All Other AI Models on Cybersecurity
Google launched Gemini 4 Argon as its frontier model for coding, office work, and cyber defense. It scored 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 and GPT-6 Astra, and can generate up to 1 million tokens in a single response. Its benchmark claims should be treated cautiously because Google used its own DeepSWE evaluation, though Argon led on 12 of 18 benchmarks overall. The standout feature is cybersecurity. On an indirect prompt-injection test, Argon scored 0.7%, far better than GPT-6 Astra and weaker competitors. Google is releasing it first to vetted security teams through its Fairwind program and says it ships without cyber guardrails so defenders can stress-test systems. Google also reports stronger penetration-testing performance than its earlier cyber model and says Argon helped find a critical healthcare software flaw. Wider release starts with paid API and Google AI Ultra users, with introductory pricing of $2 per million input tokens and $10 per million output tokens.
