Google yesterday unveiled its new flagship Gemini 4 model with some impressive scores, but a report suggests that it might “struggle” in some cases.
Bloomberg reports that, internally, Google employees have found that Gemini 4’s real-world performance doesn’t always line up with its benchmark scores. During the unveiling of Gemini 4 Argon, Google showed that its new model outpaced competitors including OpenAI’s GPT-6 Astra in many key areas.
The report cites Googlers who found that the model “does less well when employees actually put it to work,” and that it “struggles to handle certain coding tasks.” Google, speaking to Bloomberg on the claims, said that it “would be inaccurate to say that Gemini 4 is underperforming in areas such as coding,” also referencing comments from Koray Kavukcuoglu, the head of Google DeepMind, from late September, when he said that “it’s a certainty that we are always gonna be at the frontier.” The report cites one Googler who said that there’s a “large consensus” that Gemini 4 is at the frontier amongst employees, while others are concerned that Gemini “will still lag behind” competing models from Anthropic and OpenAI. Clearly, there’s a split in how employees view the model’s performance.
Gemini 4 Argon brings welcome upgrades to the model, including an expanded output limit of up to 1 million tokens, while Google also explained that the model is “highly capable at cybersecurity defense” and shows “leading performance” in other use cases. Gemini 4 Argon comes at a cost of “$4 per 1M input tokens and $20 per 1M output tokens” but will see introductory pricing which is half of that.
More on Gemini:
- Google details Gems to skills migration, including free access
- Gemini app starts rolling out big side panel simplification, settings redesign
- Pixel 11 starts testing ‘Call for Me’ to have Gemini phone on your behalf
Follow Ben: Twitter/X, Threads, Bluesky, and Instagram
FTC: We use income earning auto affiliate links. More.
Comments