LLM benchmark race
AI labs are racing to overtake each other on key industry benchmarks. But this intense race has stripped the benchmarks of most of their value.
computer use agent
WALT abstracts away the chaos of dynamic layouts, allowing AI to focus on high-level planning instead of low-level clicks.
AI puzzle solving
The verified solution achieves 54% accuracy on the semi-private test set, outperforming Gemini 3 Deep Think at less than half the cost.
OpenAI code red
OpenAI’s problem is not that it doesn't have the best model anymore but that the general feeling is that it has fallen behind.
LLM reinforcement learning
Reinforcement learning from verifiable rewards (RLVR) ushered in a new generation of reasoning models. Now, researchers are looking beyond RLVR to create the next breakthrough in AI.
vulnerable IDE
An indirect prompt injection turns the AI agent in Google's Antigravity IDE into an insider threat, bypassing security controls to steal credentials.
World models
One of the most accomplished AI scientists is departing his long-time role at Meta. What do we know about Yann LeCun's vision for the future of AI?
Nano Banana Pro
By combining advanced reasoning with real-time data, Google's Nano Banana Pro redefines what's possible in image-generation AI.
Google Gemini 3.0
Google took quite a bit to release the next version of its Gemini models. And it didn't disappoint.
malicious ai actor
By breaking down complex attacks into seemingly innocent steps, the hackers bypassed Claude's safety guardrails and unleashed an autonomous agent.