DeepMind has agents testing Gemini. That is the part to notice.
Google DeepMind executives told Business Standard the lab is running AI agents alongside its scientists to find gaps in Gemini. Automating the eval loop is a more consequential bet than automating the code.

Google DeepMind has started running AI agents alongside its own scientists — not to build product, but to find the gaps in Gemini. Two of the lab's senior executives laid out the setup to Business Standard: Manish Gupta, senior director for India and Asia-Pacific, and Seshu Ajjarapu, senior director of applied AI.
What the agents are doing is the part worth slowing down for. They've been handed routine testing. The models are being evaluated, in part, by models.
What was said
DeepMind's India teams have changed how they work. AI isn't only the thing being built any more, it's part of the apparatus doing the building, and the stated ambition runs past shipping software to charting a route toward artificial general intelligence. Demis Hassabis, DeepMind's chief executive and a joint winner of the 2024 Nobel Prize in Chemistry, has put AGI as early as 2029.
The near-term claim is narrower and easier to believe. Handing routine testing to agents should let DeepMind move faster on specialised applications, and the two named were drug discovery and materials science. There's a commercial edge as well: India's developer market, where DeepMind is up against OpenAI and Anthropic for the same builders.
I should say that the Business Standard piece is a subscriber article and I've worked from the free portion. The facts above are what's on the public page. The rest of this is mine.
Automating the eval loop is the actual move
Most agent deployments I see are pointed at the wrong half of the work. Everyone wants agents writing features. Far fewer want them doing the unglamorous job of finding where the system already falls over.
The second job has the cleaner feedback signal. A generated feature needs a human to judge whether it was any good. A found gap either reproduces or it doesn't — you can verify it cheaply, which means you can let it run without a person supervising every cycle.
It's also where teams actually lose time. Testing is what gets cut under deadline, because the cost of skipping it arrives later and usually lands on somebody else. Aiming capacity at that is a more interesting bet than another code-generation demo.
What I'd want to ask them
Models evaluating models has a known failure mode. The evaluator inherits blind spots from the thing it's evaluating. If the testing agents are built on the same family under test, the gaps they're worst at spotting are the ones that matter most — the failures the architecture is systematically bad at seeing in the first place.
Not a reason to avoid the approach. A reason to keep an independent check on whatever the automated pass concludes. Whether DeepMind does that, the free portion doesn't say, and I'd want an answer before copying the pattern.
There's a version of this that works: agents doing wide, cheap, boring coverage, humans reading the aggregate for what's conspicuously missing. And a version that quietly rots, where the eval suite drifts toward whatever the models find easy to check. The difference sits in the oversight layer, not in the agents.
Where it leaves things
Treat AGI-by-2029 as a leader stating a position rather than a date to plan against. Set it aside and the concrete claim underneath still stands: a frontier lab is using agents inside its own research loop and expects that to compound into faster progress on drug discovery and materials science.
If it does compound, the consequence isn't a better Gemini. It's that a lab's rate of improvement starts depending on how well it automates its own evaluation — a very different advantage from having the most compute or the best researchers, and much harder to observe from outside.
The India angle is more immediate. Competing for developers in that market tends to sharpen the tooling and the pricing for everyone building there, and that shows up long before 2029.
Source: Business Standard