Pichai Acknowledges Coding Gap Ahead of Gemini 4

For practitioners, coding and agentic-coding performance has become a decisive evaluation dimension for frontier models, alongside cost, latency, and general reasoning. Publicly reported model delays can therefore affect how engineering teams sequence model testing and vendor selection. Reuters reports that Alphabet CEO Sundar Pichai acknowledged coding and agentic coding as areas for improvement during the company's Q2 2026 earnings call. Gemini 3.5 Pro remains in testing after a planned June release was delayed, according to Reuters and Google's earnings-call transcript. Google stated that it has begun its "most ambitious pre-training run yet" for Gemini 4, while Reuters reported that Pichai described a roadmap with releases at almost a monthly cadence.
Frontier coding is the immediate issue
Gemini 3.5 Pro remains delayed
Reuters reported that Gemini 3.5 Pro, originally slated for a June release, remained in partner testing after its delay. Google's published earnings-call transcript likewise states that Gemini 3.5 Pro is "currently in testing." Reuters described the postponed model as one expected to strengthen Google's standing in AI coding and autonomous agent tasks.
Search Engine Journal reported that Pichai characterized Gemini 4 as necessary to compete at the next frontier, and said the next base model would be larger. Google's transcript confirms the underlying development milestone: Pichai stated, "We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress we are seeing at the frontier."
Reuters also reported that Pichai described the Gemini 4 roadmap as involving releases at almost a monthly cadence. The retrieved sources do not provide a release date, model weights, context-window size, benchmark results, pricing, or API specifications for Gemini 4.
Scale does not resolve workload-specific evaluation
Google reported that more than 9 million developers build monthly with its models across APIs and developer products, and that its model APIs process about 22 billion tokens per minute, up from 16 billion in the prior quarter. The company also reported that demand for its models remains supply constrained.
What engineering teams can monitor
For practitioners
Coding and agentic-coding workloads are now a high-stakes model-selection category because they combine code generation, tool use, iterative debugging, and long-running task reliability. Industry context: teams evaluating frontier APIs commonly need workload-specific tests rather than relying on general-purpose benchmarks, especially where an agent can modify repositories or invoke production tools.
Reuters reports that Alphabet CEO Sundar Pichai acknowledged a gap in coding and agentic coding during the Q2 2026 earnings call. "We've had clearly frontier models. There are many attributes on which we are still at the frontier; there are areas where we've acknowledged we need to improve and coding and agentic coding is an example of that," Pichai said, according to Reuters.
The distinction matters because coding quality is not a single capability. Reliable software agents require competent code synthesis, repository-level context handling, test execution, error recovery, structured tool calls, and appropriate permissions. A model can perform strongly on general reasoning while remaining less dependable on an end-to-end engineering workflow.
Usage scale and token throughput indicate extensive deployment activity, but they do not by themselves establish superiority on software-engineering or agentic tasks. Comparable transitions in the sector make independent evaluation important, using representative repositories, test suites, tool schemas, latency budgets, and failure-handling requirements.
Google also highlighted the Gemini Flash model family for its performance-cost tradeoff. Reuters reported that Pichai emphasized Flash models for cybersecurity, customer service, analytics, and enterprise software. This places the frontier-model discussion alongside a separate operational question for ML platform teams: whether a lower-cost model meets a workflow's quality threshold without escalation to a larger model.
Public model-roadmap statements are most useful when translated into concrete evaluation checkpoints rather than assumed capability gains. Relevant evidence from a future Gemini release would include:
- •Reproducible coding and agentic-task results on disclosed benchmarks and realistic repositories.
- •API details covering tool use, context management, rate limits, pricing, and regional availability.
- •Reliability measurements for multi-step tasks, including test pass rates, error recovery, and unsafe-action controls.
- •Clear separation between a model's raw capability and the performance of the surrounding agent runtime, retrieval system, and tools.
Reuters framed the earnings-call exchange as a response to investor concerns over the Gemini 3.5 Pro delay and competition from other AI providers. For technical buyers, the immediate reported development is not a new model release, but a public acknowledgment that coding and agentic coding remain areas Google intends to improve while Gemini 4 is in pre-training.
Key Points
- 1Pichai acknowledged coding and agentic coding gaps, making task-specific evaluation central for teams considering Gemini in software-engineering workflows.
- 2Gemini 3.5 Pro remains in testing after a reported June delay, while Google confirms Gemini 4 has entered pre-training.
- 3Industry context: token scale and developer adoption do not substitute for reproducible agent reliability, tool-use safety, and repository-level coding tests.
Scoring Rationale
Google's public acknowledgment of coding and agentic-coding areas needing improvement is notable for teams benchmarking frontier model providers. Gemini 4 is not yet released, so the story offers roadmap and competitive context rather than immediately deployable technical capabilities.
Sources
Primary source and supporting public references used for this report.
Practice with real Ad Tech data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Ad Tech problems