Gemini Rolls Out Speak to Window on macOS

Google is rolling out Speak to Window in English to all Gemini for macOS users. Holding the fn key lets a user dictate into an active app or request edits with optional screen context; intelligent dictation removes filler words and handles corrections. Google says the app requires an Apple Silicon Mac running macOS Sequoia 15.0 or later.
Google is rolling out Speak to Window in English to all users of the Gemini app for macOS. The feature lets a user hold the Mac's fn key, speak a request and release the key to place Gemini's response into the active window.
Google's product page describes two modes. Intelligent dictation is available by default and cleans filler words, handles mid-sentence corrections and formats spoken text. Reasoning tasks that use screen context require an explicit opt-in under the app's Speak to Window settings.
Voice input can act on selected context
Users can dictate into text fields, highlight text and ask Gemini to change its tone or formatting, summarize selected files, or request an image. For longer dictation, Google says users can double-tap the fn key to begin hands-free recording and tap again to finish.
Screen sharing has a separate control. Pressing both Command keys sends the foreground window as chat context. Google says broader access to full browser pages or files in a shared local folder requires enabling Accessibility for Gemini in macOS Privacy & Security settings.
Android Authority reported the rollout on July 29 and independently confirmed the voice shortcut, filler-word cleanup and optional screen-aware reasoning. The official Google page now documents the same behavior and is the originating source for the product's current capabilities.
Availability and design implications
Gemini for macOS is available in supported Gemini markets, runs only on Apple Silicon Macs and requires macOS Sequoia 15.0 or later. Google cautions that compatibility and feature availability vary.
For teams building desktop assistants, Speak to Window shows how voice, active-app context and text transformation can be combined without making screen reasoning mandatory. The important product boundary is user control: dictation works by default, while screen-aware reasoning and broader Accessibility access require separate choices. Clear permission prompts and visible context selection matter because the assistant can act where the user is already writing rather than inside a separate chat window.
Key Points
- 1Google is rolling out Speak to Window in English to all Gemini for macOS users.
- 2Intelligent dictation is on by default, while reasoning over screen context requires an explicit opt-in.
- 3The app requires Apple Silicon and macOS Sequoia 15.0 or later, with availability varying by feature and market.
Scoring Rationale
The update combines voice input, active-window output and optional screen reasoning in a mainstream desktop assistant. It is useful to practitioners designing multimodal productivity tools and permission boundaries, but it is an incremental product rollout rather than a major model or platform release.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems
