AI in the workspace
Google Docs kills the split-screen: Gemini draws diagrams from your text
Google Docs now lets you generate and edit images, infographics, and diagrams directly using Gemini in the sidebar or bottom bar, leveraging document context for relevance. The rollout started July 28, 2026, for paid Workspace tiers.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-29 · 3 min read

You have probably built a report by splitting your screen between a doc and a separate image tool, hoping the art matches the text. That workflow is now optional. Google is rolling out the ability to generate and edit visuals directly inside Google Docs using Gemini, the company announced on July 28.
The feature goes beyond simple image generation. Gemini reads the document you are working on and can produce diagrams, infographics, or illustrations that match the content. You can ask it to Add a diagram providing an overview of the proposal at the top of my doc and get something that actually fits. You can also refine existing visuals with natural language commands like Change the aspect ratio to 16:9 or Make the style more aesthetic.
You can create or edit multiple visuals at once, telling Gemini to add infographics in each key section or to update the style of several graphics together. The capability is accessible from the bottom bar or the Gemini side panel in web-based Docs only for now.
Google is making the feature available gradually over up to 15 days starting July 28. It requires Gemini for Workspace to be enabled. On the admin side the feature is on by default if Gemini in Drive is active. End users need Workspace smart features turned on.
The availability is tiered. Business Standard and Plus, Enterprise Standard and Plus, Education Plus, and consumer Google AI Pro and Ultra subscribers all get access. Education add-ons including Google AI Pro for Education also qualify.
This addition continues a push Google has been making across its productivity suite. Earlier this year the company added Gemini capabilities to Gmail, Calendar, Chat, and Sheets. More recently it launched Vids, a tool that uses Gemini Omni to generate videos from text and lets you insert a personal avatar, effectively turning a selfie into a production studio. Every clip gets an invisible SynthID watermark. That same watermarking logic could apply to Docs generated images although Google did not confirm that. These moves echo a trend where AI tools embed generative capabilities directly into the interface, similar to how ElevenLabs turned video production into a conversation. Microsoft has also been integrating AI into its productivity suite, with its in-house MAI model matching GPT-5.6 in Excel at a lower cost.
Google is also deepening the underlying model layer. Recent releases include Gemini 3.5 Flash Cyber, a variant restricted to government defenders, and Gemini 3.6 Flash which appeared in the Gemini app. The company has said it started its most ambitious pre-training run for Gemini 4. This push toward specialized variants aligns with a broader shift: orchestration and specialization now matter as much as raw model size. Meanwhile, Microsoft's MAI models are cutting production costs by 89%, showing the competitive pressure on providers to reduce inference expenses. The Docs image generation relies on Gemini but the specific model version powering it was not disclosed.
This changes the workflow. For years, document creation meant assembling text and visuals from separate tools. Now the AI that writes the text can also generate the graphics. The context awareness keeps the output relevant. The natural language editing removes the need for design skills to polish a chart. A small change on paper; a large one in practice.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.