Box is integrating Google Cloud's Gemini Embeddings 2 into its Agentic Platform, letting enterprise AI agents search and reason over images, charts, and tables stored in Box, not just text. The move addresses a long-standing gap in retrieval-augmented generation (RAG) systems: visual content in enterprise documents has been invisible to AI search until now.

The partnership, announced August 19, 2026, targets the trillions of gigabytes of enterprise content Box already hosts, from financial models and clinical trial protocols to engineering schematics and M&A due diligence materials. Text-based RAG has handled prose well. But a chart in a slide deck, a table in a financial report, or a diagram in an engineering spec? Those remained unsearchable by automated systems.
What Gemini Embeddings 2 actually does
Google's NEW Multimodal Model - Gemini Embedding 2
Gemini Embeddings 2 creates a unified vector space where text, images, document pages, spreadsheet tables, and charts all live together. That sounds abstract, so here's what it means in practice: an analyst can type "show me the revenue growth chart from Q2 reports" and the system retrieves the actual chart, not a text description of it.

Three capabilities stand out. First, crossmodal retrieval lets natural language queries pull up specific visual components. No manual tagging required. Second, layout-aware embedding preserves the structure of document pages, keeping callout boxes, columns, and visual hierarchies intact rather than flattening everything to text. Third, the system bridges heterogeneous formats natively, so .docx, .xlsx, .pdf, .pptx, .png, and .csv files all contribute to the same searchable knowledge base.
Why spatial structure matters for enterprise AI
Consider a financial table. Converting it to plain text strips column headers from their data points. A multimodal embedding preserves that spatial relationship, so an AI agent understands that "$4.2M" belongs to "Q3 Revenue," not just that both appear somewhere in the document.

The same logic applies to flowcharts, org charts, and technical diagrams. These aren't decorative; they encode meaning that text summaries often lose. Box and Google Cloud argue this is the natural next step for enterprise RAG: extending proven text retrieval to the visual and structural elements that live alongside prose in real business documents.
Three design patterns Box identified
Box's announcement outlines three primary use cases. The first targets financial and analytical reporting. Audit teams and analysts deal with tables, growth charts, and footnotes embedded in documents. Multimodal agents can now cross-reference written summaries against the actual visual data, flagging discrepancies between what a report claims and what its charts show.

The second pattern involves clinical and scientific data, where images and schematics carry evidential weight. The third, though only partially detailed in the announcement, appears to focus on engineering and technical documentation, where diagrams and flowcharts often contain the most critical information.

What the announcement doesn't say
Box and Google Cloud haven't disclosed pricing for the multimodal capabilities, whether they require a specific Box tier, or what the compute costs look like for embedding large document libraries. The announcement also doesn't specify latency benchmarks for multimodal search compared to text-only RAG, which matters for real-time agent applications.
There's also the question of accuracy. Visual embeddings are newer than text embeddings, and edge cases abound. A hand-drawn whiteboard photo, a low-resolution scan, or a chart with unusual formatting could all trip up the system. Box positions this as extending a "powerful and highly effective baseline," but production deployments will reveal how robust that extension really is.
Logicity's Take
This is a significant infrastructure move, not a splashy product launch. Multimodal RAG has been the obvious next step since text-based RAG proved itself, and Box is smart to partner with Google Cloud rather than build embeddings in-house. For engineering leaders evaluating enterprise AI, the real question is whether your document corpus justifies the cost. If your teams spend hours manually searching for charts or cross-referencing visual data, this integration could cut that time dramatically. If your documents are mostly prose, text RAG still handles 90% of the job.
Need Help Implementing This?
Logicity helps engineering teams evaluate and deploy enterprise AI infrastructure. Reach out to discuss how multimodal RAG fits your document workflows.
Source: Cloud Blog
Manaal Khan
Tech & Innovation Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.





