All posts

Box adds Gemini Embeddings 2 for multimodal AI agents

Manaal KhanAugust 18, 2026 at 10:01 PM4 min read
Box adds Gemini Embeddings 2 for multimodal AI agents

Box is integrating Google Cloud's Gemini Embeddings 2 into its Agentic Platform, letting enterprise AI agents search and reason over images, charts, and tables stored in Box, not just text. The move addresses a long-standing gap in retrieval-augmented generation (RAG) systems: visual content in enterprise documents has been invisible to AI search until now.

Box and Google Cloud partnership for multimodal AI agents using Gemini Embeddings 2
https://storage.googleapis.com/gweb-cloudblog-publish/images/box-multimodal-agents-gemini-embeddings-he.max-2500x2500.png

The partnership, announced August 19, 2026, targets the trillions of gigabytes of enterprise content Box already hosts, from financial models and clinical trial protocols to engineering schematics and M&A due diligence materials. Text-based RAG has handled prose well. But a chart in a slide deck, a table in a financial report, or a diagram in an engineering spec? Those remained unsearchable by automated systems.

Advertisements

What Gemini Embeddings 2 actually does

Google's NEW Multimodal Model - Gemini Embedding 2

Gemini Embeddings 2 creates a unified vector space where text, images, document pages, spreadsheet tables, and charts all live together. That sounds abstract, so here's what it means in practice: an analyst can type "show me the revenue growth chart from Q2 reports" and the system retrieves the actual chart, not a text description of it.

Animated demonstration of multimodal search retrieving visual content from enterprise documents
https://storage.googleapis.com/gweb-cloudblog-publish/original_images/GIF_1_1AuFwiE.gif

Three capabilities stand out. First, crossmodal retrieval lets natural language queries pull up specific visual components. No manual tagging required. Second, layout-aware embedding preserves the structure of document pages, keeping callout boxes, columns, and visual hierarchies intact rather than flattening everything to text. Third, the system bridges heterogeneous formats natively, so .docx, .xlsx, .pdf, .pptx, .png, and .csv files all contribute to the same searchable knowledge base.

Why spatial structure matters for enterprise AI

Consider a financial table. Converting it to plain text strips column headers from their data points. A multimodal embedding preserves that spatial relationship, so an AI agent understands that "$4.2M" belongs to "Q3 Revenue," not just that both appear somewhere in the document.

Diagram showing how multimodal embeddings preserve table structure and spatial relationships
https://storage.googleapis.com/gweb-cloudblog-publish/images/2_9rTykxw.max-1700x1700.png

The same logic applies to flowcharts, org charts, and technical diagrams. These aren't decorative; they encode meaning that text summaries often lose. Box and Google Cloud argue this is the natural next step for enterprise RAG: extending proven text retrieval to the visual and structural elements that live alongside prose in real business documents.

Trillions of gigabytes
Volume of enterprise data Box stores, including financial models, clinical trials, and engineering schematics

Three design patterns Box identified

Box's announcement outlines three primary use cases. The first targets financial and analytical reporting. Audit teams and analysts deal with tables, growth charts, and footnotes embedded in documents. Multimodal agents can now cross-reference written summaries against the actual visual data, flagging discrepancies between what a report claims and what its charts show.

Example of multimodal AI agent analyzing financial reporting documents with embedded charts
https://storage.googleapis.com/gweb-cloudblog-publish/images/3_ZPNWwdP.max-1700x1700.png

The second pattern involves clinical and scientific data, where images and schematics carry evidential weight. The third, though only partially detailed in the announcement, appears to focus on engineering and technical documentation, where diagrams and flowcharts often contain the most critical information.

Multimodal enterprise agent workflow showing document analysis across multiple file formats
https://storage.googleapis.com/gweb-cloudblog-publish/images/4_IAwu96l.max-1800x1800.png

What the announcement doesn't say

Box and Google Cloud haven't disclosed pricing for the multimodal capabilities, whether they require a specific Box tier, or what the compute costs look like for embedding large document libraries. The announcement also doesn't specify latency benchmarks for multimodal search compared to text-only RAG, which matters for real-time agent applications.

There's also the question of accuracy. Visual embeddings are newer than text embeddings, and edge cases abound. A hand-drawn whiteboard photo, a low-resolution scan, or a chart with unusual formatting could all trip up the system. Box positions this as extending a "powerful and highly effective baseline," but production deployments will reveal how robust that extension really is.

ℹ️

Logicity's Take

This is a significant infrastructure move, not a splashy product launch. Multimodal RAG has been the obvious next step since text-based RAG proved itself, and Box is smart to partner with Google Cloud rather than build embeddings in-house. For engineering leaders evaluating enterprise AI, the real question is whether your document corpus justifies the cost. If your teams spend hours manually searching for charts or cross-referencing visual data, this integration could cut that time dramatically. If your documents are mostly prose, text RAG still handles 90% of the job.

ℹ️

Need Help Implementing This?

Logicity helps engineering teams evaluate and deploy enterprise AI infrastructure. Reach out to discuss how multimodal RAG fits your document workflows.

Source: Cloud Blog

M

Manaal Khan

Tech & Innovation Writer

Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.