Title: Deployment modes in RAG Engine on Gemini Enterprise Agent Platform | Google Cloud Documentation
Description: Manage Deployment modes in RAG Engine on Gemini Enterprise Agent Platform. Configure Serverless and Spanner modes to optimize RAG performance and costs.
Source: https://docs.cloud.google.com/gemini-enterprise-agent-platform/build/rag-engine/deployment-modes
You can also configure RAG Engine grounding directly through the Google Cloud UI.
- Model Settings: Select your model (e.g.,
gemini-3.1-flash-lite) and set your Thinking level and Structured output preferences. - Grounding: In the grounding section, you can select tools like Google Search, Google Maps, or RAG Engine (along with Agent Search or Elasticsearch).
- Region: Make sure your region is set appropriately (e.g.,
global). - Change the Corpus: When using the RAG Engine for grounding in the UI, you can easily change the corpus you are querying against directly from the interface.
In Google Cloud, these solve related but different problems:
| Option | Best For | Managed AI Workflow | Semantic Search | LLM Integration | Typical Product |
|---|---|---|---|---|---|
| RAG | AI chat/Q&A over documents | Yes | Yes | Built-in | Vertex AI RAG Engine |
| Vector Search | Custom similarity systems | Partial | Yes | You build it | Vertex AI Vector Search |
| Search | Traditional enterprise/document search | Yes | Limited/Hybrid | Optional | Vertex AI Search |
Use when you want:
- Chat with documents
- Gemini grounded on your data
- Minimal infrastructure
- Managed ingestion + retrieval
Google handles:
- chunking
- embeddings
- indexing
- retrieval
- grounding
You mainly do:
rag.retrieval_query(...)or attach RAG to Gemini tools.
- Internal AI assistants
- PDF Q&A
- Knowledge base chatbots
- Fastest time-to-market
- Simplest
- Native Gemini integration
- Least code
- Fully managed
- Less control
- Less customizable ranking
- Limited advanced retrieval tuning
Use when you need:
- Custom retrieval systems
- High-scale vector similarity
- Hybrid ranking
- Recommendation engines
- Full control
You manage:
- embeddings
- chunking
- metadata
- indexing pipeline
- retrieval orchestration
Typical flow:
Documents
→ embeddings
→ Vector Search index
→ retrieve neighbors
→ send context to Gemini
- Large-scale semantic search
- Recommendation systems
- Custom AI architecture
- Multi-modal retrieval
- Production ML infra
- Maximum flexibility
- High performance
- Enterprise scale
- Custom ranking/filtering
- More engineering work
- You build RAG orchestration yourself
Use when you need:
- Google-like enterprise search
- Website search
- Search portal
- Structured + unstructured search
This is closer to:
- enterprise search
- website search
- document discovery
Less focused on LLM orchestration.
- Search experiences
- Employee portals
- Ecommerce/document search
- Websites
- Strong search UX
- Search-focused features
- Easy indexing
- Less flexible for advanced GenAI
- Not optimized primarily for Gemini grounding
- You are building AI chat/Q&A
- You want Gemini + documents
- You want fastest implementation
- Your corpus is moderate size
This is probably what you want right now.
- You need custom architecture
- Millions/billions of vectors
- Advanced retrieval tuning
- Hybrid ranking pipelines
- Recommendation/search platform
- You want Google-style enterprise search
- Search UI is primary product
- Not mainly conversational AI
Gemini
+ Vertex AI RAG Engine
Gemini
+ Vertex AI Vector Search
+ Custom retrieval pipeline
Vertex AI Search
+ optional Gemini summaries
Since you're already using:
from vertexai import ragYou are using:
That is the correct choice if your goal is:
- chatbot over docs
- grounding Gemini
- internal knowledge assistant
- fast implementation
You only need Vector Search directly if you outgrow the managed RAG abstraction.