Skip to content

Instantly share code, notes, and snippets.

@ibrezm1
Created May 19, 2026 02:31
Show Gist options
  • Select an option

  • Save ibrezm1/09d8b93811d40a8673d3bdfa9fa1efb0 to your computer and use it in GitHub Desktop.

Select an option

Save ibrezm1/09d8b93811d40a8673d3bdfa9fa1efb0 to your computer and use it in GitHub Desktop.
Rag ADK

Title: Deployment modes in RAG Engine on Gemini Enterprise Agent Platform  |  Google Cloud Documentation

Description: Manage Deployment modes in RAG Engine on Gemini Enterprise Agent Platform. Configure Serverless and Spanner modes to optimize RAG performance and costs.

Source: https://docs.cloud.google.com/gemini-enterprise-agent-platform/build/rag-engine/deployment-modes

Using the UI for RAG Grounding

You can also configure RAG Engine grounding directly through the Google Cloud UI.

  1. Model Settings: Select your model (e.g., gemini-3.1-flash-lite) and set your Thinking level and Structured output preferences.
  2. Grounding: In the grounding section, you can select tools like Google Search, Google Maps, or RAG Engine (along with Agent Search or Elasticsearch).
  3. Region: Make sure your region is set appropriately (e.g., global).
  4. Change the Corpus: When using the RAG Engine for grounding in the UI, you can easily change the corpus you are querying against directly from the interface.

In Google Cloud, these solve related but different problems:

Option Best For Managed AI Workflow Semantic Search LLM Integration Typical Product
RAG AI chat/Q&A over documents Yes Yes Built-in Vertex AI RAG Engine
Vector Search Custom similarity systems Partial Yes You build it Vertex AI Vector Search
Search Traditional enterprise/document search Yes Limited/Hybrid Optional Vertex AI Search

1. Vertex AI RAG Engine

Use when you want:

  • Chat with documents
  • Gemini grounded on your data
  • Minimal infrastructure
  • Managed ingestion + retrieval

Google handles:

  • chunking
  • embeddings
  • indexing
  • retrieval
  • grounding

You mainly do:

rag.retrieval_query(...)

or attach RAG to Gemini tools.

Best for

  • Internal AI assistants
  • PDF Q&A
  • Knowledge base chatbots
  • Fastest time-to-market

Pros

  • Simplest
  • Native Gemini integration
  • Least code
  • Fully managed

Cons

  • Less control
  • Less customizable ranking
  • Limited advanced retrieval tuning

2. Vertex AI Vector Search

Use when you need:

  • Custom retrieval systems
  • High-scale vector similarity
  • Hybrid ranking
  • Recommendation engines
  • Full control

You manage:

  • embeddings
  • chunking
  • metadata
  • indexing pipeline
  • retrieval orchestration

Typical flow:

Documents
 → embeddings
 → Vector Search index
 → retrieve neighbors
 → send context to Gemini

Best for

  • Large-scale semantic search
  • Recommendation systems
  • Custom AI architecture
  • Multi-modal retrieval
  • Production ML infra

Pros

  • Maximum flexibility
  • High performance
  • Enterprise scale
  • Custom ranking/filtering

Cons

  • More engineering work
  • You build RAG orchestration yourself

3. Vertex AI Search

Use when you need:

  • Google-like enterprise search
  • Website search
  • Search portal
  • Structured + unstructured search

This is closer to:

  • enterprise search
  • website search
  • document discovery

Less focused on LLM orchestration.

Best for

  • Search experiences
  • Employee portals
  • Ecommerce/document search
  • Websites

Pros

  • Strong search UX
  • Search-focused features
  • Easy indexing

Cons

  • Less flexible for advanced GenAI
  • Not optimized primarily for Gemini grounding

What most teams should use

Use Vertex AI RAG Engine if:

  • You are building AI chat/Q&A
  • You want Gemini + documents
  • You want fastest implementation
  • Your corpus is moderate size

This is probably what you want right now.


Use Vector Search if:

  • You need custom architecture
  • Millions/billions of vectors
  • Advanced retrieval tuning
  • Hybrid ranking pipelines
  • Recommendation/search platform

Use Vertex AI Search if:

  • You want Google-style enterprise search
  • Search UI is primary product
  • Not mainly conversational AI

Recommended Architecture

Small/Medium AI Assistant

Gemini
 + Vertex AI RAG Engine

Advanced Enterprise AI

Gemini
 + Vertex AI Vector Search
 + Custom retrieval pipeline

Enterprise Search Portal

Vertex AI Search
 + optional Gemini summaries

For your current code

Since you're already using:

from vertexai import rag

You are using:

✅ Vertex AI RAG Engine

That is the correct choice if your goal is:

  • chatbot over docs
  • grounding Gemini
  • internal knowledge assistant
  • fast implementation

You only need Vector Search directly if you outgrow the managed RAG abstraction.

from google import genai
from google.genai import types
import base64
import os
def generate():
client = genai.Client(
vertexai=True,
project="user-xpkurowsfvgo",
location="us-south1",
)
model = "gemini-2.5-flash"
contents = [
"What is RAG and why is it helpful?"
]
tools = [
types.Tool(
retrieval=types.Retrieval(
vertex_rag_store=types.VertexRagStore(
rag_resources=[
types.VertexRagStoreRagResource(
rag_corpus="projects/user-xpkurowsfvgo/locations/us-south1/ragCorpora/2305843009213693952"
)
],
similarity_top_k=20,
)
)
)
]
generate_content_config = types.GenerateContentConfig(
temperature = 1,
top_p = 0.95,
max_output_tokens = 65535,
safety_settings = [types.SafetySetting(
category="HARM_CATEGORY_HATE_SPEECH",
threshold="OFF"
),types.SafetySetting(
category="HARM_CATEGORY_DANGEROUS_CONTENT",
threshold="OFF"
),types.SafetySetting(
category="HARM_CATEGORY_SEXUALLY_EXPLICIT",
threshold="OFF"
),types.SafetySetting(
category="HARM_CATEGORY_HARASSMENT",
threshold="OFF"
)],
tools = tools,
)
for chunk in client.models.generate_content_stream(
model = model,
contents = contents,
config = generate_content_config,
):
if not chunk.candidates or not chunk.candidates[0].content or not chunk.candidates[0].content.parts:
continue
print(chunk.text, end="")
generate()
from vertexai import rag
# gcloud services enable cloudresourcemanager.googleapis.com
import vertexai
vertexai.init(project='user-xpkurowsfvgo', location="us-south1")
for corpus in rag.list_corpora():
print(f"Name: {corpus.name} | Display Name: {corpus.display_name}")
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment