Can you tell me what kind of and how many data are collected for building generative AI models, including yours?
Building modern generative AI models—especially natively multimodal ones like Gemini—requires feeding them an extraordinarily diverse and massive "diet" of information. Instead of just reading text, modern models are built to understand the world through words, code, sights, and sounds simultaneously. Here is a breakdown of what kind of data is collected and the staggering scale behind it. What Kind of Data is Collected?
Because top-tier models are multimodal, their training data covers nearly every medium of human communication and expression:
- Public Web Data: Articles, blog posts, digitized books, academic papers, and forums. This provides the foundational grammar, general knowledge, and cultural context of human language across over 100 languages.
- Source Code: Massive repositories of public and open-source code, technical documentation, and scripts written in dozens of programming languages (Python, Java, C++, HTML, etc.). This is what teaches the model logic, structured reasoning, and programming.
- Visual Data: Billions of images, photos, artwork, diagrams, and charts. Crucially, these are paired with descriptions or alt-text so the AI learns how to connect visual concepts to language.
- Audio and Video: Audio files (like podcasts and speech recordings) and videos (including video frames and transcripts, such as filtered YouTube content). This helps the model map out temporal data, tone of voice, and real-world physics.
- Synthetic & Human-Curated Alignment Data: Beyond raw internet scraping, models undergo specialized training using Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF). This includes highly curated Q&As, human preferences, and high-quality AI-generated (synthetic) data designed to teach complex math, logical reasoning, and safe behavior.
AI researchers measure text-based data volume in tokens (chunks of characters or sub-words). For visual or audio data, those modalities are also converted into an equivalent token count during processing. While leading AI labs (including Google) keep the exact proprietary recipe and total petabyte counts closely guarded for competitive and security reasons, industry standards provide a clear picture of the scale:
- The Trillion-Token Club: Frontier generative AI models are trained on datasets ranging from 15 to over 30 trillion tokens.
- The Scale Comparison: To put 20 trillion tokens into perspective, it is equivalent to roughly 25 to 30 million books. If a human read one book a week, it would take them over 500,000 years to read what a frontier model processes during pre-training.
- Multimedia Scale: This text corpus is backed by billions of images and millions of hours of video and audio data.
In AI development, a massive "data wall" has been reached. Simply scraping more of the internet introduces noise and low-quality data. Because of this, the focus has shifted heavily toward heavy deduplication, aggressive filtering (to remove explicit content, personal identifying information, and low-quality text), and generating pristine, high-logic synthetic data to make the models smarter rather than just bigger.
Are you looking at this from a developer's perspective to figure out how much data you need to fine-tune a model for a specific project, or are you just curious about the broader ethics and mechanics of how AI learns?
- https://discuss.google.dev/t/simplifying-data-preparation-for-gen-ai-with-googles-gemini-models-and-dataloop/174980
- https://blog.google/innovation-and-ai/technology/ai/google-gemini-ai/#:~:text=It%20was%20built%20from%20the,%2C%20audio%2C%20image%20and%20video.
- https://theanilbajar.medium.com/all-about-gemini-models-and-training-process-989fc3e25602
- https://www.reddit.com/r/Bard/comments/18mv4bk/is_this_public_knowledge_size_of_dataset_palm2_36/?solution=a82f33fe28c0db4ba82f33fe28c0db4b&js_challenge=1&token=7afd7253fec22262ff1c52b1703fe9ec52d9ddeb100b28e78e1318aa368e100b&jsc_orig_r=
- https://milvus.io/ai-quick-reference/what-training-data-was-used-for-google-embedding-2