๐Ÿ“ฐStalecollected in 13m

Google to use search interactions for AI training

Google to use search interactions for AI training
PostLinkedIn
๐Ÿ“ฐRead original on The Verge

๐Ÿ’กUnderstand how Google is leveraging user-generated multimodal data to fuel its AI training pipeline.

โšก 30-Second TL;DR

What Changed

Google will store Lens photos, voice searches, and Translate audio under a new 'Search Services History' category.

Why It Matters

This policy shift highlights Google's aggressive push to secure multimodal training data from its massive user base. It raises significant privacy concerns regarding how personal user interactions are repurposed for proprietary model development.

What To Do Next

Review your Google account 'Search Services History' settings to ensure your personal interaction data is not being used for AI training if you prefer to opt out.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขGoogle will store Lens photos, voice searches, and Translate audio under a new 'Search Services History' category.
  • โ€ขCollected media data is explicitly designated for training Google's AI models.
  • โ€ขUsers maintain control via the 'Save Media' toggle and the ability to disable the history setting entirely.

๐Ÿง  Deep Insight

Web-grounded analysis with 28 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขThe new 'Search Services History' setting expands Google's AI training data sources to include user interactions beyond just publicly available information, encompassing search queries, generative AI responses, and browsing activity within Search services.
  • โ€ขGoogle has historically adjusted its privacy policies to incorporate user data for AI training, with a notable update in July 2023 explicitly stating the use of publicly available information for models like Google Translate, Bard, and Cloud AI capabilities.
  • โ€ขWhile the 'Search Services History' allows users to opt out of future data collection for AI training, Google also employs automated filters to remove identifying or sensitive personal information from collected data before it is used.
  • โ€ขGoogle's approach to AI training data collection has faced scrutiny, particularly regarding the default opt-in nature of some features and the perceived complexity of privacy settings, as seen with past controversies around Gmail and Workspace data.
  • โ€ขThe data collected through Search Services History is intended to improve Google's AI models, including those powering personalized recommendations and features like Google Lens and Translate, which rely on vast datasets for their functionality.
๐Ÿ“Š Competitor Analysisโ–ธ Show

Competitor AI Training Data Practices

Feature/CompanyGoogle (Search Services History)Meta (AI Chatbots, Social Platforms)Microsoft (Copilot, Bing, Office)Apple (Siri, Apple Intelligence)
Data Sources for AI TrainingLens photos, voice searches, Translate audio, search queries, generative AI responses, browsing activity within Search services.Public posts, photos, captions, and chatbot interactions across Facebook, Instagram, and WhatsApp. Claims not to use private messages.De-identified search and news data, interactions with ads, voice and conversation activity with Copilot, including uploaded images/files.Emphasizes on-device processing; for cloud tasks, uses 'Private Cloud Compute' (PCC) with specialized personal data for requests, not general training.
User Control/Opt-outOpt-out via 'Save Media' toggle and disabling 'Search Services History' setting. When off, future activity not used for AI training (unless feedback provided).Complex opt-out process, often requiring detailed reasons; may not fully remove data already ingested. No universal opt-out for Meta AI.Provides opt-out controls in Copilot, Bing, and Microsoft Start. Users can disable 'Connected Experiences' in Office.Allows users to auto-delete conversations and object to URL crawling for AI training. Opt-out available for Apple Intelligence.
Privacy Claims/SafeguardsUses filters to automatically remove identifying/sensitive personal information.Claims not to use private messages; data is depersonalized.Claims not to use personal account data, identifying info in uploaded images/files, or sensitive personal data. Data is de-identified.Stresses privacy-first, on-device processing, and stateless computation in PCC; contractual bar on Google training on Apple user data for Siri.
Third-Party PartnershipsN/A (for this specific feature)N/AN/APartnered with Google for some Siri AI functionality (custom Gemini model on Google Cloud), with contractual restrictions on data use.

๐Ÿ› ๏ธ Technical Deep Dive

  • Google Lens utilizes convolutional neural networks (CNNs) for image recognition and natural language processing (NLP) for text extraction, trained on extensive datasets to identify objects, text, and scenes.
  • Google Translate employs Neural Machine Translation (NMT) models, specifically Google Neural Machine Translation (GNMT), which learn from billions of real sentences through tokenization, vectorization, and Transformer models to produce natural-sounding translations.
  • Google's AI training infrastructure incorporates robust annotation systems for metadata, policy engines to evaluate data usage, and de-identification/anonymization systems to ensure data privacy and compliance.
  • For handling unstructured data like images in custom AI model training on Vertex AI, Google Cloud Storage FUSE allows training jobs to access data directly from Cloud Storage buckets as if they were local files.
  • Google's Gemini models are designed with multi-modality, enabling them to process and cross-reference various data types such as text, images, and videos to achieve a richer, more contextual understanding.

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Increased pressure for explicit opt-in consent models for AI training data.
Ongoing public and regulatory scrutiny of tech companies' default opt-out data collection practices may lead to demands for more transparent and proactive user consent mechanisms.
Further integration of personal search history will lead to highly personalized, but potentially echo-chambered, AI responses.
As AI models leverage individual search and media history, responses will become more tailored, which could enhance relevance but also limit exposure to diverse perspectives.
Heightened competition among tech giants to differentiate on AI privacy features.
With varying approaches to user data for AI training, companies will increasingly use privacy as a key differentiator to attract and retain users concerned about their digital footprint.

โณ Timeline

2006-04
Google Translate launched with statistical machine translation.
2015-02
Google's voice search recording history became widely known, with a user portal introduced to access recordings.
2016-11
Google introduced the Neural Machine Translation (GNMT) system for Google Translate, improving translation quality.
2023-07
Google updated its privacy policy to explicitly state the use of publicly available information for training AI models like Translate, Bard, and Cloud AI capabilities.
2025-01
Google announced easier ways for users to modify privacy settings to block AI access to private content within Workspace Smart Features.
2026-06
Google introduces 'Search Services History' setting to collect images, audio, and video from Lens, Search Live, and Translate for AI model training.
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: The Verge โ†—