跳到正文
原文
Tavily News Discovery·· 事件 2026-10-06AI 评分71

Google DeepMind 发布 EmbeddingGemma 2 端侧多模态嵌入模型

EmbeddingGemma 2 is a best-in-class open model for natively multimodal embeddings

官方原文核对 · 关键事实加强核对

依据原始来源与逐字段引文;不代表独立实测,厂商性能声明仍是厂商自述。

  • 原文依据:Today, we’re launching EmbeddingGemma 2 , expanding beyond text to unify code, images, video, and audio in a shared embedding space.
  • 原文依据:Built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference.
  • 原文依据:It can help find a specific video clip from a voice memo, or search through hours of audio recordings based on a text query, all processed by a single, natively multimodal model.
  • 原文依据:Built from the same technology as Gemini Embedding models, EmbeddingGemma 2 is: Best-in-class for its size: Achieves leading scores among sub-1B multimodal embedders for its size across benchmarks like MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark), while matching or outperforming many larger models across text, vision, and audio tasks.
  • 原文依据:Modular by design: Requires as little as 270M parameters for text-only workloads with optional vision (170M) and audio (300M) encoders for full multimodal support.
  • 原文依据:Storage-efficient: Using Matryoshka Representation Learning (MRL), developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions.
  • 原文依据:This provides up to 6x storage reduction for local vector databases and memory usage.
  • 原文依据:Optimized for on-device performance: Runs efficiently within tight resource constraints.
  • 原文依据:With quantization, on a Google Pixel 11 Pro, EmbeddingGemma 2 requires as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model.
  • 原文依据:Extended context ready: Features an 8K token context window (4x larger than EmbeddingGemma 1), allowing it to process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations thereof directly on local hardware.
  • 原文依据:Achieving top-tier quality for code, vision, and audio EmbeddingGemma 2 matches the strong multilingual text performance of EmbeddingGemma while delivering a significant 9.92-point improvement on code performance (in MTEB Code, from 68.76 to 78.68), making it well-suited for local codebase indexing, semantic code search, and coding agent retrieval.
  • 原文依据:Across image, video, documents, and audio, it sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size.
  • 原文依据:Enabling semantic search, routing, and retrieval, fully on-device EmbeddingGemma 2 brings robust capabilities directly to edge hardware.
  • 原文依据:Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and empowers developers to build cross-modal search and retrieval that works entirely offline.
  • 原文依据:When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that understand complex multimodal data.
  • 原文依据:Because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.
  • 原文依据:Getting started with EmbeddingGemma 2 We worked closely with the following partners to ensure EmbeddingGemma 2 works immediately where you build: Download the models: Find the model weights on Hugging Face and Kaggle , with Gemini Enterprise Agent Platform Model Garden availability coming soon.
  • 原文依据:Visit LiteRT Community on Hugging Face for models optimized for on-device.
  • 原文依据:On-device deployment: Develop cross-platform apps with Google AI Edge MediaPipe for turnkey embedding, retrieval & decision tasks or LiteRT for custom model integration.
  • 原文依据:Build for the browser with transformers.js or WebGPU .
  • 原文依据:Use your favorite development tools : Serve the model efficiently using transformers, sentence-transformers, MLX , vLLM, llama.cpp , SGLang, Ollama , and LMStudio.
  • 原文依据:Store your embedding vectors with Qdrant .
  • 原文依据:Fine-tuning: Follow guidance by Unsloth for how to fine-tune EmbeddingGemma 2 for your use cases.
  • 原文依据:Explore our developer guide , documentation , and guides for inference and fine-tuning .
AI 导读

Google DeepMind 发布 EmbeddingGemma 2,官方称这是面向端侧多模态嵌入的模型,可将文本、图像、音频和视频映射到统一嵌入空间。该模型基于 Gemma 4 架构,采用 Apache 2.0 许可,参数量 740M,官方表示文本任务最低需 270M 参数,视觉与音频编码器分别为 170M 和 300M。

来源:Tavily News Discovery · blog.google