Current outcome
Google DeepMind has open-sourced EmbeddingGemma 2, an Apache 2.0-licensed 740M-parameter multimodal embedding model unifying text, code, image, video and audio retrieval, with flexible on-device loading and vector dimension reduction, improving code retrieval benchmarks over the prior generation.
Progress timeline
2 material updates- #01
EmbeddingGemma 2: An open, lightweight multimodal embedding model
Google released EmbeddingGemma 2, an Apache 2.0-licensed lightweight multimodal embedding model for on-device and local use, prompting technical discussion about its license, MRL-based training and comparisons to rival embedding models.
Source evidence: ilreb
- #02
740M跑手机,谷歌EmbeddingGemma 2一次搜遍文字、图片和音视频
Adds technical details: full model is 740M parameters unifying text, code, image, video and audio retrieval; supports flexible loading (270M text/code, 440M with vision, 740M full); context increased from 2K to 8K; MTEB Code score rose from 68.76 to 78.68; 768-dim vectors can be compressed to 512/256/128; quantized memory on Pixel 11 Pro about 191MB/567MB.
State after update: Google DeepMind has open-sourced EmbeddingGemma 2, an Apache 2.0-licensed 740M-parameter multimodal embedding model unifying text, code, image, video and audio retrieval, with flexible on-device loading and vector dimension reduction, improving code retrieval benchmarks over the prior generation.
Source evidence: theblockbeats