EmbeddingGemma 2
Willison praises EmbeddingGemma 2's Apache 2.0 license and warns proprietary embeddings create costly re-embedding lock-in.
Simon Willison highlighted EmbeddingGemma 2 and welcomed its Apache 2.0 license in a Hacker News comment. He argued that embedding applications often compute and store thousands or millions of vectors, so a closed, hosted-only model creates lock-in if the vendor later withdraws it. A replacement model would still require paying to recompute existing vectors. He noted that in April 2024 OpenAI offered to cover the financial cost of some users re-embedding content.
- EmbeddingGemma 2 is released under the Apache 2.0 license.
- Embedding workloads can require millions of stored vectors.
- A discontinued proprietary model can force a costly full re-embed.
- OpenAI previously offered to pay some users' re-embedding costs in April 2024.
My comment on EmbeddingGemma 2 — Hacker News. I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license. For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model. Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison. If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors. (In April 2024 OpenAI offered to "cover the financial cost of users re-embedding…
This source does not provide full text. Read it at simonwillison.net.