Do OpenAI's text-embedding-3 line and nomic-embed-text-v1.5 actually implement Matryoshka Representation Learning?
claim-matryoshka-representation-learning-truncatable-embeddings rests, for
its mechanism, on the Tier-1 MRL paper (Kusupati et al., arXiv:2205.13147). But
its production-adoption sub-claim — that OpenAI's text-embedding-3 models (their
dimensions shortening parameter) and nomic-embed-text-v1.5 both implement MRL —
carries an [unverified-mechanism] flag. The arXiv paper predates both models
and cannot attest to their adoption; the claim entered the vault via the
capture's hop findings, not a primary that states it directly.
Why it matters. This is a specific technical-mechanism claim (how these
systems work, not merely that they exist), which sources.md floors at Tier 1–2.
Confirming it lifts the flag and lets the MRL note move off seedling; it also
firms up the connection between MRL and
claim-openai-embedding-price-fell-5x-ada-002-to-3-small (the newer, cheaper
generation being the one that adopted truncatable dimensions).
What would answer it (Tier 1–2 primaries):
- OpenAI's own January 2024 announcement / API docs for text-embedding-3,
specifically the "shortening embeddings" /
dimensionsparameter description — does OpenAI attribute it to Matryoshka representation learning, or is that a third-party inference? - Nomic AI's nomic-embed-text-v1.5 release notes / model card / technical report stating that its truncatable dimensions are MRL-based.
Candidate next move. Fetch the OpenAI text-embedding-3 launch post and the Nomic v1.5 model card directly; quote the exact adoption language if present. If neither vendor names MRL, narrow the note to "MRL is the technique commonly cited behind these models' truncatable dimensions" rather than asserting adoption.