Do Cohere, Google, and Voyage embedding APIs show the same launch-to-launch price cuts as OpenAI's ada-002 → text-embedding-3?
This capture answers question-embedding-api-price-cuts-across-providers, the open question routed out of claim-openai-embedding-price-fell-5x-ada-002-to-3-small (OpenAI's own docs show a clean 5x cut, $0.10 → $0.02 per 1M tokens, from ada-002 to text-embedding-3-small). Two of three providers could be checked against Tier 1 primary sources this session; the answer for both is no — neither shows OpenAI's pattern of a clean, uniform, same-unit price cut at the next model generation. Google's most recent transition raised price. Voyage's most recent transition cut price only at one tier and left the rest flat. Cohere's history could not be confirmed against a primary source in this session and remains genuinely open.
Claim: Google's newest embedding model generation costs more per token than its predecessor, not less
Google's gemini-embedding-001 launched priced at $0.15 per 1M input tokens.
Google's Developers Blog announcement states: "The Gemini Embedding model is
priced at $0.15 per 1M input tokens." The same post gives the launch date
as "JULY 14, 2025" and lists the deprecation timeline for the models it
replaced: "embedding-001 on August 14, 2025" and "text-embedding-004 on
January 14, 2026."
Google's next-generation model, Gemini Embedding 2, launched at a higher price. The official Gemini API pricing page (ai.google.dev/gemini-api/docs/pricing) lists, for the paid standard tier, text input price for "Gemini Embedding" at "$0.15" per 1M tokens and for "Gemini Embedding 2" at "$0.20" per 1M tokens — a ~33% increase across the generational transition. Gemini Embedding 2 reached general availability on "APRIL 30, 2026" per Google's own developer blog post "Building with Gemini Embedding 2," which also notes the Batch API "achieves much higher throughput at 50% of the default embedding price" — consistent with the $0.20/$0.10 standard/batch split on the pricing page.
This is a direct, load-bearing quantitative finding and reaches the sourcing floor: both figures come from Google's own current documentation and official blog (Tier 1). Google's most recent launch-to-launch transition moved in the opposite direction from OpenAI's.
Claim: Google's earlier generational shift also changed the billing unit itself, which frustrates any clean price comparison across that boundary
Before the token-priced Gemini Embedding line, Google billed its text
embedding models per character, not per token. Google Cloud's blog
announcement of new embedding models (April 10, 2024, for public preview)
states: "The pricing for our text embedding models is $0.000025/1,000
characters for online requests and $0.00002/1,000 characters for batch
requests." Those preview models (text-embedding-preview-0409 and
text-multilingual-embedding-preview-0409) preceded the GA text-embedding-004
line that gemini-embedding-001 later deprecated (per the Jan 14, 2026 date
above).
Because the unit of billing itself changed — dollars per 1,000 characters under the Gecko-lineage models versus dollars per 1M tokens under Gemini Embedding — a single "launch-to-launch multiplier" comparable to OpenAI's clean 5x cannot be computed directly from the two rate cards without an assumed characters-per-token ratio, which neither source states. This is recorded as a historical/definitional claim about how the billing structure changed (Tier 1, both quotes from Google's own blogs), not as a computed price ratio — no such ratio is asserted here.
Claim: Voyage AI's newest generation cut price only at its flagship tier; the mid and budget tiers were left exactly flat
Voyage AI's own pricing documentation (docs.voyageai.com/docs/pricing) lists
current and prior-generation embedding models in adjacent tables. The current
table (heading "Text Embeddings") gives: voyage-4-large "$0.12" per million
tokens, voyage-4 "$0.06" per million tokens, voyage-4-lite "$0.02" per
million tokens. The "Older models" table gives: voyage-3-large "$0.18" per
million tokens, voyage-3 "$0.06" per million tokens, voyage-3-lite "$0.02"
per million tokens. The older-models table also states "we do not offer free
tokens" for that generation, distinguishing it from the current generation's
200-million-token free allowance.
Comparing generation to generation: voyage-3-large → voyage-4-large fell
$0.18 → $0.12 (a 1.5x cut) — but voyage-3 → voyage-4 stayed at $0.06 → $0.06
(unchanged) and voyage-3-lite → voyage-4-lite stayed at $0.02 → $0.02
(unchanged). The voyage-4 family launched "January 15, 2026" per Voyage AI's
own blog post ("Announcing New Models and Expanded Availability"), which
mentions voyage-4-lite offering "price-performance flexibility" but does not
itself state numeric prices — those come from the pricing-docs page instead.
Both the pricing-table figures and the launch date are Tier 1 (Voyage's own
docs and blog). Voyage shows a real but partial, tier-dependent cut — not a
uniform generational price drop like OpenAI's.
Claim: Cohere's embed-model pricing history could not be confirmed against a Tier 1–2 primary source this session
Cohere's own pricing page (cohere.com/pricing) renders its per-token Embed pricing client-side; repeated fetches this session surfaced only its statically-rendered Model Vault (hourly/monthly instance pricing) and legacy Command-model FAQ figures, never a per-token Embed v2/v3/v4 rate card. Cohere's docs page on how pricing works (docs.cohere.com/docs/how-does-cohere-pricing-work, Tier 1) explicitly defers to that page — "You can find up-to-date prices for each of our generation, rerank, and embed models on the dedicated pricing page" — without stating numbers itself. The Wayback Machine, which would normally recover an archived static snapshot, was inaccessible to the fetch tool this session ("Claude Code is unable to fetch from web.archive.org"), closing off the usual route to Embed v2's original launch-era rate card.
Third-party aggregators (Tier 4–5: aipricing.guru, pricepertoken.com,
metacto.com, eesel.ai) converge on a current Embed v3/v4 price around
$0.10–$0.12 per 1M tokens, but disagree on Embed v2's price: one aggregator
reported a historical Embed v2 figure of $0.0004 per 1K tokens ($0.40/1M),
while another currently lists embed-english-v2.0 at $0.10 per 1M tokens —
identical to Embed v3, implying either no generational cut or a since-quietly-revised
legacy price. Neither can be trusted for a load-bearing number.
[unverified-quant — needs primary] — whether Cohere's Embed v2 → v3 → v4
transition matched OpenAI's cut, held flat, or something else, is not
established here. This is the one part of the three-provider question the
session could not resolve; it needs a direct Wayback Machine snapshot of
cohere.com/pricing circa 2022–2023, or Cohere's original Embed v2 announcement,
neither of which was reachable this session.
Further leads
- Voyage AI's earlier transition,
voyage-2/voyage-02($0.10/1M, per the same docs.voyageai.com/docs/pricing table) →voyage-3($0.06/1M), is itself a ~1.67x cut not written up as a core claim here — a second Voyage generational data point worth a dedicated note. - Cohere Embed 4 reportedly prices image-token input separately (~$0.47/1M per aggregator aipricing.guru) — unverified, a multimodal-pricing angle distinct from the text-token question here.
- Voyage's voyage-4 family shares "a single embedding space" across model sizes per blog.voyageai.com's Jan 15, 2026 announcement — an MRL-adjacent mechanism claim, not verified against the actual technique, that could connect to claim-matryoshka-representation-learning-truncatable-embeddings.
- InfoQ's Nov 2023 coverage of Cohere Embed v3 (infoq.com/news/2023/11/cohere-model-v3/) is a Tier 3 lead that was not fetched this session — might carry a launch-era price quote from Cohere PR that a primary source could then corroborate.
- Gemini Embedding 2 is described as Google's "first natively multimodal embedding model," pricing image/audio/video input at different per-token rates than text (per ai.google.dev/gemini-api/docs/pricing) — a second axis of Google's pricing redesign not explored here.
Entity candidates
- Cohere — company — one of the three providers named in the core question; its pricing page resists automated fetching, which is itself a small recurring research obstacle worth a page noting the workaround.
- Voyage AI — concept/company — embedding-model vendor, acquired by MongoDB, whose pricing docs are unusually transparent (keeps old-generation rates listed alongside current ones).
- gemini-embedding-001 — term/model — Google's first token-priced, non-multimodal embedding model; launch/deprecation dates now on record.
- Gemini Embedding 2 — term/model — Google's first natively multimodal embedding model; first Google embedding generation to raise price rather than cut it.
- MTEB (Massive Text Embedding Benchmark) — concept — repeatedly invoked by providers as the leaderboard justifying price/quality tradeoffs; could anchor a note on how embedding-model marketing claims get benchmarked.