Skip to content
AI-Daily-Builder

Tags · #inference

Nvidia, AMD, and CoreWeave all backed Tensormesh — KV-cache reuse becomes an inference primitive

Read this because Three rivals — Nvidia, AMD, CoreWeave — co-investing is the tell: KV-cache reuse (don't recompute what you already computed) is being treated as a foundational, neutral inference-stack layer. The economics of the inference era, in one round.

Tensormesh raised $20M from Nvidia, AMD, and CoreWeave and shipped Tensormesh Inference — productized KV-cache reuse claiming up to 10x lower latency/GPU cost.

OpenRouter raises $113M to own the switchboard between every AI model

Read this because The bet isn't a model — it's the layer above all of them. OpenRouter sells optionality: route to whatever model is best/cheapest, lock into none. The tell is the cap table — Nvidia plus Snowflake, Databricks, MongoDB, ServiceNow all want a seat at the routing layer.

OpenRouter raised a $113M CapitalG-led Series B at ~$1.3B (May 26) — routing 100T tokens/month across 400+ models for 8M+ users.

Anthropic in talks to rent Microsoft's Maia 200 AI chips — compute-crunch hedge

Read this because Silicon diversification, not a chip win. Anthropic already runs on Nvidia, Google TPUs, and AWS Trainium — adding Maia 200 makes it the first lab spanning all four silicon families. Optionality is the moat when compute is the bottleneck.

Anthropic is in talks to run Claude inference on Microsoft's Maia 200 chips via Azure (no deal signed, per CNBC May 21) — a hedge away from Nvidia + TPUs.

Tip