Skip to content
AI-Daily-Builder

Archive

Google Puts Gemini 3.5 Pro Into Final Preview Stretch With 2M-Token Context and Deep Think Reasoning

Read this because The pricing and the free-trial window are what builders should check first. At $15/$60 per million tokens, Pro is priced to be used seriously — not sampled.

Gemini 3.5 Pro is in final Vertex enterprise preview with a 2M-token context window and Deep Think reasoning; GA launch expected imminently in June 2026.

Intel and Foxconn Launch Rack-Scale AI Partnership at Computex, Taking Aim at Nvidia GB200

Read this because Intel is not trying to beat Nvidia on GPU performance — it is betting on a full-stack alternative-supply play backed by Foxconn's manufacturing scale. The +4.43% stock reaction shows investors see it as credible.

Intel and Foxconn unveiled a rack-scale AI co-development partnership at Computex 2026, targeting Nvidia GB200 NVL with a 100 kW liquid-cooled Xeon rack.

Ramp Raises $750M Series F at $44B Valuation, Bets AI Agents Will Own Corporate Spend

Read this because The most telling data point: one customer burned its entire 2026 AI budget in four months. Ramp is building the observability and governance layer for AI spend — a category that barely existed a year ago.

Ramp raised $750M at a $44B valuation, betting AI agents will own corporate spend and launching the first corporate card purpose-built for AI agents.

Anthropic ships Claude Fable 5: a Mythos-class frontier model the public can finally touch

Read this because One frontier model, two releases: a safeguarded public version and a restricted full version. The pricing, the free window, and the safety-classifier fallback are what builders should check first.

Anthropic released Claude Fable 5 — its first publicly available Mythos-class model — at $10/$50 per 1M tokens, with a restricted Mythos 5 for vetted partners.

Apple rebuilds Siri on new Foundation Models at WWDC 2026, with a Google Gemini assist and a 12GB hardware floor

Read this because Tim Cook's last keynote pairs a genuinely rebuilt Siri with two asterisks practitioners should not skip: a 12GB-unified-memory device floor and EU/China launch exclusions.

WWDC 2026 unveiled Siri AI on next-gen Apple Foundation Models built with Google's Gemini, a dedicated app, an expanded developer framework, and EU/China gaps.

Microsoft Build 2026: 'Microsoft IQ' makes context the platform — Work IQ, Foundry IQ and a new Web IQ go live

Read this because We skipped the obvious MAI-models headline because the durable Build 2026 story is structural: Microsoft wants grounded context to be a default platform layer, not a feature. Watch Web IQ — a low-token web-grounding broker.

At Build 2026 Microsoft folded grounding into one brand, Microsoft IQ: Work IQ, Foundry IQ, Fabric IQ and a new Web IQ — context as managed infrastructure.

Microsoft launches seven in-house MAI models, including its first reasoning model, to cut OpenAI dependence

Read this because A big platform owner shipping its own reasoning model and openly framing it as a way to pay OpenAI less is a structural story, not just a product drop — it reprices the whole frontier-model supply chain.

At Build 2026 on June 2, Microsoft shipped seven MAI models built from scratch, led by reasoning model MAI-Thinking-1, to lower costs and reduce OpenAI

Trump's June 2 AI order asks labs for 30-day early access to frontier models

Read this because A "voluntary 30-day look" sounds modest, but it quietly establishes the federal government as a pre-release gatekeeper for the most capable models — a precedent that matters more than the day-one mechanics.

On June 2, 2026, President Trump signed "Promoting Advanced Artificial Intelligence Innovation and Security," asking AI developers to voluntarily give the

Anthropic Confidentially Files for IPO at a $965B Valuation, Eyeing a Fall 2026 Debut

Read this because A confidential S-1 is an on-ramp, not a commitment. But filing the year it crossed a ~$50B run-rate reframes Anthropic from research lab to public-scale enterprise software vendor, racing OpenAI to the listing window.

Anthropic confidentially filed a draft S-1 with the SEC on June 1, 2026 at a $965B valuation, with revenue run-rate nearing $50B.

NVIDIA and TSMC Push AI Deep Into the Fabs: cuLitho Cuts Lithography Cost 20-50%, FabTwin Goes Digital

Read this because Not the capex headline: NVIDIA is selling compute into the supply chain that builds NVIDIA's own chips. cuLitho's 20-50% litho cost/cycle-time cut is the load-bearing number — it gates how fast and cheaply sub-2nm wafers reach volume. A vertical loop.

At GTC Taipei, TSMC adopted NVIDIA's CUDA-X stack across lithography, simulation and inspection, with cuLitho cutting litho cost up to 50%.

xAI finishes training Grok V9-Medium, a 1.5T-param model tuned on Cursor developer data

Read this because The headline isn't the 1.5T parameter count — it's the corpus. Tuning a frontier model on Cursor's real developer workflows is a direct bid for the coding layer Claude and Codex dominate. Treat the benchmarks and timeline as vendor-sourced until weights or an API ship.

Musk says xAI's 1.5T-param Grok V9-Medium finished training (May 25), ~3x its production model and trained on Cursor dev data — mid-June release expected.

Cognition raises $1B at $26B — the agent-as-headcount bet, with 90% of its own code AI-written

Read this because A ~53x ARR multiple is a bet on agent-as-headcount, not agent-as-tool. The flywheel is the proof and the risk: Cognition writes ~90% of its own code with Devin, so its growth and its demo are the same thing — until growth slows.

Cognition, maker of the Devin coding agent, raised $1B+ at a $26B valuation (May 27) — 2.5x in 8 months, $492M ARR, ~90% of its own code AI-written.

Google's AI Threat Defense auto-patches vulnerabilities at machine speed — the defensive answer to AI attackers

Read this because Same week one model hunts vulnerabilities, Google ships one that auto-patches them. When attack and defense both run at machine speed, the patch window collapses from weeks to minutes — and the human moves from operator to auditor of agent-written fixes.

Google launched AI Threat Defense (May 27): a Gemini platform fusing Wiz, CodeMender, and Mandiant to find and auto-patch vulnerabilities at machine speed.

Nvidia, AMD, and CoreWeave all backed Tensormesh — KV-cache reuse becomes an inference primitive

Read this because Three rivals — Nvidia, AMD, CoreWeave — co-investing is the tell: KV-cache reuse (don't recompute what you already computed) is being treated as a foundational, neutral inference-stack layer. The economics of the inference era, in one round.

Tensormesh raised $20M from Nvidia, AMD, and CoreWeave and shipped Tensormesh Inference — productized KV-cache reuse claiming up to 10x lower latency/GPU cost.

Micron joins the $1 trillion club as HBM stops being a commodity

Read this because The number to watch isn't the trillion — it's the word "structural." UBS's 204% target hike rests on the thesis that AI memory has de-cyclicalized: HBM4 sold out, DRAM up 58-63%, pricing locked through 2029. If memory has escaped its cycle, the entire semi playbook changes.

Micron crossed a $1T market cap on May 26 — the 5th chipmaker — after UBS tripled its target to $1,625 and 2026 HBM4 capacity sold out under long-term deals.

An OpenAI reasoning model disproved an 80-year-old Erdős conjecture — and it wasn't a math-specific model

Read this because The headline is "AI does math." The real signal: it came from a general-purpose reasoning model, not a math-specific system — and it disproved an 80-year belief by constructing a counterexample no human had found. One result is not a revolution, but the generality is the story.

OpenAI says an internal reasoning model autonomously disproved the Erdős 1946 unit-distance conjecture (May 20) — a first for AI, verified by mathematicians.

OpenRouter raises $113M to own the switchboard between every AI model

Read this because The bet isn't a model — it's the layer above all of them. OpenRouter sells optionality: route to whatever model is best/cheapest, lock into none. The tell is the cap table — Nvidia plus Snowflake, Databricks, MongoDB, ServiceNow all want a seat at the routing layer.

OpenRouter raised a $113M CapitalG-led Series B at ~$1.3B (May 26) — routing 100T tokens/month across 400+ models for 8M+ users.

AMD's 256-core EPYC "Venice" is the first HPC chip to ramp on TSMC 2nm

Read this because Everyone watches the GPU. But AI clusters still need a host CPU to feed them, and AMD just put a 256-core server part on the most advanced node before anyone else. The lever here is efficiency at the power wall — and a quiet Arizona on-shoring story riding alongside it.

AMD is ramping EPYC "Venice" on TSMC 2nm (May 21) — a 256-core/512-thread part, the industry's first HPC product on the node, 70%+ uplift over Turin.

Japan's megabanks to get Anthropic's Mythos — frontier-model access as statecraft

Read this because The story isn't the model — it's the channel. Mythos access arrives via a US Treasury visit, not a sales call. A frontier model that hunts software vulnerabilities is now economic diplomacy: gated, allied-only, governed by a national working group before a query runs.

Japan's 3 megabanks will get Anthropic's vulnerability-hunting model Mythos by end-May — access conveyed in Tokyo by US Treasury's Bessent.

Tip