Skip to content
AI-Daily-Builder

Benchmarks

Benchmarks run
8
head-to-head
Latest run
2026-05-29
Most recent benchmark
Models compared
4+
Claude · GPT · Gemini · open-source
Metrics tracked
4
Latency · tokens · cost · verdict
Benchmarks run per month
NaN NaN NaN

2026-05-29 · benchmark

SWE-bench Verified — May 2026 agentic-coding leaderboard (pass@1 %)

Show prompt
SWE-bench Verified is 500 human-verified real GitHub issues from popular open-source Python repos. For each, the model-driven agent must read the issue, locate the files to change, write a patch, apply it, and pass the repository's hidden test suite — no hints, scored as the percentage of issues fully resolved (pass@1). This card reports PUBLISHED accuracy from public leaderboards as of 2026-05-28; it is not a latency benchmark.
Model Latency Cost Verdict
Claude Mythos Preview (restricted) 0ms Win
GPT-5.5 0ms Win
Claude Opus 4.8 0ms Win
Claude Opus 4.7 (Adaptive) 0ms Tie
GPT-5.3-Codex 0ms Tie
Gemini 3.1 Pro 0ms Loss
DeepSeek V4 Pro Max (open-weight) 0ms Loss
Show responses

Claude Mythos Preview (restricted)

93.9% resolved · #1 · public leaderboard (restricted-access model)

GPT-5.5

88.7% resolved · public leaderboard

Claude Opus 4.8

88.6% resolved · public leaderboard

Claude Opus 4.7 (Adaptive)

87.6% resolved · public leaderboard

GPT-5.3-Codex

85.0% resolved · public leaderboard

Gemini 3.1 Pro

80.6% resolved · public leaderboard

DeepSeek V4 Pro Max (open-weight)

80.6% resolved · public leaderboard · best open-weight

Published accuracy leaderboard, NOT a measured-latency run: `latency_ms` is set to 0 (not applicable) and token/cost fields are omitted on every row — the verified datum is the pass@1 % in each `response`. Scores are compiled from public SWE-bench Verified leaderboards (swebench.com, llm-stats.com, benchlm.ai, andrew.ooo, marc0.dev) snapshotted ~2026-05-28; exact numbers vary by harness, scaffold, and snapshot date, so treat ±1-2 points as noise. Verdict tiers (accuracy, not speed): win = 88%+, tie = 84-87.9%, loss = under 84%. Claude Mythos Preview is a restricted-access model; its 93.9% is reported but most teams cannot run it. Takeaways: (1) the frontier has compressed — the top three (Mythos 93.9, GPT-5.5 88.7, Opus 4.8 88.6) are within ~5 points; (2) agentic-coding pass rates above 88% mean the benchmark itself is saturating and SWE-bench Pro / Terminal-Bench Hard are becoming the better discriminators; (3) open-weight DeepSeek V4 Pro Max at 80.6 trails the closed frontier by ~13 points but is closing.

2026-05-24 · benchmark

DGX Spark (GB10) local-model throughput — prefill & decode tok/s across 13 model/quant/engine combos

Show prompt
Standardized single-stream (batch size 1) inference on one DGX Spark (GB10, 128 GB LPDDR5X unified memory, ~273 GB/s bandwidth, ~1 PFLOP FP4): 2,048-token input, 128-token output (ISL/OSL 2048/128). Each row is a model + quantization + inference engine. We report prompt-processing throughput (prefill, 'pp') and token-generation throughput (decode, 'tg') in tokens/sec. Latency shown is the modeled time to emit 128 tokens at the published decode rate (128 / tg × 1000).
Model Latency Cost Verdict
GPT-OSS-20B · MXFP4 · llama.cpp 1547ms Win
Qwen3.5-35B-A3B · MXFP4 · llama.cpp 2207ms Win
GPT-OSS-120B · MXFP4 · llama.cpp 2312ms Win
Qwen2.5-VL-7B · NVFP4 · TRT-LLM (vision) 3069ms Win
Llama 3.1 8B · NVFP4 · TRT-LLM 3312ms Win
Qwen3-Coder-30B-A3B · Q8_0 · llama.cpp 4129ms Win
Qwen3.6-27B · Q4_K_M +MTP · llama.cpp 4523ms Tie
Gemma 4 26B-A4B · F16 · llama.cpp 4830ms Tie
Qwen3-14B · NVFP4 · TRT-LLM 5637ms Tie
Llama 3.1 8B · FP8 · SGLang 6244ms Tie
Qwen3.6-27B · Q4_K_M · llama.cpp 9771ms Tie
Llama 3.1 70B · FP8 · SGLang 47407ms Loss
Qwen3-235B · NVFP4 · TRT-LLM (DUAL Spark) 10912ms
Show responses

GPT-OSS-20B · MXFP4 · llama.cpp

3670.42 pp / 82.74 tg tok/s · llama.cpp · NVIDIA-official

Qwen3.5-35B-A3B · MXFP4 · llama.cpp

prefill n/p / ~58 tg tok/s · llama.cpp · community (MoE A3B; theoretical ceiling ~91)

GPT-OSS-120B · MXFP4 · llama.cpp

1725.47 pp / 55.37 tg tok/s · llama.cpp · NVIDIA-official (canonical official 120B decode; engine spread 35 llama.cpp deep-ctx → 41 Ollama → ~50 SGLang)

Qwen2.5-VL-7B · NVFP4 · TRT-LLM (vision)

65831.77 pp / 41.71 tg tok/s · TRT-LLM · NVIDIA-official

Llama 3.1 8B · NVFP4 · TRT-LLM

10256.9 pp / 38.65 tg tok/s · TRT-LLM · NVIDIA-official

Qwen3-Coder-30B-A3B · Q8_0 · llama.cpp

1308 pp / 31 tg tok/s · llama.cpp · community (llama.cpp #16578; MoE A3B)

Qwen3.6-27B · Q4_K_M +MTP · llama.cpp

719 pp / 28.3 tg tok/s · llama.cpp +MTP (5 draft) · community (2.16x decode vs no-MTP)

Gemma 4 26B-A4B · F16 · llama.cpp

prefill n/p / ~26.5 tg tok/s · llama.cpp · community (MoE A4B; theoretical ~34)

Qwen3-14B · NVFP4 · TRT-LLM

5928.95 pp / 22.71 tg tok/s · TRT-LLM · NVIDIA-official

Llama 3.1 8B · FP8 · SGLang

7991 pp / 20.5 tg tok/s · SGLang · community (FP8 decode ~half of NVFP4 — same model)

Qwen3.6-27B · Q4_K_M · llama.cpp

1084 pp / 13.1 tg tok/s · llama.cpp · community (single-stream, no spec-decode)

Llama 3.1 70B · FP8 · SGLang

~803 pp / ~2.7 tg tok/s · SGLang · community (barely fits 128 GB; KV+weights thrash — avoid dense 70B FP8 on one unit)

Qwen3-235B · NVFP4 · TRT-LLM (DUAL Spark)

23477.03 pp / 11.73 tg tok/s · TRT-LLM · NVIDIA-official · DUAL DGX Spark over ConnectX-7 (does not fit one unit at usable quant)

Single DGX Spark GB10 (128 GB LPDDR5X, 273 GB/s). 'pp' = prompt-processing / prefill tok/s; 'tg' = token-generation / decode tok/s. Verdict tiers by single-stream decode: win = 30+ tok/s (snappy interactive), tie = 10-30 (usable), loss = under 10 (impractical). Source class is tagged per row: NVIDIA-official = developer.nvidia.com 'How DGX Spark Performance Enables Intensive AI Tasks' (ISL/OSL 2048/128, BS=1); community = NVIDIA dev forums / llama.cpp issues / SGLang hands-on. Takeaways: (1) Decode is memory-bandwidth-bound — tg tok/s roughly equals active-param-bytes-per-token divided by 273 GB/s, so MoE (A3B) and lower-bit quant lift it. (2) Prefill is compute-bound on Blackwell FP4 cores — routinely thousands of tok/s regardless of model size. (3) Quant format matters: NVFP4/MXFP4 is about 2x FP8 decode (Llama 3.1 8B: 38.65 NVFP4 vs 20.5 FP8). (4) Speculative MTP roughly doubles single-stream decode (Qwen3.6-27B: 13.1 to 28.3) but regresses under concurrency. (5) Dense 70B at FP8 barely fits 128 GB and thrashes (~2.7 tg) — avoid on a single unit. (6) 235B needs two Sparks over ConnectX-7. Compiled from published benchmarks; all figures single-unit unless flagged DUAL.

2026-05-24 · benchmark

LEFT JOIN silently becomes INNER JOIN — SQL correctness head-to-head

Show prompt
Two tables: customers(id, name) and orders(id, customer_id, total). A teammate wrote this query to list every customer with their number of orders, INCLUDING customers who have placed zero orders:

SELECT c.id, c.name, COUNT(o.id) AS order_count
FROM customers c
LEFT JOIN orders o ON o.customer_id = c.id
WHERE o.total > 0
GROUP BY c.id, c.name;

But customers with zero orders are missing from the result. Explain exactly why, and rewrite the query so zero-order customers appear with order_count = 0. Return the corrected SQL and a one-sentence explanation.
Model Latency Cost Verdict
claude-opus-4-7 4180ms $0.016 Win
gpt-5 3360ms $0.007 Win
gemini-3-pro 2980ms $0.002 Tie
qwen3.6-35b-a3b-nvfp4 (local DGX Spark) 9100ms Loss
Show responses

claude-opus-4-7

The WHERE o.total > 0 runs AFTER the LEFT JOIN. For a customer with zero orders every o.* column is NULL, and NULL > 0 is not TRUE, so those rows are dropped — collapsing the LEFT JOIN into an INNER JOIN. Move the predicate into the ON clause so it filters the join, not the result:

SELECT c.id, c.name, COUNT(o.id) AS order_count
FROM customers c
LEFT JOIN orders o ON o.customer_id = c.id AND o.total > 0
GROUP BY c.id, c.name;

Unmatched customers keep one all-NULL row and COUNT(o.id) correctly returns 0.

gpt-5

A post-join WHERE on a right-table column nullifies the outer join: zero-order customers produce NULL o.total, and NULL > 0 fails, so they are filtered out. Relocate the filter to the join condition:

SELECT c.id, c.name, COUNT(o.id) AS order_count
FROM customers c
LEFT JOIN orders o ON o.customer_id = c.id AND o.total > 0
GROUP BY c.id, c.name;

COUNT(o.id) skips NULLs, so zero-order customers report 0.

gemini-3-pro

Right — the WHERE clause removes the NULL rows produced by the LEFT JOIN, so customers with no orders disappear. Keep them by also allowing the NULL case:

SELECT c.id, c.name, COUNT(o.id) AS order_count
FROM customers c
LEFT JOIN orders o ON o.customer_id = c.id
WHERE o.total > 0 OR o.id IS NULL
GROUP BY c.id, c.name;

This restores the zero-order customers.

qwen3.6-35b-a3b-nvfp4 (local DGX Spark)

The problem is COUNT with GROUP BY: COUNT(o.id) returns NULL when there are no matching orders, so wrap it in COALESCE to show 0:

SELECT c.id, c.name, COALESCE(COUNT(o.id), 0) AS order_count
FROM customers c
LEFT JOIN orders o ON o.customer_id = c.id
WHERE o.total > 0
GROUP BY c.id, c.name;

This defaults missing counts to 0.

The trap: a WHERE predicate on a column from the right table of a LEFT JOIN runs AFTER the join. Unmatched (zero-order) customers get all-NULL order columns, and `NULL > 0` is not TRUE, so those rows are filtered out — silently collapsing the LEFT JOIN into an INNER JOIN. The clean fix is to move the predicate into the ON clause (`LEFT JOIN orders o ON o.customer_id = c.id AND o.total > 0`), so it filters the join rather than the result; COUNT(o.id) then ignores the remaining NULL and returns 0. Verdict criteria: win = correct diagnosis (WHERE nullifies the outer join) AND clean ON-clause fix; tie = correct diagnosis but a clunky/partially-correct fix (e.g. `WHERE ... OR o.id IS NULL`, which mishandles customers whose only orders have total <= 0); loss = wrong diagnosis or a fix that still drops zero-order rows. Latency / token / cost figures are representative for a short single-turn task; the local model row reflects a Qwen3.6-35B-A3B NVFP4 on a DGX Spark.

2026-05-09 · benchmark

Fix off-by-one in cursor pagination — agentic code edit head-to-head

Show prompt
Here is a Node.js handler for cursor pagination:

```js
function paginate(items, cursor, limit = 20) {
  const start = cursor ? items.findIndex(i => i.id === cursor) : 0;
  return {
    data: items.slice(start, start + limit),
    next: items[start + limit]?.id ?? null,
  };
}
```

It has a subtle off-by-one bug: when called with a cursor, the first item returned is the cursor itself (it should be the item *after* the cursor). Fix the function. Preserve the same signature and return shape. Add a single inline comment only where the bug was. Return only the corrected function, no preamble.
Model Latency Cost Verdict
claude-opus-4-7 2240ms $0.011 Win
gpt-5 1890ms $0.009 Win
gemini-3-pro 2050ms $0.007 Tie
qwen3.6-35b-a3b-nvfp4 920ms $0.001 Win
Show responses

claude-opus-4-7

function paginate(items, cursor, limit = 20) {
  const start = cursor ? items.findIndex(i => i.id === cursor) + 1 : 0; // bug fix: skip cursor itself
  return {
    data: items.slice(start, start + limit),
    next: items[start + limit]?.id ?? null,
  };
}

gpt-5

function paginate(items, cursor, limit = 20) {
  // bug fix: cursor item itself was being returned; advance past it
  const start = cursor ? items.findIndex(i => i.id === cursor) + 1 : 0;
  return {
    data: items.slice(start, start + limit),
    next: items[start + limit]?.id ?? null,
  };
}

gemini-3-pro

function paginate(items, cursor, limit = 20) {
  let start = 0;
  if (cursor) {
    const idx = items.findIndex(i => i.id === cursor);
    if (idx >= 0) start = idx + 1; // bug fix: cursor was inclusive
  }
  return {
    data: items.slice(start, start + limit),
    next: items[start + limit]?.id ?? null,
  };
}

qwen3.6-35b-a3b-nvfp4

function paginate(items, cursor, limit = 20) {
  const start = cursor ? items.findIndex(i => i.id === cursor) + 1 : 0; // off-by-one fix
  return {
    data: items.slice(start, start + limit),
    next: items[start + limit]?.id ?? null,
  };
}

Tests whether the model spots the off-by-one (start should be findIndex(...) + 1) without breaking the cursor=null branch. Latency client-side. Cost from public May 2026 pricing. All models given the same exact prompt.

2026-05-09 · benchmark

JSON extraction from messy meeting notes — schema fidelity test

Show prompt
Extract the following from these meeting notes into strict JSON matching this exact schema:

{
  "date": "YYYY-MM-DD",
  "attendees": ["name1", "name2"],
  "action_items": [{"owner": "name", "task": "...", "due": "YYYY-MM-DD or null"}]
}

Notes:
---
Met today (it's the 9th of May 2026, finally Friday) — me, Sarah Chen, Raj P, and Tomás joined late from his car. Sarah said she'd ship the auth migration by next Wed (so 13th?). Raj is still chasing the SSO config from the vendor — no firm date, said "sometime next week". Tomás owes a draft of the launch post but said no rush. I'll handle the customer interview write-up by EOD Monday.
---

Return ONLY the JSON. No preamble, no markdown fences, no trailing comments.
Model Latency Cost Verdict
claude-opus-4-7 1620ms $0.014 Win
gpt-5 1340ms $0.012 Tie
gemini-3-pro 1560ms $0.008 Loss
qwen3.6-35b-a3b-nvfp4 720ms $0.001 Tie
Show responses

claude-opus-4-7

{"date":"2026-05-09","attendees":["Sarah Chen","Raj P","Tomás"],"action_items":[{"owner":"Sarah Chen","task":"Ship the auth migration","due":"2026-05-13"},{"owner":"Raj P","task":"Chase SSO config from vendor","due":null},{"owner":"Tomás","task":"Draft of the launch post","due":null},{"owner":"me","task":"Customer interview write-up","due":"2026-05-11"}]}

gpt-5

{"date":"2026-05-09","attendees":["Sarah Chen","Raj P","Tomás"],"action_items":[{"owner":"Sarah Chen","task":"Ship the auth migration","due":"2026-05-13"},{"owner":"Raj P","task":"Chase SSO config from vendor","due":null},{"owner":"Tomás","task":"Draft launch post","due":null},{"owner":"narrator","task":"Customer interview write-up","due":"2026-05-12"}]}

gemini-3-pro

```json
{"date":"2026-05-09","attendees":["Sarah Chen","Raj P","Tomás"],"action_items":[{"owner":"Sarah Chen","task":"ship auth migration","due":"2026-05-13"},{"owner":"Raj P","task":"SSO config from vendor","due":null}]}
```

qwen3.6-35b-a3b-nvfp4

{"date":"2026-05-09","attendees":["Sarah Chen","Raj P","Tomás"],"action_items":[{"owner":"Sarah Chen","task":"Ship auth migration","due":"2026-05-13"},{"owner":"Raj P","task":"Chase SSO config from vendor","due":null},{"owner":"Tomás","task":"Draft launch post","due":null},{"owner":"author","task":"Customer interview write-up","due":"2026-05-12"}]}

Tests: date inference (today = 2026-05-09), name normalization (Sarah Chen / Raj P → Raj Patel?? leave as-is), null due dates ('sometime next week' / 'no rush' = null), and adherence to no-prose constraint. Reject responses with markdown fences.

2026-05-09 · benchmark

Needle in a haystack at 500K tokens — long-context recall

Show prompt
You will receive a 500,000-token document containing the full text of 12 NeurIPS papers concatenated. Buried at character offset 1,847,392 is a single sentence: 'The secret access code for the May 2026 builder-daily benchmark is QUARTZ-7392-DELTA.'

Question: What is the secret access code? Return only the code itself, nothing else.

Document follows below the marker.

--- DOCUMENT START ---
[~500,000 tokens of NeurIPS paper text]
--- DOCUMENT END ---
Model Latency Cost Verdict
claude-opus-4-7 18420ms $7.535 Win
gpt-5 22180ms $6.279 Win
gemini-3-pro 14620ms $0.628 Win
qwen3.6-35b-a3b-nvfp4 (262K cap) 0ms $0.000 Error
Show responses

claude-opus-4-7

QUARTZ-7392-DELTA

gpt-5

QUARTZ-7392-DELTA

gemini-3-pro

QUARTZ-7392-DELTA

qwen3.6-35b-a3b-nvfp4 (262K cap)

 

Pure recall test at the 70% depth point of 500K input. Latency includes prompt processing (which is dominant at this scale). Cost based on full 500K tokens in. Verdict 'win' = exact match 'QUARTZ-7392-DELTA'. Models tested at vendor-claimed max context.

2026-05-09 · benchmark

Pick the right tool — multi-step task with 6 available functions

Show prompt
You have these tools available: `web_search(query)`, `fetch_url(url)`, `read_file(path)`, `run_shell(cmd)`, `query_database(sql)`, `send_email(to, subject, body)`.

Task: Find the latest published date of Anthropic's Claude 4.7 release announcement, fetch the full text of that page, count how many times the word 'tool' appears, and email the count to ops@example.com with subject 'Claude 4.7 tool count'.

Respond with ONLY a JSON array of tool calls in execution order: [{"tool": "...", "args": {...}}, ...]. No prose, no markdown, no other text.
Model Latency Cost Verdict
claude-opus-4-7 1480ms $0.013 Win
gpt-5 1320ms $0.011 Tie
gemini-3-pro 1810ms $0.009 Loss
kimi-k2.6 980ms $0.001 Win
Show responses

claude-opus-4-7

[
  {"tool": "web_search", "args": {"query": "Anthropic Claude 4.7 release announcement date site:anthropic.com"}},
  {"tool": "fetch_url", "args": {"url": "<url-from-prior-result>"}},
  {"tool": "send_email", "args": {"to": "ops@example.com", "subject": "Claude 4.7 tool count", "body": "<count>"}}
]

gpt-5

[
  {"tool": "web_search", "args": {"query": "Claude 4.7 release announcement Anthropic"}},
  {"tool": "fetch_url", "args": {"url": "<top result url>"}},
  {"tool": "run_shell", "args": {"cmd": "echo $TEXT | grep -o tool | wc -l"}},
  {"tool": "send_email", "args": {"to": "ops@example.com", "subject": "Claude 4.7 tool count", "body": "<count>"}}
]

gemini-3-pro

I'll help you with that. Here's the plan: [
  {"tool": "web_search", "args": {"query": "Claude 4.7"}},
  {"tool": "web_search", "args": {"query": "Anthropic announcement"}},
  {"tool": "fetch_url", "args": {"url": "..."}},
  {"tool": "send_email", "args": {"to": "ops@example.com", "subject": "Result", "body": "..."}}
]

kimi-k2.6

[
  {"tool": "web_search", "args": {"query": "Anthropic Claude 4.7 release announcement"}},
  {"tool": "fetch_url", "args": {"url": "<from search>"}},
  {"tool": "send_email", "args": {"to": "ops@example.com", "subject": "Claude 4.7 tool count", "body": "<count>"}}
]

Tests whether the model picks the right 4 tools in the right order: web_search → fetch_url → (count locally, no tool) → send_email. Penalize: extra tools, wrong order, missing send_email, prose preamble.

2026-05-02 · benchmark

Summarize 1000-word post into 3 bullets — frontier head-to-head

Show prompt
Summarize this 1000-word blog post about Anthropic's Claude 4.7 release into exactly 3 bullets, each ≤ 25 words. Return only the bullets, no preamble.
Model Latency Cost Verdict
claude-opus-4-7 1240ms $0.022 Win
gpt-5 980ms $0.018 Tie
gemini-3-pro 1120ms $0.014 Loss
Show responses

claude-opus-4-7

• Anthropic shipped Claude 4.7 with 1M-token context now extended to the Opus tier.
• Tool-use determinism improved on high-fan-out parallel calls; cache TTL stays at 5 min.
• Early users report ~12% p95 latency drop on Sonnet/Haiku — Opus latency improvement claimed but unverified.

gpt-5

• Claude 4.7 released by Anthropic, headline: 1M context arrives on the Opus tier.
• Anthropic claims ~12% p95 latency improvement across the model family.
• Tool-call behavior more deterministic in parallel/high-fan-out scenarios.

gemini-3-pro

• Anthropic released Claude 4.7 today with several updates.
• The 1M context window is now available on Opus tier in addition to Sonnet and Haiku, which is a major change for long-context agentic workflows.
• Latency was reduced by approximately 12% according to Anthropic, and tool-use behavior was improved.

Same prompt, same input. Latency measured from API call start to last byte. Cost from official $/Mtok pricing (May 2026).

Tip