Live evaluation · Terminal-Bench 4.0

GLM-5.3 MXFP4

A public operations view of the running five-attempt evaluation. The board distinguishes scored model outcomes from infrastructure faults, and refreshes without redeploying the site.

正在运行的五次尝试评测公开看板。 它把模型得分结果与基础设施故障分开记录, 无需重新部署网站即可刷新。

0%scored已计分

Attempt ledger尝试台账

The denominator is fixed at 63 tasks × 5 attempts. An attempt advances this bar only when a verifier returns a reward; infrastructure failures remain visible but do not consume the target.

分母固定为 63 个任务 × 5 次尝试。 只有 verifier 返回 reward 的尝试才推进这条进度; 基础设施故障保留记录, 但不占用目标次数。

0 / 315—
Scored attempts已计分尝试—
of 315 fixed target固定目标 315
Complete tasks完整任务—
five scored attempts each每项均有 5 次计分
Observed rate实测速率—
terminal trials per hour终态 trial / 小时
Main-run ETA主跑预计剩余—
at observed rate, regrade excluded按实测速率, 不含 regrade

Verified correctnessVerifier 正确率

Binary task rewards · live二元任务 reward · 实时
Current scored attempts当前已计分尝试 — —
Completed-task matched完整任务同项比较
—MXFP4 current
—Official API
—
Official same 63 CPU tasks官方同一组 63 个 CPU 任务 — —
Official full 66-task row官方完整 66 项结果 41.82% 138 / 330 · 95% CI ±3.23 pp
The comparison always uses a task-matched denominator. Before completion, the all-attempt rate is biased by task completion order and the matched card uses only finished tasks. At 63/63 complete it becomes the final CPU-subset comparison. The official row used an opaque hosted API; this MXFP4 run uses explicit quantized weights, so the line is an external end-to-end reference, not a same-runtime quantization A/B. 比较始终使用同 task 分母。 完成之前, 全部 attempt 比例会受到任务完成顺序的偏置, 同项卡片只使用已完成 task; 达到 63/63 后, 它才成为最终 CPU 子集比较。 官方结果来自不透明的 hosted API, 当前运行使用明确的 MXFP4 量化权重, 因此这是一条端到端外部参考线, 不是同 runtime 的量化 A/B。
Task任务 Current MXFP4当前 MXFP4 Official GLM-5.3 API官方 GLM-5.3 API Δ when complete完成项差值 State状态
No tasks match this view.没有符合条件的任务。

Serving pools模型服务池

Configured concurrency 24配置总并发 24
Replica A—
Terminal终态—
Running运行中—
Pending等待—
Retries重试—
Active requests活跃请求—
Waiting / tokens等待 / token—
Replica B—
Terminal终态—
Running运行中—
Pending等待—
Retries重试—
Active requests活跃请求—
Waiting / tokens等待 / token—

Issue ledger问题台账

Faults are not silently scored故障不会被静默计分

Completion path完成路径

One owner per attempt每次尝试只有一个归属
Evaluation completion flow ARCHIVED R4 39 SCORED TWO LIVE POOLS 12 + 12 SLOTS CURRENT RUN REPAIR / REGRADE ONLY MISSING REWARDS 315 SCORED

Why the count is not the Harbor terminal count. A model outcome may score zero and still count. A container or verifier fault without a reward does not count and is replayed from the narrowest durable boundary.

为什么这里不直接使用 Harbor 的终态数量。 模型结果即使得 0 分也可以计数; 没有 reward 的容器或 verifier 故障不计数, 并从最窄的持久化边界补跑。

What this public view excludes. Task prompts, agent trajectories, internal addresses, credentials, per-request content, and unpublished benchmark artifacts never enter the feed.

公开看板不会包含什么。 Task prompt、agent trajectory、内部地址、凭据、逐请求内容以及未发布的 benchmark artifact 都不会进入数据源。