Attempt ledger尝试台账
The denominator is fixed at 63 tasks × 5 attempts: Terminal-Bench 4.0 ships 66, and three of them declare gpus = 1, which this node cannot honour while all eight GPUs are serving the model under test. An attempt advances this bar only when a verifier returns a reward; infrastructure failures remain visible but do not consume the target.
分母固定为 63 个任务 × 5 次尝试: Terminal-Bench 4.0 共 66 个任务, 其中 3 个声明了 gpus = 1, 而本节点的 8 张卡都在为被测模型提供服务, 无法满足。 只有 verifier 返回 reward 的尝试才推进这条进度; 基础设施故障保留记录, 但不占用目标次数。
Verified correctnessVerifier 正确率
| Task任务 | Current FP8当前 FP8 | Official GLM-5.3 API官方 GLM-5.3 API | Δ when complete完成项差值 | State状态 |
|---|
Serving pools模型服务池
Issue ledger问题台账
Completion path完成路径
Why the count is not the Harbor terminal count. A model outcome may score zero and still count. A container or verifier fault without a reward does not count and is replayed from the narrowest durable boundary.
为什么这里不直接使用 Harbor 的终态数量。 模型结果即使得 0 分也可以计数; 没有 reward 的容器或 verifier 故障不计数, 并从最窄的持久化边界补跑。
What this public view excludes. Task prompts, agent trajectories, internal addresses, credentials, per-request content, and unpublished benchmark artifacts never enter the feed.
公开看板不会包含什么。 Task prompt、agent trajectory、内部地址、凭据、逐请求内容以及未发布的 benchmark artifact 都不会进入数据源。