HATT v2.0 Shaped Best/Latest Trace 分析
对象:hattv2.0_coalition_delivery_shaped 的 best checkpoint 与 latest checkpoint。trace 输出目录:runs/eval/hattv2_shaped_best_latest_trace_20260703。
结论
latest checkpoint 的退化不是“没有发现任务”,也不是 mask 把动作空间错误屏蔽掉。典型失败场景中,任务已经全部发现,但策略在前 200 多步之后主动滑入 idle/explore 状态,几乎不再选择 delivery/resupply,剩余需求长期不下降,最终超时。
40 / 64best 成功但 latest 失败的场景数
0.674这些失败场景 latest 平均 sent ratio
1.225这些失败场景 latest 平均死亡数
代表场景 Trace
| scenario | checkpoint | len | completed | discovered | sent ratio | dead | final remaining | last demand drop | tail no progress |
|---|---|---|---|---|---|---|---|---|---|
| 15 | best | 638 | 10 | 10 | 1.000 | 0 | 0 | 637 | 0 |
| 15 | latest | 750 | 2 | 10 | 0.226 | 1 | 2594 | 115 | 634 |
| 16 | best | 737 | 10 | 10 | 1.000 | 2 | 0 | 736 | 0 |
| 16 | latest | 750 | 4 | 10 | 0.290 | 3 | 2053 | 190 | 559 |
| 37 | best | 656 | 10 | 10 | 1.000 | 0 | 0 | 655 | 0 |
| 37 | latest | 750 | 4 | 10 | 0.392 | 3 | 2351 | 230 | 519 |
| 53 | best | 541 | 10 | 10 | 1.000 | 0 | 0 | 540 | 0 |
| 53 | latest | 750 | 4 | 10 | 0.443 | 0 | 1532 | 224 | 525 |
| 7 | latest | 630 | 10 | 10 | 1.000 | 1 | 0 | 629 | 0 |
| 29 | latest | 429 | 10 | 10 | 1.000 | 0 | 0 | 428 | 0 |
Hard Cases 的阶段行为
| checkpoint / phase | remaining demand drop | delivery agents | idle agents | explore agents | resupply agents | zero-load agents |
|---|---|---|---|---|---|---|
| best early 0-250 | 1571.3 | 6.12 | 0.00 | 2.32 | 3.56 | 4.81 |
| best mid 250-500 | 1144.8 | 6.02 | 0.05 | 1.88 | 4.06 | 4.98 |
| best late 500-end | 476.8 | 3.56 | 1.88 | 1.76 | 4.31 | 4.39 |
| latest early 0-250 | 1081.5 | 3.12 | 2.80 | 3.69 | 2.39 | 2.69 |
| latest mid 250-500 | 0.0 | 0.00 | 8.66 | 3.25 | 0.09 | 0.08 |
| latest late 500-end | 0.0 | 0.00 | 8.75 | 2.48 | 0.00 | 0.00 |
代码层面对照
_get_available_actions()只按task.discovered and not task.completed生成 mask;idle/resupply 本身是普通合法任务。trace 中的 idle 锁死是策略选择,不是 mask 引导。- shaped reward 对剩余任务时的 idle 有惩罚,对全发现后继续 explore 也有惩罚;但惩罚强度没有压住 late checkpoint 的 idle/explore 偏置。
- best checkpoint 的行为闭环是:delivery 和 resupply 同时维持,remaining demand 持续下降。latest checkpoint 在 hard cases 中前期能送一部分,但中期起 delivery agent 变成 0,配送链断裂。
判断与下一步
- 当前 reward shaping 方向成立,不建议回到旧版能力匹配式 reward,也不建议用 action mask 去“教”策略。
- HATT v2.0 的主要剩余问题是训练后期策略退化。应优先做算法稳定化:降低 PPO 更新强度、加入 KL/clip 监控、熵退火或保存更密集 checkpoint 后做早停选择。
- 奖励侧可以小幅增强“有剩余需求时 idle/explore 的机会成本”,但这只能作为辅助;trace 显示根因更像策略分布漂移到 idle/explore 吸引子。
- 后续实验建议固定 reward shaped,先比较 HATT v2.0 的
clip_param、ppo_epoch、lr、entropy_coef schedule,不要同时大改环境。