At $N=32$, the shortfall of every planner is dominated by collisions (STLPY 34.4%, single-agent diffusion 27.5%, ours 10.9%), whereas the entire STL violation umbrella stays at or below 6.9% for every method. Joint multi-agent sampling keeps the collision channel roughly $3\times$ below both baselines, which is the source of the gap. At $N\le16$ all methods reach at least 86% success and the residual failures are about spec satisfaction rather than safety.
Setting: Dubins Car environment, mixed heterogeneous specification, where each agent samples one component (sequence, branch, cover, loop, or signal) per episode, with randomized predicate locations. The fixed-goal variant of the same sweep, at goal scale 1.0, is milder (STLPY at $N=32$ reaches 85.9% success with 8.8% collisions), so the collapse at $N=32$ is specific to randomizing predicate locations.