Nebula/AI Optimizer

AI Optimizer

Live1s
JC
Preview module
Nebula's optimizer learns per-workload signatures and proposes placement, power-cap, and thermal-envelope adjustments.
Reduce atlas-01 mean temperature by 3.2 °C
Recommendation

Rebalance training-job-8817 from R02 to R06. Estimated 4% throughput cost, 12% fewer thermal alerts.

Cap orion-02 GPU power at 620 W
Recommendation

Inference batch p95 latency changes < 1.8 ms. Saves ~38 kW at peak.

Consolidate idle inference replicas
Recommendation

12 nodes averaging < 30% util over 24h. Estimated 9 nodes reclaimable for training bursts.

Preempt vega-03-n004 checkpoint IO contention
Recommendation

Move checkpoint destination to /mnt/scratch; est. 18% write latency reduction.