mirror of
https://github.com/Crosstalk-Solutions/project-nomad.git
synced 2026-07-29 11:44:38 +02:00
The AI benchmark evicted resident models (forcing the benchmark model cold) and then went straight into the timed median-of-N loop with no warm-up. On a cold box the model-load + GPU spin-up cost landed inside the timed runs (observed: 173s TTFT / 5.83 tok/s vs a ~80 tok/s warm steady-state), and consecutive runs weren't isolated (a prior run left the model warm). Because the AI channel is uncapped and ~30% of the composite, the same machine could post a ~2x-different NOMAD Score depending on warm/cold state (888 vs 1958 observed back-to-back). Add one discarded warm-up inference after eviction and before the timed loop so every timed run measures warm, steady-state throughput. Cold and warm invocations now converge on the same score. Best-effort: a warm-up hiccup never fails the run. Closes #1139 |
||
|---|---|---|
| .. | ||
| controllers | ||
| data | ||
| exceptions | ||
| jobs | ||
| middleware | ||
| models | ||
| services | ||
| utils | ||
| validators | ||