mirror of
https://github.com/Crosstalk-Solutions/project-nomad.git
synced 2026-07-28 11:14:37 +02:00
Phase 4 of the NOMAD Score v2 effort (stacked on #1089). Captures the raw channel values the leaderboard scores from, computes the uncapped v2 score in-app (byte-matching the leaderboard's server-side recompute), submits the v2 payload, and surfaces it in the benchmark UI. - Capture raws: per-channel sysbench values + O_DIRECT disk (W4), single/multi thread CPU + memory, W6 consistency companions, and run-environment metadata (run_environment, storage_path_type, gpu_compute_detected, #1016). - AI hardening: reference model llama3.1:8b, num_predict=256, VRAM eviction before the run + unload own model after, pre-flight disk check (~6.5GB, only when the model is uncached), median-of-3 run (W7). - Score: _calculateNomadScoreV2 (uncapped, reference->1000) byte-matches the leaderboard recompute. nomad_score_v2 stays null unless a full run with AI. - Payload: submitToRepository sends the v2 raws + score alongside legacy v1. - UI: v2 headline score, legacy v1 as secondary, raws/env in details. - Migration: 13 nullable columns; double (not knex float(8,2)) so millions-scale raws keep full precision and don't break the leaderboard byte-match. Reference constant REFERENCE_SCORES_V2.ai_tokens_per_second is still the stale 40.5 placeholder (TODO in source); locking it to the measured 13.2 is a follow-up gated on the AMD iGPU fix (#1074). Live prod v2 submit is separately blocked by the leaderboard TTFT floor (tracked in the leaderboard follow-up). Verified on NOMAD3 (RTX 5060) dev env: migration applies, raws recorded, v2 score computed, backend + benchmark.tsx typecheck clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| migrations | ||
| seeders | ||