Atlas Weather Labs improved 48-hour forecast accuracy 18% with nightly multi-node fine-tunes
“A forecast model you retrain quarterly is always slightly wrong about the present. Nightly fine-tuning on fresh observations was the obvious fix — we just couldn't afford the cluster to do it. Reviosa made the nightly run cheaper than our old quarterly one.”
Dr. Samuel Ibarra
Chief Science Officer, Atlas Weather Labs
Atlas Weather Labs sells high-resolution weather intelligence to energy traders, logistics networks, and agricultural insurers. Its core product is an ML weather model — a graph neural network trained on decades of reanalysis data — that produces 48-hour regional forecasts at kilometer scale. In markets where a wind-generation forecast moves real money, forecast skill is the entire product.
The challenge
Atlas had a model that was good and a training cadence that made it worse every day it aged. The team fine-tuned quarterly on accumulated observations, because that was all their four-GPU on-prem setup could support: each fine-tune took eleven days, monopolizing the hardware researchers needed for experiments. Between refreshes, the model drifted — seasonal transitions, evolving observation networks, and shifting sensor coverage all degraded skill in ways the next quarterly run only partially recovered.
The science team's own ablations showed what continuous adaptation was worth: a model fine-tuned nightly on the previous day's global observations would deliver a double-digit accuracy improvement at the 48-hour horizon. What they lacked was a way to run a multi-node training job every single night without buying a cluster that would sit idle twenty hours a day.
Why Reviosa
Reviosa's HGX H100 nodes gave Atlas the arithmetic it needed. Each night at 02:10 UTC, after the day's observation feeds close, an orchestration job launches four 8×H100 nodes — 32 GPUs on 3.2 Tbps InfiniBand at $18.32/hr per node — pulls the latest checkpoint and fresh observations from object storage, and runs a distributed fine-tune. The job completes in under four hours, publishes the updated checkpoint to the inference fleet, and terminates every instance before the morning forecast cycle begins.
The nightly run costs roughly $290 in compute — about $106,000 a year for 365 training runs. The equivalent owned cluster would have cost several times that annually in hardware amortization, power, and staffing, while sitting idle between runs.
"InfiniBand was the detail that made it work. Our graph network is communication-heavy, and on commodity interconnects the nightly job wouldn't fit the window. On Reviosa's HGX nodes it finishes with ninety minutes to spare." — Dr. Samuel Ibarra, Chief Science Officer
The results
Atlas has now run the nightly pipeline for over a year, and the verification statistics are unambiguous:
- 18% improvement in 48-hour forecast accuracy across its verification suite, measured against the quarterly-refresh baseline over two full seasons.
- 365 training runs per year instead of four, eliminating the model drift that used to build up between quarterly refreshes.
- 4-hour nightly window from data close to published checkpoint, with capacity released completely between runs.
The accuracy gain converted directly into revenue. Two energy-trading customers expanded contracts after Atlas's wind-ramp forecasts beat their incumbent vendor in a head-to-head trial — a trial the science team credits to the nightly cadence, since ramp events are exactly where stale models fail. Churn among logistics customers fell to zero over the year.
Just as valuable, the on-prem cluster went back to being a research machine. Freed from production training, Atlas's scientists now run three times as many experiments, and promising architecture changes graduate to a full-scale Reviosa training run within days instead of waiting for a quarterly slot. The company's next model generation — currently training on burst H200 capacity — exists because the hardware finally stopped being the bottleneck.
The results, by the numbers
18%
better 48h forecast accuracy
365
training runs per year
4hrs
nightly training window
More teams on Reviosa
Loomline AI cut diffusion training costs 55% without changing a line of code
Generative AI / Fashion55%
lower training cost
Helixon Bio compressed a 6-month screening campaign into 3 weeks on burst H200 capacity
Biotech / Drug Discovery2M
compounds screened
Northfork Render finished a feature film on 400 on-demand L40S GPUs — with zero hardware owned
Media & Entertainment / VFX400
L40S GPUs at peak
Start training in minutes
Create an account, add a card, and launch your first GPU instance. Per-second billing means you only pay for what you use.
No minimum commitment · Cancel anytime · $10 free credit for new accounts