Loomline AI cut diffusion training costs 55% without changing a line of code
“We assumed cutting our training bill in half meant months of re-architecture. Instead we pointed our existing PyTorch stack at a reserved Reviosa cluster and the savings showed up on the very first invoice. The engineering team barely noticed the migration — finance definitely did.”
Mara Ellison
VP of Engineering, Loomline AI
Loomline AI builds generative design tools for fashion brands. Its flagship product turns a mood board and a fabric library into photorealistic garment renders, powered by a family of latent diffusion models fine-tuned on tens of millions of licensed apparel images. Every seasonal drop from a customer means a fresh round of large-scale training — and until last year, a painful conversation about the cloud bill.
The challenge
Loomline trained on on-demand A100 instances at a hyperscaler, and the economics were working against them. Training a new base checkpoint took roughly three weeks across 32 GPUs, and because capacity was scattered across zones, jobs regularly ran on slower interconnects than the team had benchmarked. Preemptions forced conservative checkpointing, which added overhead of its own. By the time the finance team modeled the next year of roadmap, GPU spend was on track to double.
"We were paying premium prices for capacity that behaved like spot," said Mara Ellison, Loomline's VP of Engineering. "Our schedulers spent more time working around the infrastructure than using it."
Why Reviosa
Loomline evaluated Reviosa's HGX H100 nodes — 8×H100 SXM with NVLink and 3.2 Tbps InfiniBand at $18.32/hr on demand, an effective $2.29 per GPU per hour. A two-day proof of concept reproduced their largest training run on four nodes with no code changes: same containers, same PyTorch FSDP configuration, same experiment tracking.
The benchmark results settled the decision. NVLink within each node and InfiniBand between them lifted scaling efficiency from 71% on their previous setup to 94%, cutting per-epoch wall-clock time by more than half. Loomline signed a reserved commitment for four full nodes — 32 H100s — at a committed rate below the on-demand price, with the option to burst additional on-demand H100s at $2.49/hr during seasonal peaks.
"The proof of concept was the whole sales cycle. We rsynced our training container over on a Tuesday, and by Thursday we had numbers that made the decision for us." — Mara Ellison, VP of Engineering
Reviosa's engineering team also helped Loomline move its 60 TB dataset into object storage colocated with the cluster, eliminating the cross-region egress fees that had quietly become one of the largest line items on the old bill.
The results
The migration took one sprint. Measured across the first two full training cycles on Reviosa:
- 55% lower training cost per base checkpoint, combining the reserved rate, better scaling efficiency, and zero egress fees.
- 2.4x faster epochs, shrinking a full training run from three weeks to under nine days.
- Zero preemptions on reserved capacity, which let the team relax checkpoint frequency and reclaim roughly 6% of every run previously lost to save/restore overhead.
The strategic payoff is bigger than the invoice. Because a full retrain now fits inside a two-week sprint, Loomline moved from two base-model refreshes per year to a monthly cadence, and customer-specific fine-tunes that once queued for days now start within the hour on burst capacity. The roadmap that finance flagged as unaffordable is now shipping — on the same budget line Loomline had a year ago.
The results, by the numbers
55%
lower training cost
32
reserved H100 GPUs
2.4x
faster epoch time
More teams on Reviosa
PixelPatch scaled AI photo enhancement to 4M requests a day while cutting infra spend 40%
Consumer Apps4M
requests per day
Helixon Bio compressed a 6-month screening campaign into 3 weeks on burst H200 capacity
Biotech / Drug Discovery2M
compounds screened
Northfork Render finished a feature film on 400 on-demand L40S GPUs — with zero hardware owned
Media & Entertainment / VFX400
L40S GPUs at peak
Start training in minutes
Create an account, add a card, and launch your first GPU instance. Per-second billing means you only pay for what you use.
No minimum commitment · Cancel anytime · $10 free credit for new accounts