Threadmere AI cut diffusion training costs 55% without changing a line of code
“We assumed cutting our training bill in half meant months of re-architecture. Instead we pointed our existing PyTorch stack at a reserved Reviosa cluster and the savings showed up on the very first invoice. The engineering team barely noticed the migration — finance definitely did.”
ML Platform Team
Threadmere AI
Threadmere AI builds generative design tools for fashion brands. Its flagship product turns a mood board and a fabric library into photorealistic garment renders, powered by a family of latent diffusion models fine-tuned on tens of millions of licensed apparel images. Every seasonal drop from a customer means a fresh round of large-scale training — and until last year, a painful conversation about the cloud bill.
The challenge
Threadmere trained on on-demand A100 instances at a hyperscaler, and the economics were working against them. Training a new base checkpoint took roughly three weeks across 32 GPUs, and because capacity was scattered across zones, jobs regularly ran on slower interconnects than the team had benchmarked. Preemptions forced conservative checkpointing, which added overhead of its own. By the time the finance team modeled the next year of roadmap, GPU spend was on track to double.
"We were paying premium prices for capacity that behaved like spot," said Threadmere's VP of Engineering. "Our schedulers spent more time working around the infrastructure than using it."
Why Reviosa
Threadmere evaluated Reviosa's HGX H100 nodes — 8×H100 SXM with NVLink and 3.2 Tbps InfiniBand at $19.92/hr on demand, an effective $2.49 per GPU per hour. A two-day proof of concept reproduced their largest training run on four nodes with no code changes: same containers, same PyTorch FSDP configuration, same experiment tracking.
The benchmark results settled the decision. NVLink within each node and InfiniBand between them lifted scaling efficiency from 71% on their previous setup to 94%, cutting per-epoch wall-clock time by more than half. Threadmere signed a reserved commitment for four full nodes — 32 H100s — at a committed rate below the on-demand price, with the option to burst additional on-demand H100s at $2.69/hr during seasonal peaks.
"The proof of concept was the whole sales cycle. We rsynced our training container over on a Tuesday, and by Thursday we had numbers that made the decision for us." — VP of Engineering, Threadmere AI
Reviosa's engineering team also helped Threadmere move its 60 TB dataset into object storage colocated with the cluster, eliminating the cross-region egress fees that had quietly become one of the largest line items on the old bill.
The results
The migration took one sprint. Measured across the first two full training cycles on Reviosa:
- 55% lower training cost per base checkpoint, combining the reserved rate, better scaling efficiency, and zero egress fees.
- 2.4x faster epochs, shrinking a full training run from three weeks to under nine days.
- Zero preemptions on reserved capacity, which let the team relax checkpoint frequency and reclaim roughly 6% of every run previously lost to save/restore overhead.
The strategic payoff is bigger than the invoice. Because a full retrain now fits inside a two-week sprint, Threadmere moved from two base-model refreshes per year to a monthly cadence, and customer-specific fine-tunes that once queued for days now start within the hour on burst capacity. The roadmap that finance flagged as unaffordable is now shipping — on the same budget line Threadmere had a year ago.
The results, by the numbers
55%
lower training cost
32
reserved H100 GPUs
2.4x
faster epoch time
More teams on Reviosa
Quill & Query serves contract analysis at sub-300ms p95 — and pays nothing when lawyers sleep
Legal Tech287ms
p95 response latency
Atlas Weather Labs improved 48-hour forecast accuracy 18% with nightly multi-node fine-tunes
Climate & Weather Intelligence18%
better 48h forecast accuracy
Helixmoor Bio compressed a 6-month screening campaign into 3 weeks on burst H200 capacity
Biotech / Drug Discovery2M
compounds screened
Start training in minutes
Create an account, add a card, and launch your first GPU instance. Per-second billing means you only pay for what you use.
No minimum commitment · Cancel anytime · $10 free credit for new accounts